Software-Defined Network Solution System
The software-defined network solution system addresses the complexity of dynamic network environments by integrating physical and virtual resources for efficient, automated policy application and fault recovery, enhancing operational efficiency and stability.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- WOORIDULIT CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-07-27
AI Technical Summary
Conventional network technologies struggle with the complexity of applying network policies quickly and consistently in dynamic environments involving virtual machines and containers, leading to increased operational burden and complexity in fault response and security policy application.
A software-defined network solution system that includes a management server for application-centric network control, integrating physical and virtual network resources, with features like automatic policy application, fault detection, and recovery, and centralized management.
Enhances operational efficiency, network stability, scalability, and flexibility by automating policy application and fault recovery, while ensuring consistent security and traffic control across dynamic environments.
Smart Images

Figure 112026012801880-PAT00006_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a software-defined network solution system. Background Technology
[0002] With the proliferation of data centers and cloud environments, networks are moving away from static configurations centered on physical equipment and evolving into structures that change dynamically at the application level, such as virtual machines and containers.
[0003] However, conventional network technology relied on manually configuring settings and policies at the individual device level, such as switches and routers, which presented a problem in that it was difficult to apply network policies quickly and consistently when applications were created, moved, or expanded.
[0004] Furthermore, as physical networks and virtualization environments are operated separately, fault response, security policy application, and traffic routing have become more complex, and the operational burden has increased. In such environments, there is a need for new network control technologies that enable integrated central network control, application-centric policy definition, and automated network operation by linking physical and virtual network resources.
[0005] Meanwhile, the aforementioned background technology is technical information that the inventor possessed for the derivation of the present invention or acquired during the process of deriving the present invention, and it cannot be considered as prior art disclosed to the general public prior to the filing of the present invention. Prior art literature
[0006] Korean Registered Patent No. 100977488 The problem to be solved
[0007] The objective of the present invention is to provide a software-defined network solution system that defines network policies based on application units in a data center environment and centrally controls the physical network fabric and the virtualization environment.
[0008] The technical problems of the present invention are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by those skilled in the art from the description below. means of solving the problem
[0009] A software-defined network solution system according to one embodiment of the present invention may include a management server that provides application-centric network control in a data center environment.
[0010] A software-defined network solution system according to one embodiment of the present invention may further include a router, which is one of a plurality of components that perform the role of a boundary node responsible for traffic between subnets or interaction with an external network within the data center, and constitute a physical network fabric in which physical network equipment is interconnected to form a data transmission path of the data center network.
[0011] According to one embodiment, the management server may be characterized by controlling the data transmission path, traffic flow, and policy application status of the physical network fabric by transmitting control commands to a plurality of network devices constituting the physical network fabric.
[0012] According to one embodiment, a management server may include: an application-centric policy unit that defines and applies network policies including communication allowance, blocking, priority, or security rules based on application units on a network of a plurality of network devices constituting the physical network fabric; a central control node unit that interprets the network policies defined by the application-centric policy unit and transmits commands to the physical network fabric and network devices connected to the physical network fabric so that settings are converted to the interpreted network policies; a virtualization linkage unit that links with a public cloud environment including virtual machines, containers, and virtual network functions, and automatically applies the network policies defined by the application-centric policy unit when the virtual machine or container is created, deleted, or moved; and an operation management unit that detects and collects the operational status of a plurality of network devices constituting the physical network fabric, the central control node unit, and the network devices in real time, detects whether a failure has occurred in the network devices and provides a notification or performs an automatic recovery operation, and manages expansion by automatically applying policies when network resources are added or changed.
[0013] According to one embodiment, the application-centric policy unit may be characterized by being configured to automatically select and apply a network policy corresponding to the application when the application is deployed or a service instance is created.
[0014] According to one embodiment, the operation management unit may be characterized by being configured to monitor network traffic, connection status, or equipment status in real time to detect an abnormal state, and when an abnormal state is detected, to provide a notification to an administrator or to automatically perform at least one of setting a bypass path, switching backup equipment, or reapplying a policy according to a predefined recovery policy.
[0015] According to one embodiment, the operation management unit may be characterized by selecting a recovery scenario based on a detected abnormal state among predefined recovery policies and transmitting a setting change command to the corresponding network equipment.
[0016] According to one embodiment, the operation management unit may be characterized by being configured to automatically apply a previously defined network policy to the added resource and perform load balancing or path reconfiguration when network equipment, a server, a virtual machine, or a container is added.
[0017] According to one embodiment, the path reconstruction may be characterized by collecting traffic and link status of the network equipment, calculating a new transmission path using a path calculation algorithm based on the collected information, and reconstructing the data transmission path by updating the forwarding information or flow rules of the network equipment through a central control node according to the calculated path.
[0018] According to one embodiment, the application-centric policy unit comprises a difficulty score (D) output from an application method determination algorithm. way It can be characterized by reducing the possibility of misdistribution, policy conflicts, and large-scale impact spread by quantifying the difficulty level at the pre-policy application stage for network policies defined based on ) and selecting the application procedure to apply complex and high-risk policies more conservatively.
[0019] According to one embodiment, the application method determination algorithm comprises the number of policy rules (N r The more ) increases, the more the difficulty score (D way The number of policy rules (N) increases so that ) r After adding 1 to ) and inputting it into the logarithmic function, the difficulty adjustment weight (w) to the result of the logarithmic function way Includes the first element configured to multiply by ), and the number of special conditions (N x The more ) increases, the more the difficulty score (Dway The number of special conditions (N) increases so that ) x After adding 1 to ) and inputting it into the logarithmic function, the result of the logarithmic function is 1 to the difficulty adjustment weight (w way Includes a second element configured to multiply the result of subtracting ), and the number of applicable targets (N e The more ) increases, the more the difficulty score (D way To increase the number of applicable targets (N) e Includes a third element configured to input into a logarithmic function by adding 1 to ), and the number of policy changes (N c The more ) increases, the more the difficulty score (D way Number of policy changes (N) to increase ) c Includes a fourth element configured to input into a logarithmic function by adding 1 to ), and operational importance (U way ) difficulty score (D way It may include a fifth element configured to handle important service policies more conservatively, even with the same complexity, in addition to ) directly.
[0020] According to one embodiment, the application method determination algorithm comprises the number of policy rules (N r ), number of the above special conditions (N x ), the number of applicable targets mentioned above (N e ) and the number of times the above policy was changed (N c It can be characterized by applying ) to a logarithmic function so that score changes are reflected smoothly even if the range of values for each factor rises or falls sharply.
[0021] According to one embodiment, the application method determination algorithm comprises the number of policy rules (N r ) and the number of the above special conditions (N x Although ) is a factor of the policy item, the difficulty adjustment weight (w) is designed to reflect changes in the contribution to operational risk. way It can be characterized by being configured to adjust the contribution using ). Effects of the invention
[0022] According to one aspect of the present invention described above, the software-defined network solution system proposed by the present invention can improve operational efficiency by automatically applying network policies according to application deployment and service instance changes, and can increase network stability by automating fault detection and recovery.
[0023] In addition, network scalability and flexibility are enhanced through a centralized control structure that links physical networks and virtual resources, and consistent policy-based security and traffic control can be enabled.
[0024] The effects of the present invention are not limited to those mentioned above, and various effects may be included within the scope obvious to a person skilled in the art from the contents described below. Brief explanation of the drawing
[0025] FIG. 1 is a conceptual diagram of a software-defined network solution system according to one embodiment of the present invention. FIG. 2 is a conceptual diagram of a management server according to one embodiment of the present invention. Specific details for implementing the invention
[0026] The following detailed description of the invention refers to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention. It should be understood that various embodiments of the invention are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the invention in relation to one embodiment.
[0027] When it is stated that one component is "connected" or "contracted" to another component, it should be understood that while it may be directly connected or contracted to that other component, there may also be other components in between. Conversely, when it is stated that one component is "directly connected" or "directly contracted" to another component, it should be understood that there are no other components in between.
[0028] Furthermore, it should be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the invention. Accordingly, the following detailed description is not intended to be taken in a limiting sense, and the scope of the invention is limited only by the appended claims, including all equivalents thereof, provided appropriately described. Similar reference numerals in the drawings refer to the same or similar functions across various aspects.
[0029] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the drawings.
[0031] FIG. 1 is a conceptual diagram of a software-defined network solution system according to one embodiment of the present invention.
[0032] Referring to FIG. 1, a software-defined network solution system according to one embodiment of the present invention may include a management server (100) and a router (500).
[0033] The management server (100) can control the data transmission path, traffic flow, and policy application status of the physical network fabric by transmitting control commands to a plurality of network devices that constitute the physical network fabric.
[0034] The management server (100) can define communication allowance, blocking, and security policies on an application-by-application basis, convert this into network equipment settings, and can automatically apply policies when resources change by linking with virtual machine and container environments.
[0035] In addition, the management server (100) can monitor the network and equipment status in real time and provide notifications or perform automatic recovery and path reconfiguration in the event of a failure.
[0036] A router (500) acts as a boundary node responsible for traffic between subnets or interaction with an external network within the data center, and may represent one of a plurality of components that constitute a physical network fabric in which physical network equipment is interconnected to form a data transmission path for the data center network.
[0037] The management server (100) and the router (500) may be their own servers or cloud servers for providing services according to the present invention, or they may be a set of distributed nodes in a peer-to-peer (P2P) manner.
[0038] The management server (100) can perform one or more of the operations, storage, reference, input / output, and control functions of a general computer, and may include an artificial neural network described later based on input data.
[0039] The management server (100) may include a processor and memory. The processor may include devices capable of analyzing the operating status of the exercise equipment (500) according to the present invention, the user's bio-response, and fluctuations in exercise load in real time, and transmitting control commands. The processor may execute a program or control the management server (100). Program code executed by the processor may be stored in memory. The memory may store relevant information for performing a service according to the present invention or a program for implementing a method. The memory may be volatile memory or non-volatile memory.
[0040] The management server (100) can send data to an external device or receive data from an external device using a network.
[0041] The management server (100) can train an artificial neural network and can also use an artificial neural network that has been trained. The processor can train or execute an artificial neural network stored in memory, and the memory can store an artificial neural network that has been trained. The electronic device that trains the artificial neural network and the electronic device that uses it may be the same, but they may also be separate.
[0042] Artificial intelligence is a computer system that partially implements the functions of the human brain and is capable of learning, speculating, and making judgments on its own. As learning progresses, the probability of extracting the correct answer can increase. Artificial intelligence can be composed of learning and component technologies that utilize it. The learning aspect of AI is an algorithmic technology that classifies and learns features based on input data, while the component technologies may be techniques that utilize these learning algorithms to partially implement the functions of the human brain.
[0043] Artificial intelligence is a technology that facilitates the approach to problems where multiple probabilistic answers are possible, enabling it to logically and probabilistically infer optimal cycles, methods, and plans based on input data. AI inference techniques can include evaluating input data, optimization prediction, knowledge and probability-based reasoning, and preference-based planning.
[0044] Artificial neural networks are learning algorithms in the field of machine learning that programmatically implement the connections between neurons and synapses in the brain. By creating a neural network structure through programming and then training it, artificial neural networks can acquire desired functions. Although errors may exist, they can learn from massive datasets to produce appropriate output data from input data. They have the advantage of being able to obtain output data that has yielded statistically good results and are similar to human reasoning.
[0045] The management server (100) can infer individual characteristics and interests by analyzing consumers' online behavior data, social media activities, search history, etc., using an artificial intelligence algorithm built based on big data, and may include a number of pre-trained artificial neural networks for this purpose.
[0046] The network is a high-speed backbone network of a large-scale communication network capable of high-capacity, long-distance voice and data services, and may be a next-generation wired and wireless network for providing the Internet or high-speed multimedia services.
[0047] If the network is a mobile communication network, it may be a synchronous mobile communication network or an asynchronous mobile communication network. As an example of an asynchronous mobile communication network, a WCDMA (Wideband Code Division Multiple Access) network may be cited. In this case, although not shown in the drawing, the network may include an RNC (Radio Network Controller). Meanwhile, although a WCDMA network was given as an example, it may be a 3G LTE network, a 4G network, a next-generation communication network such as 5G, or other IP-based IP networks.
[0048] The management server (100) and router (500) can include any terminal capable of exchanging data over a network, such as a desktop computer, laptop, tablet, or smartphone.
[0049] The management server (100) and the router (500) may include one or more of the computational function, storage function, reference function, input / output function, and control function of a computer to perform the service according to the present invention.
[0050] The management server (100) and the router (500) may access a website or install an application to receive the service according to the present invention. The management server (100) and the router (500) may exchange data through the website or the application.
[0051] The network is a high-speed backbone network of a large-scale communication network capable of high-capacity, long-distance voice and data services, and may be a next-generation wired and wireless network for providing the Internet or high-speed multimedia services.
[0052] If the network is a mobile communication network, it may be a synchronous mobile communication network or an asynchronous mobile communication network. As an example of an asynchronous mobile communication network, a WCDMA (Wideband Code Division Multiple Access) network may be cited. In this case, although not shown in the drawing, the network (300) may include an RNC (Radio Network Controller). Meanwhile, although a WCDMA network was given as an example, it may be a 3G LTE network, a 4G network, a 5G network, or other next-generation communication networks, or other IP-based IP networks.
[0053] A system (1) according to one embodiment of the present invention can define network policies based on application units in a data center environment and centrally control the physical network fabric and the virtualization environment, thereby resolving the complexity of network configuration and operation and enabling automated network management.
[0055] FIG. 2 is a conceptual diagram of a management server according to one embodiment of the present invention.
[0056] Referring to FIG. 2, a management server (100) according to one embodiment of the present invention may include an application-centric policy unit (110), a central control node unit (130), a virtualization linkage unit (150), and an operation management unit (170).
[0057] The application-centric policy unit (110) can define and apply network policies including communication allowance, blocking, priority, or security rules based on application units on a network of multiple network devices that make up a physical network fabric.
[0058] In addition, the application-centric policy unit may be characterized by being configured to automatically select and apply a network policy corresponding to the application when deploying the application or creating a service instance.
[0059] Meanwhile, the application-centric policy department (110) has a difficulty score (D) output from the application method determination algorithm. way It can be characterized by reducing the possibility of misdistribution, policy conflicts, and large-scale impact spread by quantifying the difficulty level at the pre-policy application stage for network policies defined based on ) and selecting the application procedure to apply complex and high-risk policies more conservatively.
[0060] The above-described algorithm for determining the application method is based on the number of policy rules (N r The more ) increases, the more the difficulty score (D way The number of policy rules (N) increases so that ) r After adding 1 to ) and inputting it into the logarithmic function, the difficulty adjustment weight (w) to the result of the logarithmic function way Includes the first element configured to multiply by ), and the number of special conditions (N x The more ) increases, the more the difficulty score (D way The number of special conditions (N) increases so that ) x After adding 1 to ) and inputting it into the logarithmic function, the result of the logarithmic function is 1 to the difficulty adjustment weight (w way Includes a second element configured to multiply the result of subtracting ), and the number of applicable targets (N e The more ) increases, the more the difficulty score (D way To increase the number of applicable targets (N) e Includes a third element configured to input into a logarithmic function by adding 1 to ), and the number of policy changes (N c The more ) increases, the more the difficulty score (D way Number of policy changes (N) to increase ) c Includes a fourth element configured to input into a logarithmic function by adding 1 to ), and operational importance (U way ) difficulty score (D way It may include a fifth element configured to handle important service policies more conservatively, even with the same complexity, in addition to ) directly.
[0061] In addition, the algorithm for determining the application method is the number of policy rules (N) mentioned above. r ), number of the above special conditions (N x ), the number of applicable targets mentioned above (N e ) and the number of times the above policy was changed (N c It can be characterized by applying ) to a logarithmic function so that score changes are reflected smoothly even if the range of values for each factor rises or falls sharply.
[0062] In addition, the algorithm for determining the application method is the number of policy rules (N) mentioned above. r ) and the number of the above special conditions (N x Although ) is a factor of the policy item, the difficulty adjustment weight (w) is designed to reflect changes in the contribution to operational risk. way It can be characterized by being configured to adjust the contribution using ).
[0063] More specifically, the application-centric policy unit (110) quantifies the difficulty level of a defined network policy at the pre-policy application stage and uses a difficulty score (D) to be used as a judgment indicator for selecting the application procedure. way You can utilize an algorithm that determines the application method to output ).
[0064] More specifically, the application-centric policy unit (110) stores the policy in a defined policy data structure, and then counts the rule / exception / target information in the stored policy data structure in a defined manner to determine the number of policy rules (N r ), number of special conditions (N x ), number of applicable targets (N e ) can be derived.
[0065] In addition, the application-centric policy unit (110) has a policy change count (N) from the history of network policy changes. c It can derive ) and operational importance (U) directly entered through the policy registration UI way ) and difficulty adjustment weights (w way) can be derived.
[0066] Finally, the application-centric policy department (110) outputs a difficulty score (D) through an application method determination algorithm. way Based on ), a difficulty level such as low, medium, or high can be determined and transmitted to the central control node (130).
[0067] The application-centric policy department (110) can reduce the possibility of misdistribution, policy conflicts, and large-scale spread of impact by using an algorithm to determine the application method, thereby applying complex and high-risk policies more conservatively.
[0068] In addition, since verification intensity can be differentiated based on difficulty scores, simple policies can be deployed quickly and safeguards for complex policies can be strengthened, thereby improving both operational efficiency and stability.
[0069] The algorithm for determining the application method may be characterized by being configured using input values for judgment and control that determine the policy application method (immediate application, dry run, phased application, approval request, etc.).
[0070] More specifically, the algorithm for determining the application method uses a difficulty score (D) for application-level network policies. way It can be characterized by being designed as a summation structure that adds a total of five elements to output ).
[0071] The first element is the number of policy rules (N r The more ) increases, the more the difficulty score (D way The number of policy rules (N) increases so that ) r After adding 1 to ) and inputting it into the logarithmic function, the difficulty adjustment weight (w) to the result of the logarithmic function way It can be characterized by being configured to multiply )
[0072] The second element is the number of special conditions (N x The more ) increases, the more the difficulty score (D wayThe number of special conditions (N) increases so that ) x After adding 1 to ) and inputting it into the logarithmic function, the result of the logarithmic function is 1 to the difficulty adjustment weight (w way It can be characterized by being configured to multiply the result of subtracting ).
[0073] The third element is the number of applicable targets (N e The more ) increases, the more the difficulty score (D way To increase the number of applicable targets (N) e It can be characterized by being configured to input into a logarithmic function by adding 1 to ).
[0074] The fourth factor is the number of policy changes (N c The more ) increases, the more the difficulty score (D way Number of policy changes (N) to increase ) c It can be characterized by being configured to input into a logarithmic function by adding 1 to ).
[0075] The fifth element is operational importance (U), a value directly entered by the manager. way ) difficulty score (D way It can be characterized by being configured to handle important service policies more conservatively, even with the same complexity, by adding directly to ).
[0076] The number of each policy rule mentioned above (N r ), number of special conditions (N x ), number of applicable targets (N e ), Number of policy changes (N c The reason for adding 1 before inputting ) into the logarithmic function is to prevent an operation error from occurring when the input value of the logarithmic function becomes 0.
[0077] In addition, the first to fourth elements of the application method determination algorithm may be characterized by using a logarithmic function to prevent the score from increasing too rapidly even as each number increases, thereby preventing a single specific variable from excessively dominating the score.
[0078] Since policy difficulty can result from the combination of different factors such as "many rules," "many exceptions," "large scope of application," "frequent recent changes," and "it is a critical service," the algorithm for determining the application method separates each factor into four elements and then assigns a single score, the difficulty score (D way It can be output as ).
[0079] In addition, since the range of values for each factor can vary significantly (e.g., thousands of endpoints, several changes), the algorithm for determining the application method may be characterized by applying a logarithmic function instead of simply summing each factor, so that the score increase is gradual even when the values are large.
[0080] Also, the number of policy rules (N r ), number of special conditions (N x Although both are the number of policy items, since they are factors that can vary the contribution to operational risk, the difficulty adjustment weight (w way It can be characterized by being reflected in the application method determination algorithm to adjust the contribution by utilizing ).
[0081] Also, operational importance (U way Since it may not be expressed solely by the policy data structure, it can be characterized by the addition of values entered by the administrator and the reflection in the application method determination algorithm so that important service policies are classified as having a higher difficulty level.
[0082] The algorithm for determining the application method uses simple summation for the difficulty score (D) in large-scale policies. way Since ) can become excessively large, utilize a logarithmic function for each factor to calculate the difficulty score (D way It can be characterized by being configured to mitigate the runaway of ).
[0083] As described in [Table 1] below and the above description, the algorithm for determining the application method may be characterized by being designed using a logarithmic function through examples of various tests, and may be characterized by being designed to reflect actual data distribution and experimental experience.
[0084] [Table 1]
[0085]
[0086] The number of policy rules (N), each variable used in the algorithm for determining the application method r ), number of special conditions (N x ), number of applicable targets (N e ), Number of policy changes (N c ), operational importance (U way ), difficulty adjustment weight (w way The characteristics of ) are as follows.
[0087] Number of policy rules (N) r ) can mean a value indicating the total number of rule entries within a specific application policy stored in the application-centered policy unit (110), and the unit can be used as the number.
[0088] The application-centric policy unit (110) can store rule entries related to each network policy in the form of a policy table or a policy document (JSON), and can define a rule list field containing rules within the stored structure, and count the number of network policies contained in this rule list field to determine the number of policy rules (N). r ) can be derived.
[0089] Here, a rule entry may refer to a rule that allows communication, a rule that blocks communication, a rule that specifies priority, or a security rule that includes port and protocol conditions.
[0090] The application-centric policy unit (110) may mean treating a unit of rule entry that is applied independently in the network policy as one.
[0091] The application-centric policy unit (110) counts the number of policy rules (N) when there are no rule entries in the rule list field. r ) can be derived as 0, and whenever a rule entry is added or deleted, the number of rule entries in the rule list field is recounted to determine the number of policy rules (N r ) can be updated.
[0092] For example, if 12 rule entries are stored in the rule list field, the number of policy rules (N r You can set ) to 12, which allows you to identify policies with many rules in advance and strengthen pre-application verification, thereby reducing operational risk.
[0093] The algorithm for determining the application method is based on the number of policy rules (N) to reflect the fact that the possibility of conflicts between rules and review costs increase as the number of rule entries increases. r It can be characterized by using ) as an input value.
[0094] Number of special conditions (N) x ) can represent a value indicating the total number of exception or special condition entries that are applied differently from the basic rules, and can be used as a unit of count.
[0095] The application-centered policy unit (110) can define a field containing exception entries when storing rule entries related to each network policy in the form of a policy table or a policy document (JSON).
[0096] The application-centered policy department (110) counts the number of exception entries contained in the fields contained in the exception entry and the number of special conditions (N x ) can be derived.
[0097] Exception entries can refer to items that partially modify the application targets of the default rule, such as blocking only specific instances, allowing only specific subnets, or applying different rules during specific time periods.
[0098] The application-centric policy department (110) has no exception entries, and the number of special conditions (N) x ) can be derived as 0, and the count is recounted whenever an exception entry is added or deleted to determine the number of special conditions (N x ) can be updated.
[0099] Since the algorithm for determining the application method can cause confusion or increase the possibility of rule conflicts during the process where exception entries understand and operate the policy, the number of special conditions (N x It can be characterized by using ) as an input value.
[0100] For example, if 3 exception entries are stored in a field containing exception entries, the number of special conditions (N x You can set ) to 3, which can encourage more conservative deployment of policies with many exceptions.
[0101] Number of applicable targets (N) e ) can mean a value indicating the total number of endpoints to which the application policy applies, and can be used as a unit of count.
[0102] Here, an endpoint can refer to the target unit to which a policy is actually applied, such as a virtual machine, container / pod, or service instance.
[0103] The virtualization linkage unit (150) can query a system that provides a list of VMs or containers (e.g., a cloud management API or an orchestrator API) to retrieve a list of targets belonging to the corresponding application based on an application identifier (label, tag, service name, etc.), and the application-centric policy unit (110) can receive a list of endpoint IDs obtained as a result of the query (e.g., instance ID, Pod UID list, etc.) as input, and count the number of items included in the list to determine the number of applicable targets (N e It can be derived as ).
[0104] Since there may be no targets before deployment, the number of applicable targets (N) e ) can be derived as 0, and the same query procedure is repeated whenever an event occurs where an endpoint is created, deleted, or moved to determine the number of applicable targets (N e ) can be derived.
[0105] Since the application method determination algorithm may increase the burden of verification and rollback as the scope of change impact widens with a larger number of targets, the number of targets (N) e It can be characterized by using ) as an input value.
[0106] For example, if the orchestrator lookup result has 250 endpoint IDs, the number of targets to apply (N e ) can be derived as 250, and through this, large-scale target policies can be identified in advance and phase application selected.
[0107] Number of policy changes (N) c ) can represent a value indicating the total number of policy change events that occurred for the policy during the recent observation period, and can be used as a unit of count.
[0108] The operation management department (170) or the application-centric policy department (110) can record a change event in the log whenever the policy is changed.
[0109] The aforementioned change events may refer to types such as policy creation, policy modification, policy application, and policy rollback.
[0110] The aforementioned recent observation interval can be used as a fixed value, such as the last 24 hours, or an administrator can select the observation interval in accordance with operating regulations.
[0111] The application-centric policy unit (110) can query logs from the past 24 hours to the current time based on the current time, and among the queried events, only the events corresponding to the policy ID are selected and counted to determine the number of policy changes (N c It can be derived as ).
[0112] For example, if it is a new policy or there are no changes within the period, the number of policy changes (N c ) can be derived as 0, and since policies that change frequently are likely not stabilized and application failures, reapplications, and rollbacks may be repeated, the algorithm for determining the application method is the number of policy changes (N c ) can be used as an input value.
[0113] For example, if two policy modification events and one policy application event are recorded for the relevant policy ID in the last 24 hours, these can be added together to set NcN_cNc to 3, thereby strengthening the application process for frequently changing policies.
[0114] Operational Importance (U way ) may represent a numerical value expressing the business importance of the service to which the application policy applies, and may be used as a grade value without authorization.
[0115] The management server (100) can provide an importance selection input field to the administrator or operator of the management server (100) through the UI.
[0116] The aforementioned importance selection input field may provide options such as General 0, Important 1, and Core 2, for example.
[0117] When an administrator or operator of the management server (100) selects the corresponding importance level (e.g., General is 0, Important is 1, Core is 2) when approving policy registration or distribution, the application-centric policy unit (110) sets the selected value as the operational importance level (U way It can be derived as ).
[0118] Operational Importance (U way The default value of ) can be set to 0, which corresponds to general, and the administrator or operator of the management server (100) may change the importance according to criteria such as internal operation standards.
[0119] Even with policies of the same structure, core services may require a more conservative application, so operational importance (U way ) can be used as an additive factor in the algorithm for determining the application method.
[0120] For example, payment services select the core and operational importance (U way You can set ) to 2, and for the development test service, select General to set the operational importance (U way You can set ) to 0, which allows you to consistently deploy critical service policies more conservatively.
[0121] Difficulty adjustment weight (w way ) is the number of policy rules (N r ) and number of special conditions (N x It can be used as a ratio value to adjust which of the two is more heavily reflected in the difficulty, and can be utilized as a dimensionless value.
[0122] Difficulty adjustment weight (w way The derivation method of ) can be explained in two ways.
[0123] First, as a simple numerical adjustment method, the system default value can be fixed at 0.5, which is the balance value, and the application-centered policy unit (110) can adjust this value by raising or lowering it by a certain amount according to the operation policy entered by the administrator.
[0124] Second, as an administrator input method, a slider adjustable from 0 to 1 is provided on the UI through the management server (100), or selectable options such as 0.3 for exception emphasis, 0.5 for balance, and 0.7 for rule emphasis are provided, and the value selected by the administrator is the difficulty adjustment weight (w way It can be set by saving as ).
[0125] The algorithm for determining the application method initially uses difficulty adjustment weights (w way By setting ) to 0.5, rules and exceptions can be reflected with equal weight.
[0126] Since exceptions may be more dangerous or an increase in the number of rules may be more burdensome in different data center operating environments, instead of a single fixed standard, adjustable difficulty adjustment weights (w way By establishing ), decisions on the application method can be consistently aligned with the operational philosophy.
[0127] For example, if it is determined that exceptions frequently cause accidents, select the exception emphasis option to adjust the difficulty weight (w way You can set ) to 0.3, and as a result, make the influence of exception-related elements relatively larger.
[0128] For example, the manager can adjust the difficulty weight (w way ) can be set to the balance value of 0.5, and if the rule entry in the rule list field of the policy data is 5, the number of policy rules (N r ) can be derived as 5, and if the exception entry of the field containing the exception entry is 1, the number of special conditions (N x ) can be derived as 1.
[0129] If the endpoint IDs in the virtualization linkage unit (150) lookup result are 20, the number of applicable targets (N e ) can be derived as 20, and if there are 0 corresponding policy ID events in the last 24 hours, the number of policy changes (N c ) can be derived as 0.
[0130] If the administrator or operator of the management server (100) selects the importance as general, the operational importance (U way ) can be derived as 0.
[0131] The difficulty score (D) output by the application method determination algorithm using the input values derived by the above-described method way ) can be approximately 4.29, and the application-centric policy unit (110) outputs the difficulty score (D way You can select the immediate application procedure by classifying ) as "low".
[0132] As another example, the number of policy rules (N r ) is 20, number of special conditions (N x ) is 5, number of applicable targets (N e ) is 200, number of policy changes (N c ) is derived as 3, and the operational importance (U way ) is derived as 1, and the difficulty adjustment weight (w way If ) is set to 0.5, the difficulty score (D way ) can be derived as approximately 10.11.
[0133] The application-centric policy department (110) outputs a difficulty score (D) according to the example described above. way You can classify ) as "intermediate" to perform a dry run and then take measures to apply it to some targets first.
[0134] As another example, the number of policy rules (N r ) is 80, number of special conditions (N x ) is 30, number of applicable targets (N e) is 2000, number of policy changes (N c ) is derived as 20, and the operational importance (U way ) is derived as 2, and the difficulty adjustment weight (w way If ) is set to 0.3, the difficulty score (D way ) can be derived as approximately 16.37.
[0135] The application-centric policy department (110) outputs a difficulty score (D) according to the example described above. way You can classify ) as "High" and select the approval process and phased expansion application.
[0136] The algorithm for determining the application method can be implemented in software through the following procedure.
[0137] First, the application-centered policy unit (110) can read a policy object corresponding to a policy ID from the policy repository, and count the number of rule lists and exception lists included in each field of the policy object to determine the number of policy rules (N r ) and number of special conditions (N x ) can be derived.
[0138] Next, the virtualization linkage unit (150) receives the list of endpoint IDs provided by the virtualization linkage unit as input, and the number of applicable targets (N) e It is possible to derive ), and by querying logs from the past 24 hours to the current time based on the current time, the number of policy changes (N c ) can be derived.
[0139] Subsequently, the operational importance (U), a value entered by the administrator and stored in the management server UI. way ) and difficulty adjustment weights (w way After reading ) and using each value, the difficulty score (D way Print ) and the printed difficulty score (D way After selecting an application procedure according to the section of ), a command according to that procedure can be transmitted to the central control node (130).
[0140] The difficulty judgment for network policies of the application-centric policy department (110) is a single score, the difficulty score (D), which is output by an application method determination algorithm, without complexly combining multiple conditions with AND or OR. way It can be implemented by dividing the interval based on ).
[0141] For example, the operator can set two threshold values, and the difficulty score (D way If ) is smaller than the first threshold value, it can be judged as "low," and the difficulty score (D way If ) is greater than or equal to the first threshold value and less than the second threshold value, it can be judged as "intermediate".
[0142] And the difficulty score (D way If ) is greater than or equal to the second threshold value, it can be judged as "high".
[0143] The administrator or operator of the management server (100) may initially set the first standard value to 6 and the second standard value to 10 based on experience, and then adjust the standard value conservatively by referring to the score distribution of policies that have problems through logs.
[0144] The algorithm for determining the application method is the number of recent policy changes (N c By updating ), changes over time can be reflected.
[0145] For example, the application method determination algorithm queries policy change events from the past 24 hours based on the current time every 10 minutes, and counts the number of events for the corresponding policy ID to determine the number of policy changes (N c ) can be updated.
[0146] If changes over the last 24 hours increase, the number of policy changes (N c ) can increase, and as a result, the difficulty score (D way ) can also grow larger.
[0147] Conversely, if changes decrease over time, the number of policy changes (N) in the same aggregation method c ) can become smaller, and as a result, the difficulty score (D way ) can also be lowered.
[0148] Therefore, even for the same policy, the difficulty assessment and application procedures can be designed to vary according to recent trends.
[0149] The value entered by the manager is operational importance (U way ) and difficulty adjustment weights (w way It can be reflected as ).
[0150] Managers select importance as a key factor for operational importance (U way If you increase the ) value, the difficulty score (D) is the same even if other input values are identical. way ) can become larger, and as a result, the algorithm determining the application method may be made to select a more conservative procedure.
[0151] In addition, the administrator can adjust the difficulty weight (w way If you adjust ), the number of policy rules (N r The influence of ) and the number of special conditions (N x The effects of ) can differ from each other.
[0152] Number of special conditions (N) x In environments where ) is viewed as more dangerous, the difficulty adjustment weight (w way Lower the ) value to the number of special conditions (N x It can be made so that the influence of ) becomes relatively larger, and the number of policy rules (N r In environments where ) is more burdensome, the difficulty adjustment weight (w way Increase the value of ) to increase the number of policy rules (N r It can make the influence of ) relatively larger.
[0153] Reflecting such input can be implemented through specific means of selection and saving in the management server UI.
[0154] The above technical configurations and examples are sufficiently provided so that a person skilled in the art can easily implement them according to the algorithm for determining the application method and the definition of variables of the present invention. Although the present invention is not limited to specific examples and various modifications are possible within its scope, it can be confirmed that a person skilled in the art can obviously understand and realize them.
[0155] The central control node (130) can interpret a network policy defined in the application-centered policy unit and transmit a command to the physical network fabric and network equipment connected to the physical network fabric so that the settings are converted to the interpreted network policy.
[0156] Meanwhile, the central control node (130) can utilize a time reduction judgment algorithm that outputs a time score () that determines whether time reduction is necessary by reflecting the time taken for policy interpretation and the causal indicator thereof.
[0157] The time reduction judgment algorithm receives measurement values collected by the central control node (130) and the operation management unit (170) as input during the process in which a policy defined by the application-centric policy unit (110) is interpreted by the central control node unit (130) and converted into equipment settings, and a time score (S sh Can output ).
[0158] The central control node (130) is the time score (S) output from the time reduction judgment algorithm. sh ) can be used to select control recovery scenarios such as strengthening cache usage, switching to incremental analysis, and expanding workers in the operation management department (170).
[0159] The time reduction judgment algorithm may be characterized by quantifying the degradation of policy interpretation performance of the central control node unit (130) and using it as a judgment input for the operation management unit (170) to select a predefined reduction and recovery action.
[0160] In addition, since the time reduction judgment algorithm is configured using a logarithmic function, it can provide a technical effect of reducing overreaction when setting thresholds by smoothing out the increase in scores at large values.
[0161] More specifically, the time reduction judgment algorithm may be characterized by being composed of the sum of a first item configured to increase the score as the interpretation time is longer, a second item configured to increase the score as reinterpretation is more frequent or there are more rules, and a third item configured to increase the correction value as the node load is higher.
[0162] The first item is normalization time (T sh_normal After adding 1 to ) and inputting it into the logarithmic function, the first weight (w) to the result of the logarithmic function sh1 It can be characterized by being configured so that the score increases as the interpretation time increases by multiplying by ).
[0163] The second item is the reinterpretation ratio (R sh ) and policy complexity (P sh After calculating the reinterpretation burden value by multiplying by ), add 1 to the reinterpretation burden value and input it into the logarithmic function, and the second weight (w) to the result of the logarithmic function sh2 It can be characterized by being configured so that the score increases as the frequency of reinterpretation or the number of rules increases by multiplying by ).
[0164] The third item is CPU usage (C sh It can be characterized by being configured such that ) is multiplied by π and then divided by 2 to be input into the trigonometric function sin, and the correction value increases in the range of 0 to 1 as the node load increases.
[0165] The time reduction judgment algorithm may include a first item that reflects the normalized interpretation time, because it may be difficult to distinguish the cause based solely on a single interpretation time when policy interpretation is delayed.
[0166] In addition, since one of the main causes of increased interpretation time in the time reduction judgment algorithm may include frequent reinterpretation, the reinterpretation rate (R sh It can include a second item reflecting ), but since it may take longer if there are many rules to process when reinterpretation occurs, policy complexity (P sh It can be characterized by being reflected in combination with ).
[0167] In addition, since the processing time for the time reduction decision algorithm may increase even with the same policy if the central control node CPU is high, CPU utilization (C sh It can be characterized by reflecting ) as the third item.
[0168] At this time, CPU usage (C sh ) is normalized and then input into the trigonometric function sin, and since it can produce a stable value in the range of 0 to 1 when the input is 0 to 1, the load effect can be smoothly reflected without an overly complex model.
[0169] Finally, the time reduction judgment algorithm is the normalized interpretation time and the reinterpretation rate (R sh ) and policy complexity (P sh By inputting it into the logarithmic function of ), the score is made smooth so that it does not grow infinitely fast even as the input value increases, allowing the operator to make a stable judgment when setting the threshold.
[0170] The time reduction judgment algorithm is normalized time (T sh_normal ) and reinterpretation ratio (R sh ), policy complexity (P sh In the range where ) is large, the score may grow too rapidly, making it difficult to set a threshold. It can be characterized by being designed using a logarithmic function to reflect the increase gradually and perform stable judgment.
[0171] In addition, the time reduction judgment algorithm has a CPU utilization (C) having a range of 0 to 1.sh ) This time score (S sh It can be characterized by being designed using the trigonometric function sin to prevent problems such as excessive reflection or negative values in ).
[0172] As described below in [Table 2] and [Table 3] and the above description, the time reduction judgment algorithm may be characterized by being designed using functions such as logarithmic functions through examples of various tests, and may be characterized by being designed to reflect actual data distribution and experimental experience.
[0173] [Table 2]
[0174]
[0175] [Table 3]
[0176]
[0177] Policy interpretation time (T), each variable used in the time reduction judgment algorithm sh ), standard interpretation time (T ref ), normalized time (T sh_normal ), Reinterpretation ratio (R sh ), policy complexity (P sh ), CPU usage (C sh ), first weight (w sh1 ) and the second weight (w sh2 The characteristics of ) are as follows.
[0178] Policy interpretation time (T sh ) may mean the time taken for the central control node (130) to receive a specific network policy or set of policies as input and convert it into configuration information (e.g., forwarding information or flow rules) applicable to physical network equipment.
[0179] Policy interpretation time (T sh The unit of ) can be standardized and used as milliseconds (ms).
[0180] The central control node (130) can record the start time using the server's timestamp at the time it enters the processing function that interprets the network policy, and can record the end time using the server's timestamp at the time when the equipment-specific setting command object or flow rule list is completed as an interpretation result.
[0181] Subsequently, the policy interpretation time (T) from the value obtained by subtracting the start time from the end time sh ) can be derived.
[0182] Policy interpretation time (T sh ) may mean a value that is not a fixed value but is derived and recorded whenever a request for interpretation of a network policy occurs.
[0183] Since the time reduction judgment algorithm requires a value that directly reflects the actual time required to determine whether the network policy interpretation time can be reduced, the policy interpretation time (T sh It can be characterized by being configured using ) as an input value.
[0184] For example, if the start time is 100000.100ms and the end time is 100000.250ms, the policy resolution time (T sh ) can be derived as 150ms.
[0185] Since this method allows for the quantification of delay based on measurement rather than mere estimation, it enables a clearer understanding of policy interpretation delays compared to conventional, vague empirical judgments.
[0186] Reference interpretation time (T ref ) may refer to a value that serves as the standard for the interpretation time of network policies deemed acceptable in a normal state, and the unit may be milliseconds (ms).
[0187] Reference interpretation time (T ref ) can be set in at least one of the following ways.
[0188] First, in the method entered by the administrator, the operator directly inputs values such as 100ms or 200ms into the management server's configuration screen or configuration file based on Service Level Objectives (SLAs) or internal operational standards, thereby determining the reference resolution time (T ref It can be used as ).
[0189] Second, in the method derived based on record data, the operation management department records the policy interpretation time (T) in the past that the central control node (130) recorded. sh Read the log of ), and, for example, derive the median or mean of the last week, and then the result is used as the reference interpretation time (T ref It can be saved and used as ), and the aggregation period, such as "last week," can be entered by the administrator on the settings screen.
[0190] Third, in the simple numerical adjustment method, the operator uses the pre-set standard interpretation time (T ref You can select adjustment rules, such as increasing by 10ms or multiplying by 1.1 times for ), and the system applies the selected rule as is to the reference interpretation time (T ref ) can be adjusted.
[0191] For example, if the operator selects 'High Sensitivity' on the settings screen, the system references the interpretation time (T ref A predetermined rule can be applied to reduce ) to 0.9 times.
[0192] The time reduction judgment algorithm is the policy interpretation time (T sh To convert ) into a ratio with respect to the reference value and compare using the same standard even if the environment differs, the reference interpretation time (T ref It can be characterized by utilizing ).
[0193] For example, if the operator determines that the SLA is 200ms, the standard interpretation time (T refYou can input ) as 200ms, and by setting this standard value, you can make a judgment that compares "how slow it is compared to the standard" even with different equipment performance or different data center environments.
[0194] Normalized time (T sh_normal ) is the measured policy interpretation time (T sh ) Reference interpretation time (T ref It can mean a ratio value divided by ), and as it is used as a normalized ratio value, it can be used as a dimensionless value.
[0195] Normalized time (T sh_normal Since ) is a normalized ratio value, it is not a fixed value, but the policy interpretation time (T sh ) or reference interpretation time reference interpretation time (T ref It can be derived again whenever ) changes.
[0196] The time reduction judgment algorithm omits the unit of the input element to the logarithmic function to facilitate easy comparison even in different environments, and normalized time (T sh_normal It can be characterized by the use of ).
[0197] For example, policy interpretation time (T sh ) is 150ms and reference interpretation time reference interpretation time(T ref If ) is 100ms, normalization time (T sh_normal ) can be derived as 1.5.
[0198] Reinterpretation Ratio (R sh ) may represent the ratio of policy interpretation requests during a certain time interval in which the cache is not reused and interpretation processing (e.g., parsing, validation, transformation) is actually performed, and as it is used as a ratio, the unit may be a dimensionless unit between 0 and 1.
[0199] You can derive the total number of policy interpretation requests by counting whenever a network policy interpretation request comes in during a specific interval (e.g., 1 minute), and count the number of reinterpretations, which increases by 1 whenever a reinterpretation is actually executed due to a cache miss or policy change detection.
[0200] Subsequent reinterpretation rate (R sh ) can be derived by the value obtained by dividing the number of reinterpretations by the total number of policy interpretation requests.
[0201] However, if the total number of policy interpretation requests is 0, the reinterpretation rate (R sh ) can be derived as 0.
[0202] For the time reduction decision algorithm, since the interpretation processing time is likely to increase if there are many reinterpretations when the same policy is applied repeatedly, the reinterpretation ratio (R sh The delay cause signal can be reflected by utilizing ).
[0203] For example, if the total number of policy interpretation requests in one minute is 50 and the number of reinterpretations is 10, the reinterpretation rate (R sh ) can be derived as 0.2.
[0204] Policy complexity (P sh ) can mean the sum of the number of rule entries defined in the application-centric policy section, and the unit can be used as the number.
[0205] Policy complexity (P sh ) can be derived by counting the number of records in a list of rules stored in a policy repository (e.g., a database table or a JSON / YAML file) that stores network policies.
[0206] For example, if there are 40 allow rule entries, 10 deny rule entries, and 5 priority rule entries stored in the policy repository, the total sum of the number of rules is 55, so the policy complexity (P sh ) can be derived as 55.
[0207] Policy complexity (P sh ) can be updated by recounting the list of rule entries in the policy store whenever a policy is created or modified.
[0208] The time reduction judgment algorithm has a policy complexity (P) because as the number of rule entries increases, the number of targets that the central control node (130) must interpret increases, and thus the possibility of delay may increase. sh It can be characterized by utilizing ).
[0209] CPU usage (C sh ) can mean a value obtained by converting the CPU usage of the central control node into a value between 0 and 1, and can be used as a dimensionless value.
[0210] The operation management department (170) can collect the CPU usage (percentage) provided by the operating system or container runtime, and convert it into a value in the range of 0 to 1 by dividing it by 100, thereby the CPU usage (C sh ) can be derived.
[0211] For example, if the CPU usage is measured at 80%, dividing this by 100 results in 0.8, so the CPU usage (C sh ) can be derived as 0.8.
[0212] Since the time reduction decision algorithm can slow down the policy interpretation processing itself as the central control node's resources become busier, CPU utilization (C sh It can be characterized by being configured to be used as a correction signal by reflecting ).
[0213] First weight (w sh1 ) and the second weight (w sh2 ) can refer to a constant that controls the degree of influence each item has on the final score in the time reduction judgment algorithm, and the unit can be used as a dimensionless unit.
[0214] First weight (w sh1) and the second weight (w sh2 ) can be derived or set in at least one of the following ways.
[0215] First, in the method input by the administrator, if the server administrator or operator wants to detect interpretation time delays more sensitively, the first weight (w sh1 If you want to make ) larger and detect reinterpretation or policy complexity issues more sensitively, use a second weight (w sh2 ) can be entered into the management UI or configuration file.
[0216] Second, in the method derived based on record data, the operations management department can prepare past operation logs, which may include input values such as normalized time values, reinterpretation rates, policy complexity, and CPU utilization, as well as label information regarding whether the corresponding section is normal or if a delay occurred.
[0217] The operator is the first weight (w sh1 ) and the second weight (w sh2 For example, changing ) in increments of 0.5, the output time score (S sh You can select an appropriate value by comparing how well ) distinguishes the actual delay alarm occurrence interval.
[0218] Third, in the simple numerical adjustment method, the operator can select a button such as 'increase or decrease sensitivity,' and the system, according to a predetermined rule, the first weight (w sh1 ) or second weight (w sh2 It can be adjusted by increasing or decreasing ) by a certain ratio.
[0219] For example, when sensitivity is increased, the existing first weight (w sh1 The first weight (w) is the value obtained by multiplying ) by 1.1. sh1 It can be configured to update ).
[0220] The reason for using weights is that bottleneck causes can vary from data center to data center, so even with the same algorithm, it is necessary to adjust which factors are considered more important.
[0221] Such weight adjustment can help adjust alarm sensitivity to suit the operating environment, compared to using only a single threshold.
[0222] For example, the operator standard interpretation time (T ref Set ) to 100ms, and the first weight (w sh1 Set ) to 1.0, and the second weight (w sh2 ) can be set to 0.5, and in this case, if the policy resolution time of a specific request is measured as 150ms via the server's timestamp, the policy resolution time (T sh ) can be derived as 1.5.
[0223] If the Operations Management Department (170) counts the total number of interpretation requests as 50 and the number of re-interpretations as 10 during one minute, the re-interpretation rate (R sh ) can be derived as 0.2.
[0224] Policy complexity (P) by counting rule entries in the policy repository sh ) is derived as 60, and CPU usage (C sh The time reduction judgment algorithm that yields ) 0.8 utilizes this as an input value to determine the time score (S sh Can output ) and the output time score (S sh By comparing ) with a threshold or the average of the previous interval, it is possible to determine whether there is a need to shorten the network policy interpretation time.
[0225] The central control node (130) measures the start and end times of the network policy interpretation processing using a high-resolution timer and the policy interpretation time (T sh ) can be derived.
[0226] In addition, the central control node (130) counts the total number of policy interpretation requests and the number of reinterpretations at each regular period of the operation management unit (170) and the reinterpretation ratio (R sh ) can be derived.
[0227] In addition, the central control node (130) counts the number of rule entries in the policy repository of the application-centered policy unit (110) to determine the policy complexity (P sh ) can be derived.
[0228] In addition, the central control node (130) can read the CPU usage of the server and convert it to a range of 0 to 1.
[0229] The management server can read the reference analysis time and weight values from the configuration screen or configuration file. The operations management department can use these input values to execute a time reduction judgment algorithm to calculate a score, and if the score satisfies the conditions, it can select and execute predefined reduction actions (e.g., enhanced cache reuse, switching to incremental analysis, worker expansion).
[0230] The central control node unit (130) has a policy interpretation time (T sh Receive an absolute reference value for ) from the operator, and policy interpretation time (T sh If the input absolute threshold value is exceeded, it can be determined that the interpretation time of the network policy needs to be shortened.
[0231] In addition, the central control node (130) is the time score (S) output by the time reduction judgment algorithm. sh Receives a threshold value for ), and time score (S sh If ) exceeds the threshold, it can be determined that shortening is necessary.
[0232] In the case where the central control node (130) uses multiple conditions, if the operator selects OR combination, the policy interpretation time (T sh If ) exceeds the absolute threshold value or time score (S shIt can be configured to perform a shortened action if any of the cases where ) exceeds the threshold is satisfied.
[0233] Conversely, if the operator selects the AND combination, the policy resolution time (T sh ) exceeds the absolute threshold value and simultaneously the time score (S sh ) can also be configured to perform a shortened operation only when it exceeds a threshold.
[0234] At this time, policy interpretation time (T sh The absolute reference value of ) and the time score (S sh The threshold value of ) can be configured to be set to at least one of the administrator input method, the historical data-based method, and the simple numerical adjustment method.
[0235] The central control node (130) has a policy interpretation time (T) in 1-minute intervals, for example. sh ), Reinterpretation ratio (R sh ), policy complexity (P sh ), CPU usage (C sh ) can be calculated, and the time reduction judgment algorithm can be repeatedly utilized using the input value of the corresponding section.
[0236] The central control node (130) has a time score (S) of the previous section and the current section. sh The growth rate can be calculated by comparing ), and if the growth rate exceeds a standard set by the operator (e.g., 20%), it can be determined that the interpretation time of the network policy needs to be shortened.
[0237] Additionally, during time periods when application deployment is concentrated, the number of rule entries or reinterpretations may increase, in which case the time score (S sh Changes over time can be reflected in a way that ) also increases together.
[0238] The server administrator or operator has a policy interpretation time (T sh ), standard interpretation time (T ref ), first weight (w sh1 ) and the second weight (wsh2 You can adjust the sensitivity of the judgment on whether to shorten by directly inputting the absolute reference value and the score threshold.
[0239] The operator [is] the first weight (w sh1 If you set ) to a large value, changes in the normalized time value are reflected more significantly in the score, allowing you to configure it to operate sensitively to the interpretation time itself.
[0240] The operator [is] the second weight (w sh2 If ) is set to a large value, the combined change in reinterpretation rate and policy complexity is reflected more significantly in the score, allowing it to be configured to operate sensitively to the causes of reinterpretation or policy complexity.
[0241] In addition, the time reduction judgment algorithm may be characterized by being configured to enable standard correction tailored to the operating environment, as the ratio relative to the standard may change even for the same policy interpretation time if the operator changes the standard interpretation time.
[0242] The above technical configurations and examples are sufficiently provided so that a person skilled in the art can easily implement the time reduction judgment algorithm and variable definitions of the present invention. Although the present invention is not limited to specific examples and various modifications are possible within its scope, it can be confirmed that a person skilled in the art can obviously understand and realize them.
[0243] The virtualization linkage unit (150) is linked with a public cloud environment including virtual machine, container, and virtual network functions, and can automatically apply network policies defined in the application-centric policy unit when the virtual machine or container is created, deleted, or moved.
[0244] The operation management unit (170) can detect and collect the operation status of multiple network equipment constituting a physical network fabric and the central control node unit and the network equipment in real time, detect whether a failure has occurred in the network equipment and provide a notification or perform an automatic recovery operation, and manage expansion by automatically applying a policy when adding or changing network resources.
[0245] Additionally, the operation management unit (170) may be configured to monitor network traffic, connection status, or equipment status in real time to detect abnormal conditions, and when an abnormal condition is detected, provide a notification to the administrator or automatically perform at least one of setting a bypass path, switching backup equipment, or reapplying the policy according to a predefined recovery policy.
[0246] Additionally, the operation management unit (170) may be characterized by selecting a recovery scenario based on a detected abnormal state among predefined recovery policies and sending a setting change command to the corresponding network equipment.
[0247] Additionally, the operation management unit (170) may be characterized by being configured to automatically apply existing defined network policies to the added resources and perform load balancing or path reconfiguration when network equipment, servers, virtual machines, or containers are added.
[0248] Here, path reconstruction may be characterized by collecting traffic and link status of the network equipment, calculating a new transmission path using a path calculation algorithm based on the collected information, and reconstructing the data transmission path by updating the forwarding information or flow rules of the network equipment through a central control node according to the calculated path.
[0249] Meanwhile, the operations management department (170) determines whether to perform a path check based on the accumulation of the number of times a failure occurs, using a scoring system for the path check necessity score (Ssum A path quality summing algorithm that outputs ) can be utilized.
[0250] That is, the path quality summing algorithm is not a simple statistical calculation, but can be used as a judgment criterion for the operation management department (170) to decide whether to perform path checks (setting a bypass path, updating the flow, reapplying the policy, etc.).
[0251] The operation management department (170) can more sensitively reflect failures that occur in a short period of time even with the same cumulative number of times through the first item of the path quality summing algorithm, and can correct the inspection priority through the second item.
[0252] As a result, the operations management department (170) can reduce unnecessary route inspections by utilizing a route quality summing algorithm, while also accelerating inspection decisions in situations accompanied by quality deterioration.
[0253] The operation management department (170) can transmit the output result of the path quality summation algorithm to the central control node department (130), and the central control node department (130) can transmit the received path inspection necessity score (S sum It is possible to control the updating of forwarding information or flow rules of multiple network devices constituting a physical network fabric according to ).
[0254] The path inspection necessity score (S) output by the path quality summing algorithm sum ) can be output as a value that sums up how often failure events occurred in the recent time interval, how congested the current path is, and how poor the current path quality is.
[0255] More specifically, the path quality summing algorithm may be characterized by being composed of the sum of a first item reflecting the failure event occurrence rate per unit time and a second item reflecting both congestion and quality degradation.
[0256] The path quality summing algorithm assigns a fault weight (w) to each of the two items mentioned above. sum1 ) and quality weights (w sum2 The degree of reflection of each element can be adjusted by multiplying by ), and the operation management department (170) can adjust the path inspection necessity score (S) output by the path quality summing algorithm. sum ) is the threshold (θ sum If ) or more, the path check procedure can be initiated.
[0257] More specifically, the first item of the path quality summing algorithm is the number of failure events (F sum The product of ) and the reference time (t0) is accumulated in the time interval (Δt sum Add 1 to the result of dividing by ) and input it into the logarithmic function, and add the fault weight (w) to the result of the logarithmic function sum1 It can be characterized by being configured to reflect the failure event occurrence rate per unit time by multiplying by ).
[0258] In addition, the second item of the path quality summing algorithm is path utilization (U sum Input the result of normalizing ) by multiplying by π and dividing by 2 into the trigonometric function sin, and input the error rate (P sum Add ) and the error rate (P sum Quality weight (w) added to the result of adding ) sum2 It can be characterized by being configured to simultaneously reflect congestion and quality deterioration by multiplying by ).
[0259] The operations management department (170) determines the failure event using a predefined classification rule and the recent cumulative time interval (Δt sum The number of cases during ) is the number of failure events (F sum It can be tallied as ).
[0260] Next, the operations management department (170) has the same number of failure events (F sum Even if it is ), the meaning changes if the time interval is different, so the number of failure events (F sumThe product of ) and the reference time (t0) is accumulated in the time interval (Δt sum It can be converted into a unit time occurrence rate and compared by a path quality summing algorithm configured to divide by ).
[0261] The path quality summing algorithm uses the number of failure events (F) to reduce the distortion of judgment caused by the score becoming excessively large in situations where the number of failure events surges. sum The product of ) and the reference time (t0) is accumulated in the time interval (Δt sum It is configured to input the value divided by ) into a logarithmic function, which can produce a smooth result.
[0262] In addition, since deciding on inspections based solely on failure frequency may lead to missing situations of congestion and quality deterioration, path utilization (U sum ) and error rate (P sum ) can be derived and characterized as being reflected in the second item of the path quality summing algorithm.
[0263] The Operations Management Department (170) can see significant path quality issues, especially in sections with high usage rates, so it uses a path quality summing algorithm to calculate the path usage rate (U sum It can be characterized by normalizing ) and inputting it into the trigonometric function sin to reflect the high usage rate interval more sensitively.
[0264] Finally, the operations management department (170) uses a failure weight (w) to align the importance of the first and second items with the operation policy. sum1 ) and quality weights (w sum2 It can be characterized by incorporating ) into the path quality summing algorithm.
[0265] Here, the path quality summing algorithm may be characterized by the use of a logarithmic function so that the score increases as the failure rate increases, but to prevent excessive runaway and overshadowing other factors.
[0266] In addition, the path quality summing algorithm uses path utilization (Usum It may be characterized by the use of the trigonometric function sin to ensure that the score is reflected more sensitively in the high usage rate interval while maintaining that the value is converted to 0 to 1 within the range of 0 to 1.
[0267] As described in [Table 4] and [Table 5] below and the above description, the path quality summing algorithm may be characterized by being designed using functions such as logarithmic functions through examples of various tests, and may be characterized by being designed to reflect actual data distribution and experimental experience.
[0268] [Table 4]
[0269]
[0270] [Table 5]
[0271]
[0272] The number of failure events (F), each variable used in the path quality summing algorithm sum ), cumulative time interval (Δt sum ), reference time (t0), path utilization (U sum ), error rate (P sum ), fault weight (w sum1 ) and quality weights (w sum2 ), threshold(θ sum The characteristics of ) are as follows.
[0273] Number of failure events (F sum ) is the cumulative time interval (Δt sum It may mean the number of occurrences of events determined by the operation management department (170) as a failure or abnormality during ) and the unit may be used as times (number of cases).
[0274] The operation management department (170) can receive or query at least one of a link down or link up event, a BFD session down event, a port error counter surge alarm, a drop counter surge alarm, and an equipment status abnormality alarm from the management interface of a plurality of network equipment or routers (500) that constitute a network equipment or a physical network fabric, and count only the event that matches a predefined failure or abnormality classification rule as one event among them, thereby counting the number of failure events (F sum ) can be derived.
[0275] For example, a link down event or a BFD down event can each be counted as one, and an error or drop alarm can be set to be counted as one only when the error rate or drop rate exceeds a predetermined threshold.
[0276] The operation management department (170) has a cumulative time interval (Δt sum Number of failure events (F) when ) starts sum Set ) to 0, and whenever the corresponding event is additionally detected within the time window, the number of failure events (F sum It can be managed by increasing ) by 1, and the cumulative time interval (Δt sum Events older than ) are excluded from the aggregation, so the recent cumulative time interval (Δt) sum Only the number corresponding to ) can be maintained.
[0277] Number of failure events (F sum Since ) is a value that directly reflects the cumulative number of times a failure is detected, which is a requirement of the invention, it can be used as a direct basis for determining whether to check a path. For example, if 2 link downs and 1 BFD down occur within the last 5 minutes, the Operations Management Department [determines] the number of failure events (F sum ) can be set to 3.
[0278] In this way, the number of failure events (F) sumUnlike simply viewing the current failure status as 0 or 1, using ) can reflect how frequently the failure has occurred recently, which can help distinguish between temporary events and recurring failures.
[0279] Additionally, the operation management unit (170) may use AI to assist in event classification. In this case, the operation management unit may provide information such as event code, time of occurrence, equipment ID, port ID, counter snapshot, and alarm message as input values to the AI. The AI may provide classification labels such as link failure or quality deterioration as output values, along with whether the event is a failure or anomaly (e.g., 0 or 1). The operation management unit (170) may provide the number of failure events (F) only when the output value is determined to be a failure or anomaly. sum It can be configured to increase ) by 1.
[0280] Cumulative time interval (Δt sum ) is the number of failure events (F sum It can refer to the recent time interval for aggregating ), and the unit can be used as minutes.
[0281] Also, the cumulative time interval (Δt sum ) can be derived from setting values directly entered by an administrator or operator through the UI of the management server (100) by the operation management department (170).
[0282] Cumulative time interval (Δt sum First, it can be derived from a value set by the administrator directly entering it on the operation screen, such as "last 5 minutes".
[0283] Second, after examining past operational logs to determine the average duration of failures or the length of periods where events are concentrated, an accumulated time interval (Δt) is defined to sufficiently include the periods where failures are concentrated. sum It can be derived by determining ).
[0284] Third, initially, the cumulative time interval (Δtsum You can change it step by step through simple numerical adjustments, such as setting it to 5 minutes and increasing it to 10 minutes if the inspection is too frequent, or decreasing it to 3 minutes if the inspection is too late.
[0285] The path quality summing algorithm is the number of identical failure events (F sum Even if it is ), because the meaning differs depending on whether it occurred in a short period of time or accumulated over a long period, the number of failure events (F sum To compare ) as the frequency of occurrence per unit time, the cumulative time interval (Δt sum It can be characterized as being composed using ).
[0286] As a result, the operations management department (170) has a cumulative time interval (Δt sum By utilizing a path quality summing algorithm configured using ), it is possible to identify situations where failures occur in a short period of time more quickly than a method that judges based only on the accumulated count.
[0287] The reference time (t0) is a reference time constant, and the number of failure events (F sum ) and cumulative time interval (Δt sum It can mean a value to clarify the reference unit when expressing the frequency of occurrence per unit time by combining ), and the unit can be used as minutes.
[0288] The operation management department (170) can store the reference time (t0) as a fixed constant, and, for example, can use it fixed at 1 minute.
[0289] The path quality summing algorithm is the number of failure events (F sum ) accumulating time interval (Δt sum It can be characterized by being constructed using a reference time (t0) to clarify the reference unit so as not to cause confusion in the value obtained by dividing by ).
[0290] As a result, the operations management department (170) can reduce the risk of incorrect path inspection thresholds or weight settings due to unit confusion when using the path quality summing algorithm.
[0291] Path utilization (U sum ) is the usage rate of a path or a link included in that path, i.e., congestion, and can be utilized as a dimensionless ratio between 0 and 1.
[0292] The operation management department (170) can read the interface counter of the network equipment, for example, the accumulated number of bytes, at regular intervals, and can calculate the amount of bytes actually increased during the measurement period by subtracting the accumulated number of bytes read immediately before from the accumulated number of bytes read this time, calculate the throughput by dividing this increase amount by the measurement time, and then divide it by the link capacity (e.g., 10 Gbps) to obtain the path utilization rate (U) of the corresponding link. sum ) can be calculated between 0 and 1.
[0293] When a path consists of multiple links, the operation management department (170) selects one of the methods to use the maximum value of the usage rate among the links included in the path as the representative usage rate of the path, or to use the average value of the usage rate as the representative usage rate, thereby determining the path usage rate (U sum ) can be derived.
[0294] The operations management department (170) stores the "previous counter value" at the initial time and calculates the path utilization rate (U) by differential calculation from the next measurement time. sum ) can be derived.
[0295] The path quality summing algorithm is the number of failure events (F sum In a situation where ) accumulates, if the route is actually congested, the priority of route checking may increase, so route utilization (U sum It can be characterized by being configured to reflect the level of congestion together by utilizing ).
[0296] For example, if the interface accumulated bytes increased by 6GB over one minute and the link capacity is at the 1GB / s level, the corresponding one-minute average path utilization (U sum ) can be derived at a level of approximately 0.1.
[0297] Operations Management Department (170) has a path usage rate (U sum By reflecting ), the possibility of quality degradation due to congestion can be considered together with the method of judging based solely on the number of failures.
[0298] Error rate (P sum ) can represent the degree of path quality deterioration, such as the error rate or drop rate, and can be used as a dimensionless ratio between 0 and 1.
[0299] The operation management department (170) can read the error counter (e.g., number of CRC error packets) or drop counter (e.g., number of queue drop packets) of the network equipment at regular intervals and can obtain the number of error packets or drop packets that increased during the measurement period by subtracting the previous accumulated value from the current accumulated value.
[0300] The operation management department (170) calculates the total number of packets increased during the same period, and then divides the increase in errors or drops by the total increase in packets to obtain an error rate (P sum ) can be derived.
[0301] For example, if defined as the drop rate, the operations management department derives the error rate (P) from the value obtained by dividing the number of dropped packets increased during the measurement period by the total number of packets increased during the measurement period. sum ) can be derived.
[0302] Operations Management Department (170) initially has an error rate (P sum ) can be set to 0, and then updated via the counter difference at each measurement cycle.
[0303] The path quality summing algorithm includes the error rate (P) in the inspection judgment because actual service quality issues can occur if drops or errors increase even without clear failures such as link downs. sum It can be characterized by being configured using ), and through this, the operation management department (170) can reflect quality deterioration more quickly than judgment centered on down events.
[0304] Disability weight (w sum1 ) and quality weights (w sum2 ) can refer to a coefficient that controls the influence of factors based on failure frequency and factors based on congestion / quality on the final judgment score, and can be used as a dimensionless value.
[0305] Disability weight (w sum1 ) and quality weights (w sum2 ) can be derived by the three methods described below.
[0306] First, if the administrator wants failure-centric operations on the operation screen, the failure weight (w sum1 Enter a relatively large ) and if you want congestion and quality-oriented operations, the quality weight (w sum2 You can set it by inputting ) relatively large.
[0307] Second, the number of failure events (F) at the point in time when many unnecessary checks occurred in past operation logs. sum ) and path utilization (U sum ), error rate (P sum Check the ) value, and the path inspection necessity score (S) calculated at that point. sum To ensure that ) does not easily exceed the threshold, the fault weight (w sum1 ) and quality weights (w sum2 It can be derived based on recorded data by adjusting ).
[0308] Third, initially, the disability weight (w sum1 ) and quality weights (w sum2You can set each to 1 and make simple numerical adjustments by raising or lowering them in increments of 0.1 or 0.5 while observing the inspection results during operation.
[0309] Since the path quality aggregation algorithm uses failure weights (w) to adjust the reflection weight according to operational policies, as failure patterns, congestion, and quality importance vary by data center. sum1 ) and quality weights (w sum2 It can be characterized by being configured using ), and as a result, it can be adjusted to reduce false positives or missed detections according to field characteristics compared to a fixed standard method.
[0310] threshold (θ sum ) can represent a criterion value for determining whether to perform a path check, and can be used as a dimensionless unit.
[0311] threshold (θ sum ) can be derived by the three methods described below.
[0312] First, administrators can directly input inspection trigger criteria on the operation screen.
[0313] Second, it can be derived based on recorded data by collecting final score values at the time of past failures, identifying the range in which scores are distributed within the sections where actual failures or quality degradation occurred, and then determining the range based on that.
[0314] Third, initial threshold (θ sum After setting ), if the check is excessive, the threshold value (θ sum Raise ) and if the inspection is delayed, the threshold (θ sum It can be adjusted step by step through simple numerical adjustment by lowering ).
[0315] Because the path quality summing algorithm requires a judgment criterion to perform path checks only when necessary rather than always, a threshold value (θ) sumIt can be characterized by using ), and as a result of using it, unnecessary checks can be reduced compared to a method that performs checks even when a simple event occurs only once.
[0316] The Operations Management Department (170) checks the route inspection necessity score (S) at 1-minute intervals. sum Assuming that ) is updated, at each update point, the most recent accumulated time interval (Δt sum Number of failure events (F) during ) sum Re-calculating ), and the path usage rate (U at that point in time) sum ) and error rate (P sum ) can be derived again as a counter difference.
[0317] For example, if the reference time (t0) is fixed at 1 minute, and the cumulative time interval (Δt sum Set ) to 5 minutes, and the fault weight (w sum1 Set ) to 1, quality weight (w sum2 Set ) to 1.1, threshold (θ sum ) can be set to 3.0, and the number of failure events (F sum ) is 30, path utilization (U sum ) is 0.9, error rate (P sum If ) is derived as 0.02, the path check necessity score (S sum ) can be output as approximately 3.055.
[0318] threshold (θ sum When ) is 3.0, the path check necessity score (S) output according to the example above sum Since ) exceeds the threshold, the operations management department (170) may decide to perform a path check.
[0319] The operation management department (170) can transmit the results of these decisions to the central control node department (130) to request that it set a bypass path, update forwarding information, or update flow rules.
[0320] The operation management department (170) can first receive candidate events for failures and anomalies through an event receiving unit or a log collection unit to implement a path quality summing algorithm in software, and can determine whether the event corresponds to a failure or anomaly by applying a predefined classification rule or optionally an AI classifier, and only if it corresponds to a record in an aggregation structure to determine the number of failure events (F sum It can be reflected in ).
[0321] The operations management department (170) has an accumulated time interval (Δt) based on the current time. sum Remove records older than ) from the aggregation structure and use the number of remaining records to count the number of failure events (F sum ) can be derived.
[0322] Operations Management Department (170) reads the equipment interface counter and the drop or error counter to determine the path utilization (U) from the difference with the previous snapshot. sum ) and error rate (P sum ) can be derived.
[0323] The operations management department (170) can then calculate the final score by reflecting the sum of the log value for the failure frequency factor, the sine value for the congestion factor, and the quality deterioration value as weights, respectively, and the path inspection necessity score (S sum ) is the threshold (θ sum It can be configured to determine whether to perform a path check by comparing whether there is an abnormality, and if it is determined that a check is necessary, to send a command to change the setting or update the rule to the central control node (130).
[0324] The operation management department (170) can operate the criteria and conditions of the path quality summing algorithm as a single standard, which can be expressed as a condition that “if the final score is greater than or equal to the threshold, a check is performed.”
[0325] Operations Management Department (170) states that since the final score is a value that combines the failure frequency factor and the congestion and quality factor as weights, the failure weight (wsum1 ) and quality weights (w sum2 You can combine the reflection weight of the two elements by adjusting ), and the threshold value (θ sum You can adjust the overall inspection sensitivity by adjusting ).
[0326] The operations management department has a threshold value (θ sum ) can be set to at least one of an administrator input method, a record data-based method, or a simple numerical adjustment method, and can be readjusted by reflecting the frequency of inspections or failure response results during operation.
[0327] The Operations Management Department [determines] the number of failure events (F) over time sum ) is the recent cumulative time interval (Δt sum Since it is maintained within the range, if failures continue to occur, the number of failure events (F) increases as the number of events remaining in the aggregation structure increases. sum As ) increases, the final score can rise, and if the failure stops, as time passes, the event falls out of the window, and the number of failure events (F) follows. sum The final score may decrease as ) decreases.
[0328] The Operations Management Department calculates the path utilization rate (U) through the counter difference at every measurement cycle. sum ) and error rate (P sum Since it updates ), if congestion increases or drops or errors increase, the final score at that point can be reflected so that it immediately rises.
[0329] Accordingly, the operations management department can repeatedly determine whether the condition is met over a continuous time interval rather than at a single point in time.
[0330] The Operations Management Department provides operational variability based on user input through an accumulated time interval (Δt sum ), fault weight (w sum1 ) and quality weights (w sum2 ), threshold(θ sum ) can be exposed as an operational policy parameter, and if inspections are too frequent, the administrator can set the threshold (θ sumIncrease ) or cumulative time interval (Δt sum You can lower sensitivity by increasing ), and if the inspection is too late, the threshold (θ sum Lowering ) or the cumulative time interval (Δt sum You can increase sensitivity by reducing ).
[0331] In addition, if the manager places more importance on failure events, the failure weight (w sum1 You can input a relatively large ), and if you prioritize congestion and quality more, the quality weight (w sum2 ) can be entered relatively large, and each parameter can be set and adjusted in at least one of the following: an administrator input method, a record data-based method, or a simple numerical adjustment method.
[0332] The above technical configurations and examples are sufficiently provided so that a person skilled in the art can easily implement the path quality summing algorithm and variable definitions of the present invention. Although the present invention is not limited to specific examples and various modifications are possible within its scope, it can be confirmed that a person skilled in the art can obviously understand and realize them.
[0334] The embodiments described above are for illustrative purposes only, and those skilled in the art will understand that the embodiments described above can be easily modified into other specific forms without altering the technical concept or essential features of the embodiments described above. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.
[0336] The scope of protection sought through this specification is defined by the claims set forth below rather than by the detailed description, and should be interpreted to include all modifications or variations derived from the meaning and scope of the claims and the concept of equivalents. Explanation of the symbols
[0337] 100: Management Server
Claims
Claim 1 A software-defined network solution system comprises: a management server that provides application-centric network control in a data center environment; and further comprises a router, which is one of a plurality of components that constitute a physical network fabric in which physical network equipment is interconnected to form a data transmission path of the data center network, and which performs the role of a boundary node responsible for traffic between subnets or interaction with an external network within the data center; wherein the management server controls the data transmission path, traffic flow, and policy application status of the physical network fabric by transmitting control commands to a plurality of network equipment constituting the physical network fabric; and wherein the management server comprises: an application-centric policy unit that defines and applies network policies including communication allowance, blocking, priority, or security rules based on application units on the network of a plurality of network equipment constituting the physical network fabric; a central control node unit that interprets the network policies defined by the application-centric policy unit and transmits commands to the physical network fabric and network equipment connected to the physical network fabric so that settings are converted to the interpreted network policies; and wherein the system is linked with a public cloud environment including virtual machine, container, and virtual network functions, and when the virtual machine or container is created, deleted, or moved, the application-centric policy unit A virtualization linkage unit that enables the automatic application of defined network policies; and a plurality of network devices constituting the physical network fabric and the central control node unit, an operation management unit that detects and collects the operating status of the network devices in real time, detects whether a failure has occurred in the network devices and provides a notification or performs an automatic recovery operation, and manages expansion by automatically applying policies when network resources are added or changed;The system includes the application-centric policy unit, which is configured to automatically select and apply a network policy corresponding to the application when deploying the application or creating a service instance; the operation management unit is configured to monitor network traffic, connection status, or equipment status in real time to detect abnormal conditions, and when an abnormal condition is detected, to provide a notification to an administrator or to automatically perform at least one of setting a bypass path, switching backup equipment, or reapplying policies according to a predefined recovery policy; the operation management unit is characterized by selecting a recovery scenario according to the detected abnormal condition among predefined recovery policies and transmitting a configuration change command to the corresponding network equipment; the operation management unit is characterized by being configured to automatically apply an existing network policy to the added resource and perform load balancing or path reconfiguration when network equipment, servers, virtual machines, or containers are added; and the path reconfiguration is characterized by collecting traffic and link status of the corresponding network equipment, calculating a new delivery path using a path calculation algorithm based on the collected information, and reconfiguring the data delivery path by updating the forwarding information or flow rules of the network equipment through a central control node according to the calculated path. A software-defined network solution system. Claim 2 delete Claim 3 In claim 1, the application-centric policy unit comprises a difficulty score (D) output from an application method determination algorithm. way A software-defined network solution system characterized by reducing the possibility of mis-distribution, policy conflicts, and large-scale impact spread by quantifying the difficulty level at the pre-policy application stage for network policies defined based on ) and selecting an application procedure to apply complex and high-risk policies more conservatively. Claim 4 In claim 3, the above-mentioned application method determination algorithm comprises the number of policy rules (N r The more ) increases, the more the difficulty score (D way The number of policy rules (N) increases so that ) r After adding 1 to ) and inputting it into the logarithmic function, the difficulty adjustment weight (w) to the result of the logarithmic function way Includes the first element configured to multiply by ), and the number of special conditions (N x The more ) increases, the more the difficulty score (D way The number of special conditions (N) increases so that ) x After adding 1 to ) and inputting it into the logarithmic function, the result of the logarithmic function is 1 to the difficulty adjustment weight (w way Includes a second element configured to multiply the result of subtracting ), and the number of applicable targets (N e The more ) increases, the more the difficulty score (D way To increase the number of applicable targets (N) e Includes a third element configured to input into a logarithmic function by adding 1 to ), and the number of policy changes (N c The more ) increases, the more the difficulty score (D way Number of policy changes (N) to increase ) c Includes a fourth element configured to input into a logarithmic function by adding 1 to ), and operational importance (U way ) difficulty score (D way A software-defined network solution system that includes a fifth element configured to handle critical service policies more conservatively, even at the same complexity, in addition to ) directly. Claim 5 In claim 4, the application method determination algorithm comprises the number of policy rules (N r ), number of the above special conditions (N x ), the number of applicable targets mentioned above (N e ) and the number of times the above policy was changed (N c A software-defined network solution system characterized by applying a logarithmic function to ensure that score changes are reflected smoothly even if the range of values for each factor rises or falls sharply. Claim 6 In claim 4, the application method determination algorithm comprises the number of policy rules (N r ) and the number of the above special conditions (N x Although ) is a factor of the policy item, the difficulty adjustment weight (w) is designed to reflect changes in the contribution to operational risk. way A software-defined network solution system characterized by being configured to adjust the contribution using ).