An improved ibn network-based intention conflict coordination resolution method

By transforming intent conflict into a resource allocation problem and using deep reinforcement learning and game theory methods, the conflict caused by resource competition in intent networks is resolved, resource utilization is optimized, user experience and satisfaction are improved, and network state changes are adapted.

CN116170878BActive Publication Date: 2025-12-12CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211631308.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2025-12-12
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

In intent networks, existing technologies struggle to effectively resolve conflicts between multiple intents, especially when resources are limited, leading to a decline in user experience and satisfaction. Furthermore, existing solutions require significant human intervention and cannot adapt to changes in network conditions.

Method used

The intention conflict problem is transformed into a resource allocation problem. Intentions are divided into time-tolerant and non-time-tolerant types based on the time dimension. Deep reinforcement learning and game theory methods are used in combination with resource allocation strategies to coordinate intention conflicts and optimize network resource utilization.

Benefits of technology

It enables automated coordination of conflicting intentions under resource constraints, ensuring user experience and satisfaction, while improving resource utilization and adapting to changes in network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116170878B_ABST
    Figure CN116170878B_ABST
Patent Text Reader

Abstract

The application relates to an intention conflict coordination and resolution method based on an improved IBN network and belongs to the field of intention networks, and comprises the following steps: S1: mapping the intention request of a user to the level of network resources, and converting the intention conflict problem into a resource allocation problem; S2: dividing the resource allocation problem into time-tolerant intentions and non-time-tolerant intentions in the time dimension; S3: for the time-tolerant intentions, the intention is adjusted to other time slices adjacent thereto to be realized, and the influence of the previous time period on the time period is increased; and S4: for the non-time-tolerant intentions, multiple performance intention conflict resolution problems are converted into a multi-objective optimization problem about network resource deployment, a deep reinforcement learning network is used to continuously perceive and learn the network environment, the deployment strategy in the network is adjusted and updated, and the intention conflict is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of intent-based networking, and relates to an intent conflict coordination and resolution method based on an improved IBN network. BACKGROUND

[0002] With the rapid development of network technology, in addition to bringing people a colorful life, it also leads to an increasingly complex network structure. The human and time costs required for network configuration, operation and maintenance will be greatly improved. It can be understood that the automated network which can understand user demand, analyze network status, execute network strategy and closed-loop operation has become a new trend of network development. The development of software-defined networking (SDN) is an important step in the development of network automation, which separates the control plane and data plane of the network, making it easier to deploy the network uniformly and quickly. However, in the face of complex and variable network environment, it still cannot meet the diverse needs of users, and it is impossible to understand the user's intent demand. The proposal of intent-based networking (IBN) makes it possible for users to directly express their intent to the network, and also makes network automation possible. IBN is a brand-new network model, which can analyze the intent expressed by the user and the network performance intent in the network, and IBN translates the intent into the corresponding network strategy through analyzing these intents, and automatically configures the network according to the network strategy, and detects and optimizes it.

[0003] Intent is the core of IBN, and the operation process of IBN is closely related to intent. Users only need to describe their needs without managing the implementation process. The network automatically implements related intents and continuously monitors network state information, and constantly optimizes and adjusts, forming an automated closed-loop management of the network. In the intent network, multiple intents can exist at the same time. Detecting these intents and finding the relevance between them can help the network better understand the intents and the relationship between them. However, these intents are not always harmonious, and these conflicting intents in the network are closely related to intent translation, network existing state and network configuration strategy formulation, etc. Therefore, it is unreasonable to ignore the conflicts between intents, discard conflicting intents, or set priorities according to the time sequence of arrival, which will seriously affect the user experience, and will have a negative impact on user satisfaction and network status.

[0004] The Open Network Foundation NorthBound Interface Working Group (ONF NBIWG) published the Intent NBI white paper in 2015, which proposed an NBI concept for user network operation intent, introduced the use, characteristics and basic architecture of the intent northbound interface. Intent NBI converts the intent requirements expressed by users into network configuration details related to specific technologies and fields, and then completes the configuration according to the details. The ONF chairman Daivd Lenrow proposed the standard draft "Intent: Don't Tell Me What to Do! Tell Me What You Want!", which pointed out that the intent network is only to tell the network its own needs, without understanding how to configure. This clearly defines the concept of intent network, bringing new direction and opportunity to the development of network, and then the research on intent network is carried out by all parties. Most of the current research on intent network is based on software-defined network. In software-defined network, some SDN controllers (ONOS, OpenDaylight, etc.) have opened the northbound user intent interface. IBN can be regarded as a high-level and intelligent SDN, but IBN is not limited to SDN, it is a more flexible, intelligent and automated network model.

[0005] With the in-depth study of intent network, its combination with existing network research is also closer. The intent of network performance indicators and quality of experience (QoE) is constantly mentioned. Yu Hongfang et al. proposed a QoE monitoring system based on intent software-defined network, which can improve user experience and more effectively utilize network resources. Resources are always limited, whether it is bandwidth resources, cache resources or computing resources. However, in the case of limited resources, the user's expectations for the network are constantly rising. How to balance multiple intent requirements in the network under the condition of limited resources is a key problem that needs to be solved in intent network, which will greatly affect the user's experience and satisfaction.

[0006] Yang Hui et al. proposed a self-adaptive slice generation and optimization strategy based on deep reinforcement learning (SPG-RL) to generate a combined strategy that meets the intent, while using a deep neural evolutionary network (DNEN) auxiliary model (SPG-RL-DNEN) to reconfigure incompatible slices to ensure the intent. C. Prakash et al. proposed a policy graph abstraction (PGA) to represent network policies. PGA is a simple and intuitive abstraction graph that detects and resolves policy conflicts using a graph structure, models and combines service chain policies by merging multiple service chain requirements into conflict-free combined chains. Zhang Hao Di et al. proposed an intent parser based on intent-based SDN to analyze the received intent, and when there are multiple conflicting intents in the network, to decide how to reconfigure the network, optimize the running scheme in real time, and detect and resolve conflicts between intents. Li Yuheng et al. proposed a method to solve the consistency problem of intent delivery based on the new northbound interface Intent NBI. A module implementing the commitment theory is added to the Intent NBI to solve the logical consistency problem of the SDN northbound interface. Zhang Jia Ming et al. proposed QICR (Quadruple-based Intent Conflict Resolution) to solve the intent conflict between multiple users in the same network during the initial network construction. By constructing a network intent quadruple <SrcGroup, DstGroup>, <filter> , <sfc> , <constraint>and the way of constructing network information graph and using conflict resolution algorithm to resolve intention conflict.

[0007] In view of the intention conflict problem existing in the intention network, the existing solution is to first map the intention demand to the demand at the network level, so as to solve the intention conflict problem in some network conflict solving schemes in software defined network. The demand is shown in the form of a graph, and the conflict problem is resolved using a graph-based theory. However, it requires a large amount of manual participation, which is contrary to the concept of network automation and network autonomy of the intention network to some extent. Moreover, the network strategy formulated in this way is often not maintained after deployment, and when the network state changes, the user's intention request is difficult to guarantee. SUMMARY

[0008] Therefore, the purpose of the present application is to provide an intention conflict coordination and resolution method based on an improved IBN network. Considering the real-time changes of the network in the Internet of Things edge computing and the maintenance and alternation of user intention demand, the user's intention request is first mapped to the level of network resources, because most user network demands are essentially related to some demands of QoS in the network. The conflict between intentions is often due to the conflict of resource competition. Based on the above analysis, after converting the intention conflict problem into a resource allocation problem, it is first divided from the time dimension into time-tolerant intention and non-time-tolerant intention. For time-tolerant intention, the intention is adjusted to the adjacent time slice to realize, so as to resolve the intention conflict caused by network resource competition from the perspective of time resource. For the conflict that cannot be resolved from time, the multiple performance intention conflict resolution problem is converted into a multi-objective optimization problem about network resource allocation, and the method of deep reinforcement learning is used to solve it. Continuously perceive and learn the network environment, adjust and update the deployment strategy in the network, which can not only guarantee the maintenance of intention under closed loop condition, but also can optimize and adjust it. Combined with the thinking of game theory, the solution of coordination and cooperation is adopted to guarantee the resource utilization rate and formulate the optimal resource allocation strategy to avoid resolving intention conflict.

[0009] To achieve the above purpose, the present application provides the following technical scheme:

[0010] An intention conflict coordination and resolution method based on an improved IBN network, comprising the following steps:

[0011] S1: mapping the user's intention request to the level of network resources, and converting the intention conflict problem into a resource allocation problem;

[0012] S2: Divide the resource allocation problem into time-tolerant intentions and non-time-tolerant intentions from a time dimension;

[0013] S3: For time-tolerant intentions, adjust the intention to be implemented in other adjacent time slices, and increase the influence of the previous time slice on that time slice;

[0014] S4: For non-time-tolerant intents, the problem of resolving multiple performance intent conflicts is transformed into a multi-objective optimization problem of network resource allocation. A deep reinforcement learning network is used to continuously perceive and learn the network environment, adjust and update the deployment strategy in the network, and resolve intent conflicts.

[0015] Furthermore, in step S1, the user's intent request is mapped to the network resource layer, transforming the intent conflict problem into a resource allocation problem. First, an intent IoT system is constructed, including an intent management server, edge nodes, and terminal devices. The edge node consists of a micro base station (SBS) and a MEC server. The intent management server communicates wirelessly with the edge nodes, the edge nodes communicate with each other via wired links, and each terminal device communicates wirelessly with its corresponding micro base station.

[0016] The intent management server parses and translates user intents through the intent engine, then verifies network policies based on the current network status information, optimizes the policies, and then distributes them to the actual network infrastructure.

[0017] The terminal device is used to collect user intents input in various forms and unify these intents into a standard form.

[0018] The network is defined as having M edge nodes, namely M MEC servers and M base stations. This represents the set of MEC servers; the number of end-user devices is U. The set of terminal devices is represented as: the set of user equipment associated with base station m is represented as The decision-making process of the intent management server is set to take place in discrete time slots, each time slot t having a constant duration T. s ;

[0019] Assume each base station can allocate K types of network resources of the same type and quantity, where the types are represented as follows: V network resources available to each base station K The total amount is expressed as Will Divided into capacity of The amount of resources allocated to a user equipment in time slot t for a sub-resource block is represented as follows:

[0020]

[0021] The deployment vector of network resource is expressed as:

[0022]

[0023] At time slot t, the total amount of all resources that can be deployed in the network is lower than the amount that can be deployed in the network, i.e.:

[0024]

[0025] Further, the intention of a user is defined as a performance intention related to network performance, which is classified into constraint intention and optimization intention, and the constraint intention is classified into best-effort intention and guaranteed intention, wherein:

[0026] The constraint intention is an intention that requires certain network performance to be no lower or higher than a certain threshold value;

[0027] The optimization intention is an intention that requires network performance indicators to be maximized or minimized;

[0028] The best-effort intention is an intention that makes network performance indicators best effort to reach the desired threshold value;

[0029] The guaranteed intention is an intention that requires network performance indicators to reach the desired threshold value.

[0030] Further, the performance indicator described by the performance intention depends on the deployment of network resources and other variable sets x related to performance n At time slot t, the performance indicator described by the performance intention is expressed as:

[0031]

[0032] Where F n [*] is the performance function of the performance indicator described by the performance intention; the performance function of the conflicting intention is expressed as Where l represents the number of performance intentions;

[0033] The user intention satisfaction at time t is defined as follows:

[0034]

[0035] Where D(t) is the bandwidth at time t, V(t) is the delay at time t, L(t) is the packet loss at time t; α d ,α v ,α l is the calculation coefficient of each network resource;

[0036] The multi-performance intention conflict avoidance problem is expressed as an optimization problem with the optimization type performance intention and the best-effort type performance intention as optimization objectives, and the guarantee type performance intention as a constraint condition. By taking the inverse of the performance function inequality of the threshold type performance intention, the performance function of the intention greater than or not less than a specific threshold is converted into the performance function of the intention less than or not greater than a specific threshold. The optimization problem is expressed as follows:

[0037]

[0038]

[0039] y i (t)≥1,i∈I nec

[0040] where I nec is the set of guarantee type intentions.

[0041] Further, the conflict resolution for the time-tolerant type intention in the step S3 includes the following steps:

[0042] The processed intention information is unfolded according to the time sequence information, and is mapped into a resource description framework graph (RDF) by considering the operation object, operation action, network performance index and the like involved in the intention.

[0043] The information after the splitting of each intention is obtained.

[0044] The multiple conflicting intentions are generated into an RDF graph, and are merged according to the time sequence relationship.

[0045] According to the resource usage in the network and the demand of each intention on the time sequence, for the intention which does not have strict requirements on the time resource, if the intention competes with other intentions on the network resource, the intention is moved to the adjacent time slice with idle network resource to be implemented.

[0046] The influence of the previous time period on the current time period is increased, that is, the successfully configured intention in the previous time slice and the intention still existing in the current time slice are also guaranteed to be successfully configured. When the effective part of the intention which cannot be successfully configured in the previous time slice is also effective in the current time slice, the intention is no longer configured.

[0047] Further, in the step S4, for the resource competition problem converted from the intention conflict which cannot be resolved in time, the problem is converted into a multi-objective joint optimization problem, and the performance index, expected state value and resource utilization rate of the intention design are considered as optimization objectives. An optimal resource allocation strategy under the constraint condition is formulated.

[0048] The double-delay deep deterministic policy gradient algorithm TD3 is used to customize the resource allocation strategy of the conflict intention related to optimization and network performance. The TD3 model is described as an agent, which interacts with the environment by executing an action in the current state and obtaining a new state. At the same time, the agent obtains the corresponding reward from the environment, and then evaluates the action according to the immediate reward to decide whether to increase or decrease the reward. The optimal strategy is obtained by optimizing the policy to obtain the maximum reward, and it is considered that the optimal strategy is found.

[0049] The reward function is:

[0050]

[0051] Wherein γ∈(0, 1), which is a discount factor for determining the priority of short-term rewards;

[0052] The policy function is:

[0053] π(a|s)=P(A=a|S=s),π(s,a)→[0,1]

[0054] The state value function is recursively represented by the Bellman equation:

[0055]

[0056] The action value function is recursively represented by the Bellman equation:

[0057]

[0058] The TD3 algorithm uses two sets of networks to represent different Q values, and selects the smallest one as the update target;

[0059] The smaller one is taken as the Q-target or Q-value, and the update target is:

[0060]

[0061] Update the evaluation network:

[0062]

[0063] In the update of the policy network, only one evaluation function is used to provide gradient for the policy function:

[0064]

[0065] The key performance numerical level, network resource capacity limit and link state are taken as the state, the parameter setting of the network resource deployment strategy is taken as the action, and the current numerical value of the optimization target and the resource utilization rate are taken as the reward.

[0066] The application has the beneficial effects that in the research on the intention network, the case of conflict between intentions is considered, and corresponding solutions are formulated according to different intention conflict cases. At the same time, considering the limited performance of resources in the network, the intention is quantized, the network state is detected in real time, the user satisfaction is quantized based on the intention and the actual network deployment condition, and the resource utilization rate is also taken into account to reduce resource waste, which guarantees the user experience and saves resources, and guarantees the stability of the network.

[0067] Other advantages, objects and features of the application will be set forth in part in the following specification, and in part will become apparent to those skilled in the art upon examination of the following specification, or can be learned by practice of the application. The objects and other advantages of the application can be realized and attained by the methods and instrumentalities particularly pointed out in the following description. BRIEF DESCRIPTION OF DRAWINGS

[0068] In order to make the objects, technical solutions and advantages of the application clearer, the preferred detailed description of the application will be combined with the drawings to describe the application, in which:

[0069] Figure 1 It is a definition diagram of the intention network IBN;

[0070] Figure 2 It is a schematic diagram of the intention network model structure;

[0071] Figure 3 It is a schematic diagram of the TD3 network structure. DETAILED DESCRIPTION

[0072] The embodiments of the application are described below through specific specific examples, and those skilled in the art can easily understand other advantages and effects of the application from the disclosure of the specification. The application can also be implemented or applied by different specific embodiments, and the details in the specification can be modified or changed based on different views and applications without departing from the spirit of the application. It should be noted that the diagrams provided in the following examples only illustrate the basic concept of the application in a schematic manner, and the following examples and features in the examples can be combined with each other without conflict.

[0073] In the drawings, only for example, the representation is a schematic view, not a real view, and cannot be understood as a limitation of the present application; in order to better illustrate the embodiments of the present application, some components of the drawings will be omitted, enlarged or reduced, and do not represent the actual size of the product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings can be omitted.

[0074] The same or similar reference numerals in the drawings of the embodiments of the present application correspond to the same or similar components; in the description of the present application, it should be understood that if the terms "upper", "lower", "left", "right", "front", "back" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore the positional relationship described in the drawings is only for example, and cannot be understood as a limitation of the present application, for those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0075] As shown in Figure 1 The definition of intent-based networking IBN mainly includes four parts, which are translation and verification, automation implementation, network state awareness, and guarantee and automation optimization / remediation. Among them, translation and verification means that the system obtains higher-level business policies from end users, and converts them into necessary network configurations, generates and verifies the final design and configuration to ensure correctness. Automation implementation means that the system can configure appropriate network changes on the existing network infrastructure, which is completed through network automation or network orchestration. In the network state awareness part, the system provides real-time network state for the systems under its management control, and is protocol and transmission agnostic. Guarantee and optimization means that the system continuously verifies the implementation of the original intent, and takes corrective measures when the required intent cannot be implemented.

[0076] Based on the definition of IBN, the present application designs an intent-based Internet of Things system framework combined with the actual network, which mainly consists of three layers, namely the infrastructure layer, the control plane layer and the application layer. The layer where various terminal Internet of Things devices and network devices are located is the infrastructure layer. The application program plane provides network management functions such as load balancing, traffic monitoring, etc. The controller manages all access points, thereby facilitating the execution of network policies. In addition, in this system, closed-loop management of the network can be realized, the network environment is detected in real time, and the network policy and deployment are continuously optimized and adjusted, thereby guaranteeing the realization of user intent.

[0077] The application layer of the IBN is mainly responsible for collecting the intention input by the user in various forms, and unifying the intention in various forms into a standard form. In the present application, the intention expression is designed in a form similar to natural language. In order to ensure more accurate understanding of the user's intention, the user needs to express the intention demand according to certain rules. The users facing the application layer include but are not limited to ordinary users, network administrators, etc. The intention layer is the core of the IBN and is the most critical factor driving the operation of the IBN. The core component of the intention layer is the intention engine, which is mainly responsible for the analysis and translation of the user's intention, then verifies the network policy according to the current network state information, optimizes the policy and then issues it to the actual network facility. For intention expression and translation, the framework uses INDIRA. It uses a descriptive language to define the intention and uses machine reasoning to understand the user's intention.

[0078] As shown in Figure 2 , the network model is composed of an intention management server, edge nodes and terminal devices. Each edge node is composed of a micro base station (SBS) and a MEC server. Each micro base station has a fixed coverage range, and the overlapping of edge nodes is not considered in this embodiment. The edge nodes are in wired link communication, and the terminal devices use wireless link to transmit data with the corresponding micro base stations.

[0079] In this model, there are M edge nodes in the network, i.e. M MEC servers and M base stations, denote the set of MEC servers. The number of terminal user devices is U, denote the set of terminal devices. It is assumed that each user device can only access a single base station, so the set of user devices associated with base station m can be represented as The decision-making process of the intention management server is set to be carried out in discrete time slots, and each time slot t has a constant duration T s .

[0080] It is assumed that each base station can allocate K types and quantities of network resources, and the types can be represented as The total amount of network resources V K that can be used by each base station can be represented as It is considered to be divided into sub-resource blocks with a capacity of Therefore, the amount of resources allocated to the user device at time slot t can be represented as:

[0081]

[0082] The allocation vector of network resources can be represented as:

[0083]

[0084] Note that at time slot t, the total amount of all resources that can be allocated in the network should be less than the amount that can be allocated in the network, that is:

[0085]

[0086] In this embodiment, the intent referred to is primarily related to network performance. Performance intents can be categorized into constraint intents and optimization intents. Intents requiring certain network performance metrics to be no lower than or no higher than a certain threshold, such as requiring transmission latency to be no higher than a set value, are constraint intents. Intents requiring the maximization or minimization of network performance metrics, such as maximizing network transmission rate, are optimization intents. Constraint intents can be further divided into best-effort intents and guarantee intents. A guarantee intent ensures that network performance metrics reach the desired threshold, while a best-effort intent strives to achieve the desired threshold. The performance metrics described by the performance intent can be represented by a specific performance function, which depends on the allocation of network resources and other performance-related variable sets x. n Examples include link transmission distance and path loss exponent. Therefore, at time slot t, the performance described by the performance intent is expressed as:

[0087]

[0088] Where F n [*] represents the performance function of the performance metric described by the performance intent. The performance function of the conflicting intent is represented as follows: Where l represents the number of performance intents.

[0089] Intent evaluation primarily considers objective factors such as network parameters, including bandwidth, latency, and packet loss, while excluding factors that are difficult to quantify, such as subjective user experiences. Furthermore, different intent service types in an intent network have varying resource requirements, and different users have different sensitivities to resource allocation based on intents. Intent satisfaction is also influenced by the performance metrics mapped to the intent. Therefore, user intent satisfaction at time t is defined as follows:

[0090]

[0091] Where D(t) is the bandwidth at time t, V(t) is the delay at time t, and L(t) is the packet loss at time t. α d ,α v ,α l Especially the calculation coefficients for various network resources.

[0092] The multi-performance intention conflict avoidance problem can be expressed as an optimization problem with the optimization goal of optimization type performance intention and best-effort type performance intention. Meanwhile, the optimization problem takes the guarantee type performance intention as a constraint condition. In addition, by negating the performance function inequality of the threshold type performance intention, the performance function of the intention greater than (or not less than) a certain threshold can be converted into the performance function of the intention less than (or not greater than) a certain threshold. Therefore, the optimization problem is expressed as follows:

[0093]

[0094]

[0095] y i (t)≥1,i∈I nec

[0096] where I nec is the set of guarantee type intentions.

[0097] Time tolerance type intention: The conflict resolution problem of intention network, the existing researches mostly do not consider the time sequence information of intention, but only consider the operation object, behavior and the like involved therein. The processed intention information is unfolded according to the time sequence information, and the operation object, operation action, network performance index and the like involved in the intention are considered, and are mapped into a resource description framework graph (RDF graph). Thus, the time sequence relationship between intentions is more intuitively displayed. The network performance intention referred to in the embodiment is an intention related to performance index, expected state and the like, and it is assumed that the information after splitting of each intention has been obtained. A plurality of conflicting intentions are generated into an RDF graph, and are combined according to the time sequence relationship. According to the resource usage in the network, and considering the demand of each intention on the time sequence, for the intention which does not have strict requirements on time resource, if it competes with other intentions on network resource, the intention is considered to be moved to the adjacent time slice with idle network resource to be implemented. In order to increase the association of each time period of the time-varying intention, the influence of the previous time period on the current time period needs to be increased, that is, the successfully configured intention in the previous time slice, and the intention still existing in the current time slice also needs to be ensured to be successfully configured; when the effective part of the intention which cannot be successfully configured in the previous time slice is also effective in the current time slice, the intention does not need to be configured again.

[0098] Non-time tolerant intention: For the intention conflict that cannot be resolved in time, due to the limited resources, the conflicting intentions cannot be fully satisfied. In order to maximize the satisfaction of users, the resource competition problem is converted into a multi-objective joint optimization problem, considering the resource utilization rate, the performance index of the intention design, the expected state value and the resource utilization rate as the optimization objective, and formulating the optimal resource allocation strategy under the constraint condition. Due to the time-varying nature of the network, the optimization objective function, the constraint condition and the like may change dynamically with time, and the complexity of solving the objective optimization problem using numerical algorithm is high. Therefore, deep reinforcement learning (DRL) is considered to cope with the time-varying characteristics of the network, and the resource allocation strategy that meets the user's intention demand is formulated according to the optimization objective. The Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm is used to customize the resource allocation strategy between the conflicting intentions related to optimization, network performance and the like. As shown in FIG. 8, TD3 is a combination of reinforcement learning and deep learning, and the model principle of reinforcement learning can be specifically described as follows: an agent, in the current state state, interacts with the environment Environment by executing an action, obtains a new state, and at the same time obtains a corresponding reward reward from the environment. The agent evaluates the action executed according to the immediate reward, decides to increase or decrease the reward, and obtains the most reward through the optimization strategy, that is, it is considered to find the optimal strategy. Figure 3

[0099] Reward function:

[0100]

[0101] Wherein γ∈(0, 1) is a discount factor for determining the priority of short-term rewards

[0102] Policy function: π(a|s)=P(A=a|S=s), π(s, a)→[0, 1]

[0103] The state value function is recursively represented by the Bellman equation as follows:

[0104]

[0105] The action value function is recursively represented by the Bellman equation as follows:

[0106]

[0107] ​The TD3 algorithm uses two sets of networks to represent different Q values, and by selecting the minimum one as the target of the update (Target Q Value), it can suppress the continuous overestimation.

[0108] Since the two network parameters are randomly initialized, the output Q values are large or small, and the minimum value is taken as the Q-target or Q-value value, and the update target is:

[0109]

[0110] Update the evaluation network:

[0111]

[0112] In the update of the policy network, only one of the evaluation functions is used to provide gradients for the policy function:

[0113]

[0114] In the present embodiment, the key performance value level, network resource capacity limit, and link state are taken as the state, the parameter setting of the network resource deployment strategy is taken as the action, and the current value of the optimization target is taken as the reward, that is, the performance indicators involved, such as maximizing the average transmission rate of users, the average transmission delay is not higher than a certain threshold, etc. At the same time, in order to avoid excessive consideration of user satisfaction when formulating the strategy and cause resource waste, the resource utilization rate is also considered as the reward, which can reduce resource waste as much as possible while ensuring user intention satisfaction.

[0115] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the purpose and scope of the technical solutions, which should be covered in the scope of the claims of the present application.< / constraint> < / sfc> < / filter>

Claims

1. An improved IBN network-based intention conflict coordination resolution method, characterized in that: The method comprises the following steps: S1: mapping the user's intention request to the level of network resources, and converting the intention conflict problem into a resource allocation problem; S2: dividing the resource allocation problem into time-tolerant intentions and non-time-tolerant intentions in the time dimension; S3: for time-tolerant intentions, adjusting the intention to the adjacent time slice to realize, and increasing the influence of the previous time period on the time period; S4: for non-time-tolerant intentions, converting the multiple performance intention conflict resolution problem into a multi-objective optimization problem about network resource allocation, using a deep reinforcement learning network to continuously perceive and learn the network environment, adjusting and updating the deployment strategy in the network, and solving the intention conflict; The conflict resolution of the time-tolerant intention in step S3 comprises the following steps: Unfold the processed intention information according to the time sequence information, and consider the operation object, operation action and network performance index involved in the intention, and map it to a resource description framework graph RDF; Obtain the information after each intention is split; Generate an RDF graph for multiple conflicting intentions, and merge them according to the time sequence relationship; According to the resource usage in the network and considering the demand of each intention in the time sequence, for the intention that does not have strict requirements on time resources, if it competes with other intentions on network resources, the intention is moved to the adjacent time slice with idle network resources to realize; Increase the influence of the previous time period on the time period, that is, the successfully configured intentions in the previous time slice and the intentions still existing in the current time slice also need to be successfully configured; when the effective part of the intention that cannot be successfully configured in the previous time slice is also effective in the current time slice, the intention is no longer configured; In step S4, for the intention conflict problem that cannot be resolved in time, it is converted into a multi-objective joint optimization problem, considering resource utilization, taking the performance index, expected state value and resource utilization of the intention design as optimization objectives, and formulating an optimal resource allocation strategy under the constraint condition; A double-delay deep deterministic policy gradient algorithm TD3 is used to customize the resource allocation strategy of the conflicting intention related to optimization and network performance; the TD3 model is described as an agent, which interacts with the environment by executing an action in the current state, obtains a new state, and obtains a corresponding reward from the environment; the agent evaluates the action executed according to the immediate reward, decides to increase or decrease the reward, and obtains the maximum reward through the optimization strategy, that is, the optimal strategy is found; The reward function is: wherein a discount factor for determining a short-term reward priority; The policy function is: The state value function is recursively represented by the Bellman equation: The action value function is recursively represented by the Bellman equation: The TD3 algorithm uses two sets of networks to represent different Q values, and selects the smallest one as the update target; The smaller one is taken as the Q-target or Q-value, and its update target is: Update the evaluation network: In the update policy network, only one of the evaluation functions is used to provide gradients for the policy function: The key performance numerical level, network resource capacity limit, and link state are used as the state, the parameter setting of the network resource allocation strategy is used as the action, and the current value of the optimization target and the resource utilization are used as the reward.

2. The method of claim 1, wherein: The intention request of the user is mapped to the level of network resources in step S1, and the intention conflict problem is converted into a resource allocation problem. First, an intention Internet of Things system is constructed, including an intention management server, edge nodes, and terminal devices. The edge node is composed of a micro base station SBS and a MEC server; the intention management server communicates with the edge nodes wirelessly, and the edge nodes communicate with each other through a wired link; and each terminal device communicates with the corresponding micro base station wirelessly. The intention management server analyzes and translates the user's intention through the intention engine, then verifies the network policy according to the current network state information, optimizes the policy, and then issues it to the actual network facility; The terminal device is used to collect the intentions input by the user in various forms, and unify the intentions in various forms into a standard form; Definition of the network with N edge nodes, i.e. N MEC servers and M base stations, denotes the set of MEC servers; the number of end user devices is , denotes the set of end devices; the set of user devices associated with the base station is denoted by , ; the decision process of the intent management server is set to take place in discrete time slots, each time slot has a constant duration ; Assume each base station can allocate the same type and number of K kinds of network resources, whose type is represented as , the total amount of network resources that each base station can use is represented as , the total amount of network resources that each base station can use is represented as , divide into sub-resource blocks with a capacity of , and the amount of resources allocated to the user equipment in time slot is represented as: (1) The allocation vector of network resources is represented as: (2) At the time slot The total amount of all resources that can be allocated in the network at the time is lower than the amount that can be allocated in the network, i.e. (3)。 3. The method of claim 1, wherein: The user's intention is defined as a performance intention related to network performance, and the performance intention is divided into constraint type intention and optimization type intention, and the constraint type intention is divided into best effort type intention and guaranteed type intention, wherein: The constraint type intention is an intention that requires certain network performance to be no lower than or no higher than a certain threshold value; The optimization type intention is an intention that requires the network performance index to be maximized or minimized; The best effort type intention is an intention that makes the network performance index try to reach the expected threshold value; The guaranteed type intention is an intention that requires the network performance index to reach the expected threshold value.

4. The method of claim 3, wherein the improved IBN network-based intent conflict coordination resolution method is characterized by: The performance indicators described by the performance intent depend on the allocation of network resources and on other sets of variables related to performance At the time slot The performance indicators described by the performance intent are expressed as: (4) wherein is a performance function of the performance indicator described by the performance intent; the performance function of the conflicting intent is denoted as wherein denotes the number of performance intents; The user's intention satisfaction at time t is defined as follows: (5) Wherein, is Bandwidth at the moment, is Delay at the moment, is Packet loss at the moment; In addition, the calculation coefficient of each network resource; The multi-performance intention conflict avoidance problem is represented as an optimization problem with optimization type performance intention and best effort type performance intention as optimization objectives, and guaranteed type performance intention as constraint condition. By taking the inverse of the performance function inequality of the threshold type performance intention, the performance function of the intention greater than or not less than a certain threshold value is converted into the performance function of the intention less than or not greater than a certain threshold value. The optimization problem is expressed as follows: (6) wherein is a set of assurance intents.

Citation Information

Patent Citations

  • Intention-driven wireless network resource conflict resolution method and device

    CN115119332A

  • System and method for resolving conflict

    US20080008983A1