A network self-governance method and system for self-aware agents
Patent Information
- Application Number
- CN202610890554.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]本申请提供一种基于自我认知的智能体的网络自治方法和系统,能够解决现有网络智能体的行为不可控且不可靠,导致无法安全且可信地实现高阶网络自治的技术问题
[0014]本申请实施例提供的技术方案带来的有益效果包括:
Smart Images

Figure CN122824601A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network intelligent agent technology, specifically to a network autonomous method and system for intelligent agents based on self-awareness. Background Technology
[0002] With the rapid growth of network scale, the continuous complexity of service types, and the constant rise in operating costs, building autonomous networks with self-configuration, self-optimization, and self-healing capabilities has become a common goal for operators worldwide. Industry organizations such as the TM Forum (TMF) have proposed a hierarchical system for network autonomy capabilities, ranging from L0 (fully manual) to L5 (fully autonomous), providing a clear evolution path for the industry. Currently, major global operators are driving their networks from "tool-based automation" to "intelligent closed-loop" through automated configuration, intelligent alarms, and cross-domain resource orchestration. However, their autonomy capabilities are mostly at the L2-L3 level, still significantly lagging behind the higher-order (L4 / L5) fully autonomous networks. Overall, this field urgently needs to achieve a higher degree of autonomous decision-making, closed-loop control, and intelligent operation and maintenance in complex, dynamic, heterogeneous network environments to address the challenges brought by scaling and complexity. In the technological path towards higher-order autonomous networks, AI agent-based architectures are considered a key enabling technology. In existing solutions, agents can simulate a closed loop of "observation-analysis-decision-execution," achieving a certain degree of automation within specific network domains. The industry has begun exploring multi-agent collaborative frameworks, attempting to enable agents from different network domains (such as RAN, core network, and transport network) to collaborate, achieving cross-domain closed loops and end-to-end optimization. Furthermore, some leading vendors have proposed their own autonomous network solutions, such as building a unified AI model layer, intent-driven networks, or achieving higher levels of automation in specific scenarios (such as data center networks). Despite the correct direction of agent technology, existing agent-based solutions still face the following challenges when applied to complex carrier-grade networks: (1) In complex and dynamic network environments, the decision-making logic of existing intelligent agents may produce unstable or unreproducible outputs, and their behavior lacks consistency. This makes the actions and results of intelligent agents difficult to predict and fails to meet the high reliability requirements of telecommunications-grade networks; (2) Existing intelligent agent architectures usually focus on responding to and executing external tasks, which may make decisions that are effective locally but detrimental to global or long-term goals, or produce harmful operations when attacked or misled, posing a risk of “going out of bounds”, and the security and compliance of their behavior are difficult to guarantee. (3) Existing intelligent agents' decision-making is mostly based on data-driven models, and their internal reasoning processes are difficult to provide decision-making basis and action chains that human administrators can understand. This seriously hinders the establishment of necessary trust in critical network scenarios and makes problem tracing, responsibility definition, and regulatory auditing extremely difficult; (4) Existing intelligent agents often rely on pre-set, large amounts of context or fixed knowledge bases, and may fail if the environment changes slightly, showing vulnerability. At the same time, in multi-agent collaborative scenarios, there is a lack of effective mechanisms to ensure the consistency of goals and the coordination of behaviors of each agent, which can easily lead to policy conflicts and make it difficult to achieve global optimization. Summary of the Invention
[0003] This application provides a network autonomy method and system based on self-awareness intelligent agents, which can solve the technical problem that the behavior of existing network intelligent agents is uncontrollable and unreliable, making it impossible to achieve high-order network autonomy safely and reliably.
[0004] In a first aspect, embodiments of this application provide a network autonomy method for intelligent agents based on self-awareness, comprising: Constraint criteria are preset and stored in the agent; if the agent's own operating state and behavior results deviate from the constraint criteria, the first optimization requirement derived from the self-correction intention is generated. If, after anticipating the evolution trend of the network environment in which the agent is currently located, a problem that needs to be addressed is identified, then a second optimization requirement arising from changes in the external environment is generated. The first and second optimization requirements are comprehensively evaluated to obtain a unified executable goal that meets the constraint criteria. The unified executable goal is planned as a specific sequence of behaviors and executed to drive the network to autonomously evolve toward a state that meets the constraint criteria.
[0005] In conjunction with the first aspect, in one implementation, the constraint criteria include fixed constraint criteria and immediate constraint criteria; wherein, the immediate constraint criteria are criteria defined or modified based on the fixed constraint criteria during the operation of the agent.
[0006] In conjunction with the first aspect, in one implementation, if the agent's own operating state deviates from the constraint criteria of the behavioral outcome, a first optimization requirement derived from a self-correction intention is generated, which includes the following steps: The agent monitors its own operating state and behavioral results to identify whether it deviates from fixed constraint criteria or immediate constraint criteria. When a deviation is detected, the system performs a self-diagnosis to determine the type of deviation and generates a correction requirement corresponding to the deviation type as the first optimization strategy.
[0007] In conjunction with the first aspect, in one implementation, the first optimization strategy includes action correction requirements for modifying specific execution actions, target adjustment requirements for adjusting the current task objective, and criterion update requirements for updating immediate constraint criteria.
[0008] In conjunction with the first aspect, in one implementation, if the expected evolution trend of the network environment in which the agent is currently located presents a problem that needs to be addressed, a second optimization requirement arising from changes in the external environment is generated, which includes the following steps: Acquire raw data of the network environment in which the agent is located, and process the raw data to form a structured representation of the current state of the network environment; Based on structured representations and stored domain knowledge, scenario assessment is performed to construct a context describing the operational logic and relationships of the network environment. Based on context, predictive models are used to estimate future changes in the state of the network environment; Identify events that the agent needs to handle during future state changes; generate corresponding handling strategies for these events as a second optimization requirement.
[0009] In conjunction with the first aspect, in one implementation, a comprehensive evaluation of the first optimization requirement and the second optimization requirement is performed to form a unified executable objective that conforms to the constraint criteria, which includes the following steps: The first optimization requirement and the second optimization requirement are transformed into internal candidate objectives and external candidate objectives, respectively, and then merged into a set of candidate objectives; The overall effect of each objective in the candidate objective set is evaluated through prediction models and simulations. Objectives that meet the fixed constraint criteria and the immediate constraint criteria are selected from the candidate objective set as executable objectives.
[0010] In conjunction with the first aspect, in one implementation, a unified executable objective is planned as a specific sequence of behaviors and executed to drive the network to autonomously evolve toward a state that conforms to the constraint criteria, which includes the following steps: For a unified achievable objective, a sequence of actions to achieve the achievable objective is generated by combining fixed constraint criteria, immediate constraint criteria, and a predictive model. The system executes a sequence of actions, monitors the execution process and results in real time, and forms an autonomous closed loop.
[0011] In conjunction with the first aspect, one implementation also includes establishing a knowledge base, in which constraint criteria and knowledge sets are centrally stored and updated; the knowledge sets include general knowledge, agent-specific knowledge, and normative value knowledge.
[0012] In conjunction with the first aspect, in one implementation, the method further includes constructing a knowledge engine; updating the knowledge set in the knowledge base based on the knowledge engine, combining the agent's observation of the network environment, the results of its behavior execution, and external interaction information; and iteratively optimizing the prediction model and decision-making strategy in the knowledge base based on the experience collected during the agent's operation.
[0013] Secondly, embodiments of this application provide a network autonomous system based on a self-aware intelligent agent, comprising: The self-awareness module is used to preset and store constraint criteria in the agent; if there is a deviation between the agent's own operating state and the behavioral results and the constraint criteria, the first optimization requirement derived from the self-correction intention is generated. The contextual awareness module is used to generate a second optimization requirement based on changes in the external environment if there are problems that need to be addressed in the expected evolution trend of the network environment in which the agent is currently located. The decision-making and action module is used to comprehensively evaluate the first and second optimization requirements to form a unified executable goal that conforms to the constraint criteria; the unified executable goal is planned as a specific sequence of behaviors and executed to drive the network to autonomously evolve toward a state that conforms to the constraint criteria.
[0014] The beneficial effects of the technical solutions provided in this application include: A network autonomy method based on self-awareness of intelligent agents is proposed, in which constraint criteria are preset and stored in the intelligent agent, defining what the intelligent agent should achieve in the long term, and ensuring that all its behaviors have a unified and correct judgment standard. Secondly, when the agent is working, it continuously compares its own behavior with the constraint criteria. If a deviation is found, it does not take corrective action directly, but transforms the found deviation into a clear internal correction demand signal, namely the first optimization demand. At the same time, the agent also analyzes the current network environment and predicts what situations might occur. After predicting risks, it does not immediately take corrective action, but transforms the identified risks into another clear external scenario demand signal, namely the second optimization demand. That is, after discovering various problems, it does not take corrective action directly, but only generates corresponding demands. Finally, under the established constraint criteria, the generated first and second optimization requirements are comprehensively evaluated and integrated to generate a unified executable goal. This ensures that the executable goal satisfies the constraint criteria and represents a trade-off and synthesis of multiple optimization intentions under the constraints of the criteria. Subsequently, the executable goal is planned as an action and executed. The execution results are fed back, triggering further awareness of the network's own state and environment, thereby driving the network to continuously and autonomously evolve towards a better state while satisfying the constraint criteria. This solves the technical problem of uncontrollable and unreliable behavior of existing network agents, which makes it impossible to achieve high-order network autonomy securely and reliably. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the network autonomy method for self-aware intelligent agents in the embodiments of this application; Figure 2 This is a schematic diagram of the overall architecture of the network autonomous system based on self-awareness intelligent agents in the embodiments of this application; Figure 3 This is a flowchart of the network autonomous system based on self-awareness of intelligent agents in the embodiments of this application. Detailed Implementation
[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0017] To make the technical problem that this application aims to solve clearer, the causes of the technical problem will be analyzed in detail below: Existing agent architectures typically tightly couple environmental perception, task planning, and behavior execution. When abnormal or suboptimal behaviors occur in complex network environments, it is difficult to accurately determine whether the behavior stems from erroneous environmental perception, unreasonable task planning, or a flawed policy model. This indivisibility and uninterpretability of decision-making logic makes the agent's behavior appear as a whole. Once an error occurs, it is impossible to perform fine-grained attribution and correction, nor can effective and understandable constraints be imposed before the behavior occurs, thus making its behavior inherently uncontrollable. Existing solutions typically drive agent behavior from a single source, or are passively triggered entirely by external task requirements (such as manual instructions or alarm events), or are entirely driven by internal model objectives (such as maximizing the reward function). The former turns the agent into a highly automated tool, lacking initiative and long-term goals; the latter may lead to agent behavior becoming disconnected from real-time network requirements, or even harming the overall network in pursuit of internal goals. Due to the lack of a mechanism to simultaneously and equally address both the internal self-optimization needs and the external environmental adaptation needs, agent decision-making is prone to bias and short-sightedness, resulting in unreliable performance in multi-objective, long-cycle, and dynamically changing network environments. Some existing intelligent agents lack behavioral norms (such as security ethics, operational principles, and long-term missions), and their behavior is determined solely by immediate policies, easily leading to unpredictable and out-of-bounds actions. Even if rules are preset, they cannot be adjusted according to changes in the agent's own capabilities, the evolution of the network environment, and the accumulation of operational experience. This state of lacking norms or having rigid norms means that intelligent agents either act without a basis for action or behave inappropriately when the rules become invalid due to environmental changes.
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0019] In a first aspect, embodiments of this application provide a network autonomy method for intelligent agents based on self-awareness, comprising: S100. Preset and store constraint criteria in the agent; if the agent's own operating state and behavior result deviate from the constraint criteria, generate the first optimization requirement based on the self-correction intention. S200. If there are problems that need to be addressed in the expected evolution trend of the network environment in which the agent is currently located, then a second optimization requirement arising from changes in the external environment is generated. S300. Conduct a comprehensive evaluation of the first and second optimization requirements to obtain a unified executable objective that meets the constraint criteria; plan the unified executable objective as a specific sequence of behaviors and execute it to drive the network to autonomously evolve toward a state that meets the constraint criteria.
[0020] By setting up this method, where constraint criteria are preset and stored in the agent, the long-term goal of the agent is defined, ensuring that all its actions have a unified and correct judgment standard. Secondly, when the agent is working, it continuously compares its own behavior with the constraint criteria. If a deviation is found, it does not take corrective action directly, but transforms the found deviation into a clear internal correction demand signal, namely the first optimization demand. At the same time, the agent also analyzes the current network environment and predicts what situations might occur. After predicting risks, it does not immediately take corrective action, but transforms the identified risks into another clear external scenario demand signal, namely the second optimization demand. That is, after discovering various problems, it does not take corrective action directly, but only generates corresponding demands. Finally, under the established constraint criteria, the generated first and second optimization requirements are comprehensively evaluated and integrated to generate a unified executable goal. This ensures that the executable goal satisfies the constraint criteria and represents a trade-off and synthesis of multiple optimization intentions under the constraints of the criteria. Subsequently, the executable goal is planned as an action and executed. The execution results are fed back, triggering further awareness of the network's own state and environment, thereby driving the network to continuously and autonomously evolve towards a better state while satisfying the constraint criteria. This solves the technical problem of uncontrollable and unreliable behavior of existing network agents, which makes it impossible to achieve high-order network autonomy securely and reliably.
[0021] Furthermore, in one embodiment, the constraint criteria include fixed constraint criteria and immediate constraint criteria; wherein, the immediate constraint criteria are behavioral criteria defined or modified based on the fixed constraint criteria during the operation of the intelligent agent.
[0022] In this embodiment, fixed constraint criteria can be understood as a mission of the intelligent agent. These are the highest-level behavioral norms pre-set and stored by the designer during the intelligent agent system design or initialization phase. They clarify the fundamental purpose, core value orientation, and long-term operational direction of the intelligent agent, providing the ultimate value judgment standard and behavioral boundaries for the entire intelligent agent. The legitimacy and compliance of all subsequently generated immediate criteria, specific goals, and actions must ultimately be retrospectively verified against the fixed constraint criteria. This solves the fundamental questions of why the intelligent agent's behavior is controllable and where it should evolve. Immediate constraint criteria can be understood as the intelligent agent's meta-goals. These are operational norms dynamically generated, maintained, and modified during the intelligent agent's operation, based on fixed constraint criteria and combined with real-time awareness of its own state and environment. They guide recent or current specific decisions and behaviors. When the intelligent agent discovers a deviation between its current behavior, goals, or strategies and the fixed constraint criteria through self-monitoring and diagnosis, it triggers the evaluation and adjustment of the immediate constraint criteria. For example, to ensure network fairness (fixed constraint criteria), an immediate constraint criterion of "prioritizing certain types of critical business traffic" may be dynamically generated in specific congestion scenarios. This modification ensures that the immediate constraint criterion is always a specific and adaptive manifestation of the fixed constraint criterion in the current context.
[0023] Furthermore, in one embodiment, if the agent's own operating state deviates from the constraint criteria regarding the behavioral outcome, a first optimization requirement derived from a self-correction intention is generated, which includes the following steps: The agent monitors its own operating state and behavioral results to identify whether it deviates from fixed constraint criteria or immediate constraint criteria. When a deviation is detected, the system performs a self-diagnosis to determine the type of deviation and generates a correction requirement corresponding to the deviation type as the first optimization strategy.
[0024] In this embodiment, self-monitoring refers to the agent continuously comparing its own operating state and behavioral results according to fixed constraint criteria and immediate constraint criteria to identify whether there are deviations. Self-diagnosis, on the other hand, involves analyzing the root causes of deviations after they are detected and classifying them into a structured first optimization strategy.
[0025] Furthermore, in one embodiment, the first optimization strategy includes action correction requirements for modifying specific execution actions, target adjustment requirements for adjusting the current task objective, and criterion update requirements for updating immediate constraint criteria.
[0026] In this embodiment, specifically, action correction requirements address problems at the level of specific action execution; goal adjustment requirements address problems of unreasonable or conflicting task goals; and criterion update requirements address problems of contradictions between current immediate constraint criteria and fixed constraint criteria. This mechanism ensures that the agent can accurately identify its own problems at different levels of behavior, goals, and criteria, and output clear correction requirements, rather than arbitrarily executing potentially erroneous correction actions, thereby improving the interpretability of system behavior, decision stability, and long-term operational reliability.
[0027] Furthermore, in one embodiment, if the expected evolution trend of the network environment in which the agent is currently located presents a problem that needs to be addressed, a second optimization requirement arising from changes in the external environment is generated, which includes the following steps: Acquire raw data of the network environment in which the agent is located, and process the raw data to form a structured representation of the current state of the network environment; Based on structured representations and stored domain knowledge, scenario assessment is performed to construct a context describing the operational logic and relationships of the network environment. Based on context, predictive models are used to estimate future changes in the state of the network environment; Identify events that the agent needs to handle during future state changes; generate corresponding handling strategies for these events as a second optimization requirement.
[0028] In this embodiment, raw network data (such as traffic, KPIs, and alarms) is acquired and processed to form a structured representation. Next, this structured representation is deeply analyzed using domain knowledge from a knowledge base to construct a context that reveals the network's operational logic and relationships. Based on this context, a pre-built prediction model is used to predict future network state changes. Finally, the agent identifies events requiring intervention from the prediction results and transforms them into a clear second optimization requirement.
[0029] Furthermore, in one embodiment, the first optimization requirement and the second optimization requirement are comprehensively evaluated to form a unified executable objective that conforms to the constraint criteria, which includes the following steps: The first optimization requirement and the second optimization requirement are transformed into internal candidate objectives and external candidate objectives, respectively, and then merged into a set of candidate objectives; The overall effect of each objective in the candidate objective set is evaluated through prediction models and simulations. Objectives that meet the fixed constraint criteria and the immediate constraint criteria are selected from the candidate objective set as executable objectives.
[0030] In this embodiment, the process of comprehensively evaluating the first and second optimization requirements to form an executable goal is a concrete manifestation of the decision-making process of this method. This process first involves goal transformation and merging. The first optimization requirement, stemming from a self-correction intent, is interpreted as a set of internal candidate goals, while the second optimization requirement, arising from environmental changes, is derived as a set of external candidate goals. The decision-making module uses predictive models and network simulation to proactively assess the potential system evolution, benefits, costs, and risks that may result from the execution of each candidate goal in the set. Within a decision-making framework comprised of fixed and immediate constraint criteria, various goals are comprehensively weighed. Finally, the goal that best aligns with the fundamental mission, value system, and security constraints, and delivers the best overall effect, is selected from the candidate set and formally designated as the executable goal for the next stage of system action. This mechanism ensures that the agent's final decision is scientific, controlled, and globally optimal, serving as the decision engine driving the network to achieve reliable, controllable, and autonomous operation.
[0031] Furthermore, in one embodiment, a unified executable objective is planned as a specific sequence of behaviors and executed to drive the network to autonomously evolve toward a state that conforms to the constraint criteria, which includes the following steps: For a unified achievable objective, a sequence of actions to achieve the achievable objective is generated by combining fixed constraint criteria, immediate constraint criteria, and a predictive model. The system executes a sequence of actions, monitors the execution process and results in real time, and forms an autonomous closed loop.
[0032] In this embodiment, the executable objective is transformed into actual actions, forming a closed loop. For a defined, unified executable objective, under strict constraints of fixed and immediate constraints, and using a predictive model to deduce the effects of various action combinations, a safe, feasible, and orderly sequence of specific actions is generated. This sequence is executed step-by-step via API calls, with real-time monitoring of the execution status and final network result at each step. The data produced at this stage is immediately fed back to the agent, triggering necessary self-correction and re-decision-making. Thus, the endpoint of a single action becomes the starting point for a new round of cognition and optimization, forming a complete, self-driven, autonomous closed loop. This ensures that the agent not only makes excellent decisions but also executes them reliably and learns continuously from the execution results, thereby driving the network to evolve autonomously and robustly in directions that conform to the constraint criteria.
[0033] Furthermore, in one embodiment, the method further includes establishing a knowledge base to centrally store and update constraint criteria and knowledge sets; the knowledge sets include general knowledge, agent-specific knowledge, and normative value knowledge. In this embodiment, a knowledge module consisting of a knowledge base and a knowledge engine is constructed and maintained to provide intelligent support for the entire autonomous process. The knowledge base centrally stores constraint criteria and a complete set of knowledge. This set specifically includes: general knowledge, such as network protocols and domain terminology; agent-specific knowledge, such as the agent's executable actions, historical goals, and real-time self-model (recording its capabilities, state, and performance); and normative value knowledge, such as security boundaries, regulations, and value systems used for multi-objective trade-offs.
[0034] Furthermore, in one embodiment, the method further includes constructing a knowledge engine; updating the knowledge set in the knowledge base based on the knowledge engine, combining the agent's observation of the network environment, the results of its behavior execution, and external interaction information; and iteratively optimizing the prediction model and decision-making strategy in the knowledge base based on the experience collected during the agent's operation.
[0035] In this embodiment, the knowledge engine is responsible for driving the activation and evolution of knowledge. It performs two main functions: first, knowledge updating, which involves adding, deleting, and modifying factual content in the knowledge base based on the agent's observations of the environment, the results of its own actions, and interaction information with external entities (humans or other intelligent agents) to maintain its freshness; second, iterative optimization of models and strategies, which involves triggering continuous training and tuning of the prediction models and decision-making strategies within the knowledge base based on the success and failure experience data collected during operation, and writing back the optimization results. This mechanism ensures that the knowledge system supporting the agent's cognition and decision-making is continuously learning and updating, thus laying the foundation for achieving reliable, controllable, and continuously self-improving autonomy in complex networks.
[0036] Secondly, embodiments of this application also provide a network autonomous system based on a self-aware intelligent agent, comprising: The self-awareness module is used to preset and store constraint criteria in the agent; if there is a deviation between the agent's own operating state and the behavioral results and the constraint criteria, the first optimization requirement derived from the self-correction intention is generated. The contextual awareness module is used to generate a second optimization requirement based on changes in the external environment if there are problems that need to be addressed in the expected evolution trend of the network environment in which the agent is currently located. The decision-making and action module is used to comprehensively evaluate the first and second optimization requirements to form a unified executable goal that conforms to the constraint criteria; the unified executable goal is planned as a specific sequence of behaviors and executed to drive the network to autonomously evolve toward a state that conforms to the constraint criteria.
[0037] The workflow of the self-awareness module is as follows: self-monitoring, self-diagnosis, self-regulation, and goal generation; The self-monitoring function continuously tracks the agent's internal state, including: the agent's behavioral outcomes, goal progress, capability constraints, and the degree of deviation from its mission. This deviation may manifest as performance degradation, goal conflict, resource shortages, unexpected behavior, or goal misalignment. Since the agent's state and capabilities are stored in the knowledge module, the monitoring process essentially verifies whether the agent is still operating according to its mission and normative constraints. Once a deviation is detected, the agent enters a self-diagnosis phase rather than directly implementing reverse correction, thus avoiding blind error correction.
[0038] Self-diagnosis, in essence, involves intelligently diagnosing and analyzing the causes of deviations after they are detected, and then formulating a need to correct them. Self-diagnosis comprises the following layers: If the deviation stems from a failure in action, then a behavioral-level correction is needed. If the issue stems from the goal being unattainable or conflicting with existing goals, then a goal-level revision is required. If the conflict stems from a conflict of value principles or constraints, then a meta-objective-level correction requirement arises. If the demand arises from changes in capabilities or environment, then a higher-level need emerges. Requirements define what needs to be fixed, not how to fix it. Using requirements instead of direct corrective actions ensures greater interpretability, traceability, and stability of agent behavior.
[0039] The self-regulation phase determines whether the requirements align with the agent's mission, principles, rules, and constraints, preventing behavioral corrections from deviating from the overall mission. The knowledge base stores the following principles and norms concerning the agent itself: Mission and long-term constraints, operational and behavioral boundaries, regulations and normative rules, as well as security boundaries and resource capabilities.
[0040] The self-regulation module, based on constraints and norms in the knowledge base, performs consistency and feasibility analysis on requests. It then generates meta-goals that reflect the nature of the requirements and satisfy behavioral norm constraints. Meta-goals describe how agents should behave in a regulated manner, embodying the agent's normative requirements. For example, in an autonomous network environment, rules might include maintaining fairness in network resource services, prioritizing critical traffic, limiting risks, and conserving resources. The role of meta-goals is to ensure that subsequently generated goals do not deviate from normative constraints.
[0041] The goal formation function translates meta-goals into actionable task objectives. The steps include: extracting possible implementation methods from the knowledge base, evaluating the cost, risks, and resource requirements of each solution, filtering out unachievable goals, and forming a candidate goal set. Goal formation addresses the question of what specific actions can be taken to satisfy the meta-goal. Finally, it outputs a set of achievable candidate goals.
[0042] The workflow of the contextual cognition module is as follows: contextual cognition consists of three stages: perception, reasoning, and prediction.
[0043] Perception, that is, perceiving data in the environment. Perceiving data in the network environment. This data includes: Information from network devices, such as traffic, key performance indicators, alarms, and logs; and externally monitored information that may affect the network, such as weather and temperature. This raw data undergoes processing (typically based on machine learning methods) including feature extraction, anomaly detection, traffic classification, and semantic enhancement to form a structured representation of the current state of the environment. Furthermore, the perception results can be combined with domain knowledge from a knowledge base for semantic enhancement to improve the accuracy of environmental understanding.
[0044] Reasoning, in this contextual sense, involves constructing the environmental context by querying a knowledge base after receiving perceived data and then performing contextual reasoning. In networks, context represents the current network state, network behavior patterns, possible interactions between different entities, topology-related constraints, and patterns or interpretations from the network's historical knowledge. This enables agents not only to see data but also to understand its meaning, which is the true purpose of contextual reasoning. Subsequently, environmental prediction models are built based on contextual information, such as traffic evolution models, fault propagation models, topology change risk models, and performance indicator evolution models. Predictive models can be built using machine learning, rules, knowledge graphs, causal analysis, multi-agent simulation, and other methods.
[0045] Prediction, based on the predictive model built in the inference phase, forecasts future scenario states, generates estimates of future states, and provides corresponding task requirements based on these forecasts. Examples include: future traffic load prediction, risks of failure or performance degradation, trends of equipment stress or resource shortages, SLA default probability, and the future impact of environmental conditions (e.g., temperature changes) on the network. After obtaining the future state, the agent identifies potential future challenges or opportunities, thus forming corresponding task requirements, such as: "Congestion will occur in the next 10 minutes, requiring advance traffic diversion," "Link utilization will reach a threshold, requiring capacity expansion," and "KPI degradation is predicted, requiring parameter adjustment or route switching." The core task of this stage is mapping future predictions into manageable requirements. What the agent obtains from the external world are requirements; requirements have no mission and do not contain value, simply indicating "a problem that needs to be solved." Examples include: future traffic load prediction, risks of failure or performance degradation, trends of equipment stress or resource shortages, SLA default probability, the future impact of environmental conditions (e.g., temperature changes) on the network, and simulation of the effects of small-scale topology adjustments.
[0046] The workflow of the decision-making and action module is as follows: goal management, planning, and execution. The core function of the decision-making module is to comprehensively reason about external task requirements and internal self-goals, and to formulate actionable decision objectives under the constraints of rules, strategies, and value systems. The decision-making stage integrates three types of inputs: external task requirements from situational cognition, candidate objectives from self-cognition, and decision principles, including rules and norms, value systems, strategy preferences, and risk constraints.
[0047] For external task requirements, the decision-making module first interprets the requirements and derives objectives, transforming the requirements into a set of candidate objectives. Subsequently, this set of objectives, along with the internal objectives generated by the self-awareness module, enters the objective evaluation and selection process. During the objective evaluation process, the decision-making module can utilize predictive models and simulation techniques to assess the potential system evolution outcomes of different objectives, thereby selecting the objective that best aligns with the mission, value system, and security constraints.
[0048] Candidate targets must also be evaluated and ranked according to a value system. This process comprehensively considers: the benefits, costs (resource consumption, adjustment costs), risks (security, controllability), and normative constraints (environmental costs, regulatory requirements). Even if a target is beneficial to the agent, it may be rejected if it does not meet the basic thresholds of the value system assessment. Based on these conditions, the value assessment function ultimately selects suitable targets, which are then handed over to the planning and execution functions of the decision-making module for implementation. The function of the planning phase is to calculate the control strategy or action plan to achieve the goal selected in the decision-making phase, based on the comprehensive predictive model and system constraints. According to the control strategy and changes in the environmental state, the planning phase generates a series of action sequences to guide the system from its current state to the desired state. The planning process not only considers the feasibility of achieving the goal but also explicitly incorporates safety and risk constraints to prevent the system from entering a high-risk or uncontrollable state. The execution phase is responsible for executing various operations step by step according to the action sequence generated in the planning phase, calling external tools or interfaces. During execution, the agent continuously monitors the execution results and sends the execution status and feedback information back to the self-awareness module to monitor execution deviations, strategy failures, or environmental changes, thereby triggering necessary self-correction and re-decision, forming a complete autonomous closed loop.
[0049] In addition, the system also includes the knowledge module mentioned above. The knowledge module is used to store, manage and evolve various types of knowledge that support the agent's cognition, decision-making and autonomous behavior. It is the foundation for realizing self-cognition, situational cognition and rational decision-making. Among them, the knowledge base is used to store structured and unstructured knowledge, covering general knowledge, agent-specific knowledge and normative and value knowledge, providing a unified knowledge foundation for reasoning, learning and decision-making. Furthermore, general knowledge comprises domain-specific and technical knowledge shared by all network agents, including: a domain classification system, which describes the domain terminology and hierarchical relationships of networks, topologies, protocol stacks, and hardware and software components; system attributes and requirements, which are the technical attributes that agents must meet in terms of security, reliability, and performance; scientific and engineering knowledge, which includes rules, algorithm libraries, and problem-solving methods extracted from standards, literature, and engineering practices; agent-specific knowledge, which is closely related to the agent's specific role and operating state, mainly including behavioral and goal knowledge, agent self-model, and normative and value-based knowledge; and behavioral and goal knowledge, which includes executable actions and their triggering conditions, a set of pursueable goals, and external requirements and their corresponding meta-goals.
[0050] In autonomous network agents, the self-model is a structured representation of the agent's own state, behavioral patterns, and behavioral outcomes. For agents in an autonomous network, the state of network resources, the agent's current goals and candidate goal lists, and the key performance indicators (KPIs) and performance metrics generated during network operation constitute the core components of the self-model, serving as a necessary foundation for self-awareness and self-regulation. Normative and value-based knowledge includes: operational and behavioral boundaries, defining the scope and boundaries of the agent's behavior; normative rules, defining legal, ethical, and engineering norms; engineering experience knowledge, expert knowledge summarized in the form of rules of thumb; and a value system, used to evaluate the costs and benefits of actions at the economic, social, or collaborative levels. The value system provides a basis for multi-objective trade-offs, enabling agents to make rational choices between short-term gains and long-term synergy.
[0051] The knowledge engine is responsible for reasoning, updating, and evolving the knowledge base, and is the core mechanism for realizing the cognitive abilities of an intelligent agent. It does not directly make the final decision, but rather provides knowledge support and reasoning results for the self-awareness, situational awareness, and decision-making modules. Its core functions include: knowledge updating, complex reasoning, and continuous learning. Knowledge updating refers to the structured updating of the knowledge base by the knowledge engine based on environmental observations, execution feedback, and event information to maintain the timeliness, consistency, and traceability of knowledge. The knowledge updating process includes data collection, information extraction, knowledge alignment, knowledge verification, knowledge writing, and version management. Complex reasoning refers to the multi-level reasoning performed by the knowledge engine based on the knowledge base to support contextual understanding, root cause analysis, and decision support; it is the core source of cognitive intelligence for autonomous agents. Typical processes include: reasoning request input, context aggregation (extracting relevant knowledge), hybrid reasoning execution, and reasoning result verification. Hybrid reasoning execution includes LLM reasoning, knowledge graph multi-hop reasoning, rule-based reasoning, and causal reasoning. Reasoning result verification involves outputting reasoning conclusions or candidate solutions.
[0052] Continuous learning means that the agent can continuously update its model, policies, and capabilities during operation, achieving self-evolution rather than simply accumulating knowledge. A typical process includes: Experience collection: successful strategies, failed strategies, execution logs, reward signals, etc.; Environmental representation: converting experience into learnable state-action data; Learning trigger: triggering learning when deviation exceeds a threshold, policy performance deteriorates, or a new pattern emerges; Model update: fine-tuning the LLM model, updating the policy network, updating the prediction model, and updating the knowledge representation; Stability check: avoiding catastrophic forgetting; Knowledge synchronization: writing the learning results back to the knowledge base or policy base.
[0053] The functions of each module in the above-mentioned intelligent agent network autonomous system correspond to the steps in the above-mentioned intelligent agent network autonomous method embodiment, and their functions and implementation processes will not be described in detail here.
[0054] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0055] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.
[0056] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.
[0057] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.
[0058] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.
[0059] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.
[0060] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A network autonomy method for intelligent agents based on self-awareness, characterized in that, It includes: Pre-set and store constraint criteria in the intelligent agent; If the agent's own operating state and behavioral results deviate from the constraint criteria, then a first optimization requirement derived from the intention to self-correct is generated. If, after anticipating the evolution trend of the network environment in which the agent is currently located, a problem that needs to be addressed is identified, then a second optimization requirement arising from changes in the external environment is generated. The first optimization requirement and the second optimization requirement are comprehensively evaluated to obtain a unified executable goal that meets the constraint criteria; the unified executable goal is planned as a specific sequence of behaviors and executed to drive the network to autonomously evolve toward a state that meets the constraint criteria.
2. The network autonomy method for self-aware intelligent agents as described in claim 1, characterized in that, The constraint criteria include fixed constraint criteria and real-time constraint criteria; wherein, the real-time constraint criteria are criteria defined or modified based on the fixed constraint criteria during the operation of the agent.
3. The network autonomy method for self-aware intelligent agents as described in claim 2, characterized in that, If the agent's own operating state and behavioral results deviate from the constraints, a first optimization requirement derived from a self-correction intention is generated, which includes the following steps: The intelligent agent monitors its own operating state and behavioral results to identify whether it deviates from the fixed constraint criterion or the immediate constraint criterion. When a deviation is detected, a self-diagnosis is performed on the deviation to determine the type of deviation, and a correction requirement corresponding to the deviation type is generated as the first optimization strategy.
4. The network autonomy method for self-aware intelligent agents as described in claim 3, characterized in that, The first optimization strategy includes action correction requirements for revising specific execution actions, target adjustment requirements for adjusting the current task objective, and criterion update requirements for updating the immediate constraint criteria.
5. The network autonomy method for self-aware intelligent agents as described in claim 4, characterized in that, If the anticipated evolution trend of the network environment in which the agent is currently located presents a problem that needs to be addressed, a second optimization requirement arising from changes in the external environment is generated, which includes the following steps: The raw data of the network environment in which the agent is located is acquired and processed to form a structured representation of the current state of the network environment. Based on the structured representation and stored domain knowledge, scenario assessment is performed to construct a context describing the operational logic and relationships of the network environment. Based on the aforementioned context, a predictive model is used to estimate future state changes in the network environment. Identify the events that the agent needs to handle in the future state changes; generate the corresponding handling strategies for the events as the second optimization requirement.
6. The network autonomy method for self-aware intelligent agents as described in claim 5, characterized in that, A comprehensive evaluation of the first optimization requirement and the second optimization requirement is performed to form a unified executable objective that conforms to the constraints, which includes the following steps: The first optimization requirement and the second optimization requirement are respectively transformed into internal candidate targets and external candidate targets, and then merged into a candidate target set; The overall effect of each objective in the candidate objective set is evaluated by prediction model and simulation. Objectives that meet the fixed constraint criterion and the immediate constraint criterion are selected from the candidate objective set as the executable objectives.
7. The network autonomy method for self-aware intelligent agents as described in claim 6, characterized in that, The unified executable objective is planned as a specific sequence of behaviors and executed to drive the network to autonomously evolve toward a state that conforms to the constraints. This includes the following steps: For the unified executable objective, an action sequence to achieve the executable objective is generated by combining the fixed constraint criteria, the immediate constraint criteria, and the prediction model. The action sequence is executed, and the execution process and results are monitored in real time to form an autonomous closed loop.
8. The network autonomy method for self-aware intelligent agents as described in claim 1, characterized in that, It also includes establishing a knowledge base, in which the constraint criteria and knowledge set are centrally stored and updated; the knowledge set includes general knowledge, agent-specific knowledge, and normative value knowledge.
9. The network autonomy method for intelligent agents based on self-awareness as described in claim 8, characterized in that, It also includes building a knowledge engine; based on the knowledge engine, and combining the agent's observation of the network environment, the results of its behavior execution, and external interaction information, updating the knowledge set in the knowledge base; The knowledge engine iteratively optimizes the prediction models and decision-making strategies in the knowledge base based on the experience collected during the operation of the intelligent agent.
10. A network autonomous system for intelligent agents based on self-awareness, characterized in that, It includes: The self-awareness module is used to pre-set and store constraint criteria in the agent; If the agent's own operating state and behavioral results deviate from the constraints, then a first optimization requirement derived from the intention to self-correct is generated. The contextual awareness module is used to generate a second optimization requirement based on changes in the external environment if there are problems that need to be addressed in the expected evolution trend of the network environment in which the agent is currently located. The decision-making and action module is used to comprehensively evaluate the first optimization requirement and the second optimization requirement to form a unified executable goal that conforms to the constraint criteria; to plan the unified executable goal as a specific sequence of behaviors and execute it, driving the network to autonomously evolve toward a state that conforms to the constraint criteria.