Active defense system and method based on intelligent control heterogeneous network function

Through the intelligent control heterogeneous network functional system, the optimization of scheduling strategies is used to solve the problem of large resource overhead in heterogeneous redundant security architecture, and the efficient and flexible security defense of the mobile communication network is realized to adapt to complex dynamic environments.

CN120342667APending Publication Date: 2025-07-18INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510396095.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing heterogeneous redundant security architecture leads to excessive network resource overhead in mobile communication networks, increased system construction costs, and difficult to adapt to dynamic and complex network environments. Traditional control methods are easily identified and invaded by attackers.

Method used

The intelligent control heterogeneous network functional system is adopted, including the network functional heterogeneous executor set, enhanced service communication agent and heterogeneous executor perception and scheduling strategy learning center. Through deep reinforcement learning algorithms, the scheduling strategy is optimized to achieve dynamic selection and call network functional executors, reducing communication delays, and improving adaptability and flexibility.

Benefits of technology

Effectively reduce network resource overhead, improve communication efficiency, enhance system adaptability and flexibility, reduce single point failure risk, optimize resource utilization, and adapt to dynamic network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342667A_ABST
    Figure CN120342667A_ABST
Patent Text Reader

Abstract

The invention provides an active defense system and method based on an intelligent control heterogeneous network function, and the system comprises a network function heterogeneous executor set, an enhanced service communication agent, and a heterogeneous executor perception and scheduling strategy learning center. The enhanced service communication agent comprises a network function heterogeneous executor communication agent and a data collector; the network function heterogeneous executor communication agent is used for selecting a target network function heterogeneous executor for the signaling request; the data acquisition unit is used for acquiring network function heterogeneous executor state data and input and output sequence information; the heterogeneous executor sensing and scheduling strategy learning center comprises a sensing calculation module and a scheduling strategy optimization module; the sensing calculation module is used for determining comprehensive parameters of each network function heterogeneous executor, wherein the comprehensive parameters comprise a judgment result, a credibility result, an observation state and a reward function value; and the scheduling strategy optimization module is used for updating and issuing the latest probability scheduling strategy to the enhanced service communication agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer networks, and in particular, to an active defense system and method based on intelligent control heterogeneous network functions. Background Art

[0002] The classic heterogeneous redundancy architecture adopts a "single-input multiple-output mode", that is, each input is scheduled to multiple heterogeneous executors for synchronous processing, and the outputs of them are judged or voted.

[0003] However, the "single-input multiple-output mode" of the heterogeneous redundancy system increases the resource overhead by at least N times compared with the non-heterogeneous redundancy system, where N is greater than or equal to 3, representing the minimum number of heterogeneous executors that need to be synchronously executed to meet the security requirements. Inevitably, the system resource overhead is increased significantly, resulting in a significant increase in the system or network construction cost, especially for mobile communication networks with high business concurrency peaks and large fluctuations.

[0004] It can be seen that the security defense method in the related art has the technical problem of large network resource overhead. Summary of the Invention

[0005] The present invention provides an active defense system and method based on intelligent control heterogeneous network functions to solve the defect of excessive network resource overhead existing in the application of the heterogeneous redundancy security architecture in the prior art, and to improve the adaptability and flexibility of the network.

[0006] The present invention provides an active defense system based on intelligent control heterogeneous network functions, including: a set of network function heterogeneous executors, an enhanced service communication agent, and a heterogeneous executor perception and scheduling policy learning center; The enhanced service communication agent includes: a network function heterogeneous execution body communication agent and a data collector; the network function heterogeneous execution body communication agent is configured to, when the source network function of the received signaling request is a conventional network function and the destination network function of the signaling request is a network function heterogeneous execution body, select a target network function heterogeneous execution body from the network function heterogeneous execution body set for the signaling request according to a dynamically updated probability scheduling policy, so as to route and forward the signaling request; the data collector is configured to collect the status data of the network function heterogeneous execution body set and the network function heterogeneous execution body input-output sequence information of the network function heterogeneous execution body communication agent; and transmit the status data and the network function heterogeneous execution body input-output sequence information to the heterogeneous execution body perception and scheduling policy learning center; the heterogeneous execution body perception and scheduling policy learning center includes: a perception calculation module and a scheduling policy optimization module; the perception calculation module is configured to determine the comprehensive parameters of each network function heterogeneous execution body in the network function heterogeneous execution body set based on the status data and the network function heterogeneous execution body input-output sequence information, wherein the comprehensive parameters include: a judgment result, a trust degree result, an observation state, and a reward function value; the scheduling policy optimization module is configured to update the current probability scheduling policy based on the comprehensive parameters of each network function heterogeneous execution body, and send the latest probability scheduling policy to the enhanced service communication agent.

[0007] According to an active defense system based on an intelligent control heterogeneous network function provided by the present invention, the network function heterogeneous execution body communication agent is further configured to: when the source network function of the signaling request is a network function heterogeneous execution body and the destination network function of the signaling request is a conventional network function, forward the signaling request to the conventional service communication agent to which the destination network function belongs.

[0008] According to an active defense system based on an intelligent control heterogeneous network function provided by the present invention, the status data of the network function heterogeneous execution body set includes at least one of the following: the total number of signaling requests executed by each network function heterogeneous execution body in the network function heterogeneous execution body service, and the total number of signaling requests waiting in the network function heterogeneous execution body buffer; the average service time of the signaling requests of each network function heterogeneous execution body in the current time slot; the waiting duration of each signaling request of each network function heterogeneous execution body in the network function heterogeneous execution body set.

[0009] An active defense system based on intelligent control heterogeneous network functions provided by the present invention, the perception and calculation module includes: an asynchronous multi-party voting and decision-making sub-module, a trust degree calculation sub-module, a state observation sub-module, and a reward calculation sub-module; the asynchronous multi-party voting and decision-making sub-module is used to perform asynchronous multi-party voting and decision-making on the input and output sequence information of the network function heterogeneous executor for the current and historical, and obtain a decision result, wherein, the input and output sequence information of the network function heterogeneous executor includes: the service type to which the input belongs, the service type to which the output belongs, and the service operation; and send the decision result to the trust degree calculation sub-module; the trust degree calculation sub-module is used to calculate the trust degree based on the decision result, and obtain a trust degree result; and send the trust degree result to the state observation sub-module; the state observation sub-module is used to perform state perception based on the trust degree result, and obtain the observation state of each network function heterogeneous executor in the network function heterogeneous executor set; and send the observation state of each network function heterogeneous executor to the reward calculation sub-module; the reward calculation sub-module is used to perform function calculation based on the observation state of each network function heterogeneous executor, and obtain the reward function value of each network function heterogeneous executor.

[0010] An active defense system based on intelligent control heterogeneous network functions provided by the present invention, the scheduling strategy optimization module is based on a deep reinforcement learning algorithm.

[0011] An active defense system based on intelligent control heterogeneous network functions provided by the present invention, the deep reinforcement learning algorithm is an actor-critic deep reinforcement learning algorithm, and the scheduling strategy optimization module is further used to, when the decision result is abnormal or reaches a preset cycle time slot, based on the observation state of each network function heterogeneous executor and the reward function value of the observation state of each network function heterogeneous executor, perform learning optimization according to the actor-critic deep reinforcement learning algorithm.

[0012] The present invention also provides an active defense method based on intelligent control heterogeneous network functions, including the following steps: Through the network function heterogeneous execution body communication agent, when the source network function of the received signaling request is a conventional network function and the destination network function of the signaling request is a network function heterogeneous execution body, according to the dynamically updated probability scheduling strategy, select a target network function heterogeneous execution body in the network function heterogeneous execution body set for the signaling request to route and forward the signaling request; Through the data collector, collect the status data of the network function heterogeneous execution body set and the network function heterogeneous execution body input-output sequence information of the network function heterogeneous execution body communication agent; Through the perception calculation module, based on the status data and the network function heterogeneous execution body input-output sequence information, determine the comprehensive parameters of each network function heterogeneous execution body in the network function heterogeneous execution body set, where the comprehensive parameters include: judgment result, trust degree result, observation status, and reward function value; Through the scheduling strategy optimization module, update the current probability scheduling strategy based on the comprehensive parameters of each network function heterogeneous execution body and issue the latest probability scheduling strategy.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the active defense method based on intelligent control heterogeneous network functions as described in any one of the above.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the active defense method based on intelligent control heterogeneous network functions as described in any one of the above.

[0015] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the active defense method based on intelligent control heterogeneous network functions as described in any one of the above.

[0016] The active defense system and method based on the intelligent control heterogeneous network function provided by the present invention can, by introducing a set of heterogeneous network function executors, dynamically select and invoke different network function executors according to actual needs, thereby improving the adaptability and flexibility of the network; through the communication agent of the heterogeneous network function executors, it can, according to the dynamically updated probability scheduling strategy, select the most suitable heterogeneous network function executor for routing and forwarding of signaling requests, which helps to reduce communication latency and improve service response speed, thus enhancing the overall communication efficiency; the data collector is responsible for collecting the status data of the heterogeneous network function executors and the input-output sequence information of the communication agent, and this information provides a basis for subsequent perception calculation and scheduling strategy optimization; by the perception calculation module, the observation state of each heterogeneous network function executor is determined, and then the scheduling strategy optimization module updates the probability scheduling strategy based on these observation states. Thus, the strategy update mechanism based on real-time data and intelligent algorithms can ensure the intelligence and accuracy of the scheduling strategy, thereby optimizing the allocation and utilization of network resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art one by one. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic structural diagram of the active defense system based on the intelligent control heterogeneous network function provided by the present invention.

[0019] Figure 2 It is a schematic diagram of the ICH-NF active defense architecture (communication model C) provided by the present invention.

[0020] Figure 3 It is a schematic diagram of the ICH-NF active defense architecture (communication model D) provided by the present invention.

[0021] Figure 4 It is a schematic diagram of the overall working process and mechanism of the active defense system based on the intelligent control heterogeneous network function provided by the present invention.

[0022] Figure 5 It is a schematic diagram of the communication NFX scheduling and signaling interaction working process and mechanism (communication model C) provided by the present invention.

[0023] Figure 6 It is a schematic diagram of the communication NFX scheduling and signaling interaction working process and mechanism (communication model D) provided by the present invention.

[0024] Figure 7It is a schematic flow diagram of the NFX perception and scheduling policy update and optimization mechanism provided by the present invention.

[0025] Figure 8 It is a schematic flow diagram of the asynchronous multi-party voting decision mechanism provided by the present invention.

[0026] Figure 9 It is a schematic diagram of the voting decision method provided by the present invention.

[0027] Figure 10 It is a schematic diagram of the deep reinforcement learning network and scheduling policy optimization provided by the present invention.

[0028] Figure 11 It is a schematic diagram of the NFX scheduling policy update method provided by the present invention.

[0029] Figure 12 It is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed implementation manners

[0030] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] The classical heterogeneous redundant architecture adopts the "single input and multiple output mode", that is, each input is scheduled to multiple heterogeneous executors for synchronous processing, and their outputs are judged or voted. However, in the actual mobile communication network application, the following problems exist, and the specific analysis is as follows: Problem 1: The "single input and multiple output mode" of the heterogeneous redundant system increases the resource overhead by at least N times compared with the non-heterogeneous redundant system, where N is greater than or equal to 3, representing the minimum number of heterogeneous executors that need to be synchronously executed to meet the security requirements. Inevitably, the system resource overhead is greatly increased, resulting in a significant increase in the system or network construction cost, especially for the mobile communication network with high business concurrency peaks and large fluctuations.

[0032] Problem 2: The classical heterogeneous redundant defense architecture is usually based on modern control theory, relies on prior knowledge and accurate control object models, and the control modes and strategies are relatively fixed. The classical control method is difficult to adapt to dynamic, complex and rapidly evolving systems such as mobile communication networks.

[0033] Problem 3: As the running time of the core network functions increases, the probability that their vulnerabilities are identified or discovered by attackers or that the attackers successfully intrude will increase significantly, resulting in a significant increase in the probability that the traditional mode relying on the synchronous parallel output adjudication of multiple heterogeneous execution entities fails.

[0034] To address the above problems, for the service-oriented 5G / B5G mobile core network, the present invention proposes an active defense method and system based on Intelligent Control-based Heterogeneous Network Function (ICH-NF). The application scenarios and technical solutions will be introduced in detail below.

[0035] Refer to Figure 1 , Figure 1 is a schematic structural diagram of the active defense system based on intelligent control heterogeneous network functions provided by the present invention. As Figure 1 shown, the system includes the following: a set of network function heterogeneous execution entities 101, an enhanced service communication proxy 102, and a heterogeneous execution entity perception and scheduling policy learning center 103.

[0036] The enhanced service communication proxy 102 includes: a network function heterogeneous execution entity communication proxy 1021 and a data collector 1022.

[0037] The network function heterogeneous execution entity communication proxy 1021 is used to select a target network function heterogeneous execution entity in the set of network function heterogeneous execution entities for a signaling request according to the dynamically updated probability scheduling policy when the source network function of the received signaling request is a conventional network function and the destination network function of the signaling request is a network function heterogeneous execution entity, so as to route and forward the signaling request; The data collector 1022 is used to collect the status data of the set of network function heterogeneous execution entities and the input / output sequence information of the network function heterogeneous execution entity of the network function heterogeneous execution entity communication proxy; and transmit the status data and the input / output sequence information of the network function heterogeneous execution entity to the heterogeneous execution entity perception and scheduling policy learning center 103; The heterogeneous execution entity perception and scheduling policy learning center 103 includes: a perception calculation module 1031 and a scheduling policy optimization module 1032; The perception calculation module 1031 is used to determine the comprehensive parameters of each network function heterogeneous execution entity in the set of network function heterogeneous execution entities based on the status data and the input / output sequence information of the network function heterogeneous execution entity, where the comprehensive parameters include: a judgment result, a trust degree result, an observation state, and a reward function value; The scheduling policy optimization module 1032 is used to update the current probability scheduling policy based on the comprehensive parameters of each network function heterogeneous execution entity and send the latest probability scheduling policy to the enhanced service communication proxy 102.

[0038] Reference Figure 2 , Figure 2 is a schematic diagram of the ICH-NF active defense architecture (communication model C) provided by the present invention.

[0039] Reference Figure 3 , Figure 3 is a schematic diagram of the ICH-NF active defense architecture (communication model D) provided by the present invention.

[0040] In an embodiment of the present invention, based on the compatibility with the 3GPP R16 enhanced service-based architecture (eSBA) communication models C and D, an ICH-NF active defense architecture is designed, as shown in Figure 2 , Figure 3 shown below. Its architecture and core components will be described in detail below.

[0041] As shown in Figure 2 , Figure 3 shown, the ICH-NF active defense architecture includes three parts: a set of network function (NF) heterogeneous execution entities (NFX), an enhanced service communication proxy (eSCP), and a heterogeneous execution entity awareness and scheduling policy learning center. The ICH-NF active defense architecture accesses the conventional 3GPP (3rd Generation Partnership Project) core network through the standard interface between service communication proxies (SCPs), that is, the enhanced service communication proxy (eSCP) and the conventional SCP are interconnected through the standard interface between SCPs defined by 3GGP (see the F4 interface in Figure 2 ). The routing discovery and forwarding of the conventional NF in the 3GPP core network are responsible for its home SCP (see the F2 interface in Figure 2 ). Among them, the set of conventional NFs includes conventional, non-heterogeneous redundant NFs, for example, session management function (SMF), policy control function (PCF), user plane function (UPF), authentication service function (AUSF), etc. The conventional NFs can be deployed uniquely or in a homogeneous redundant manner to improve the reliability of the core network.

[0042] The functions of the core modules in the active defense system based on intelligent control heterogeneous network functions will be described in detail below.

[0043] In an embodiment of the present invention, a set of network function heterogeneous executors (NFX) is used to represent a slice for a certain key NF (such as AMF, UDM, etc.) of the mobile core network or a slice composed of multiple key NFs.

[0044] The NF heterogeneous executor is constructed by using heterogeneous software and hardware implementation methods. For example, the NF software uses different open-source softwares such as Aether, free5GC, Open5GS, etc., as well as different commercial core network softwares; at the same time, the hardware and system software use different manufacturer physical servers and virtualization platform softwares.

[0045] In order to achieve the flexible scheduling of service requests and responses by the eSCP and the data collection of the set of network function heterogeneous executors, in the ICH-NF core network active defense architecture, the context data and log data of the network function heterogeneous executors are uniformly stored and shared through the internal interface iF1. According to the network function heterogeneous executor scheduling decision, the service can be seamlessly taken over by the network function heterogeneous executors within the same set of network function heterogeneous executors.

[0046] In an embodiment of the present invention, the enhanced Service Communication Proxy (eSCP) consists of a network function heterogeneous executor communication proxy and a data collector.

[0047] The functions of the network function heterogeneous executor communication proxy include: routing discovery and forwarding of signaling requests, and online scheduling decision-making for network function heterogeneous executors.

[0048] According to an active defense system based on intelligent control heterogeneous network functions provided by the present invention, the network function heterogeneous executor communication proxy is also used for: When the source network function of the signaling request is a network function heterogeneous executor and the destination network function of the signaling request is a conventional network function, forwarding the signaling request to the conventional service communication proxy to which the destination network function belongs.

[0049] In an embodiment of the present invention, when the source network function of the signaling request is a network function heterogeneous executor and the destination network function is a conventional network function, the network function heterogeneous executor communication proxy forwards it to the conventional service communication proxy to which the destination network belongs (through the F4 interface), and this interface follows the SCP interface specification given in 3GPP 23.501. When the source network function requested by the signaling is a conventional network function and the destination network function is a network function heterogeneous executor, the Network Function Heterogeneous Executor Communication Proxy (NFX SCP) selects a network function heterogeneous executor for the arriving signaling request according to the latest probability scheduling policy, and forwards the signaling request to the selected target network function heterogeneous executor (through the F1 interface); wherein, the probability scheduling policy (i.e., each network function heterogeneous executor is associated with a probability value) is sent by the (Network Function) Heterogeneous Executor Sensing and Scheduling Policy Learning Center to the Network Function Heterogeneous Executor Communication Proxy.

[0050] Through the embodiments of the present invention, the computational overhead is extremely low, and real-time scheduling decisions can be efficiently implemented when each signaling request arrives. Especially in the case of a sharp increase in the signaling request arrival rate (a sharp increase in normal services or a denial-of-service attack), the Network Function Heterogeneous Executor (Service) Communication Proxy can still work well without falling into congestion collapse, and can effectively solve the single point of failure problem.

[0051] In the embodiments of the present invention, the data collector is responsible for collecting the status data of the network function heterogeneous executors through the F3 interface, and sending the collected data to the status observation sub-module in the sensing and computing module in the heterogeneous executor sensing and scheduling policy learning center (through the F5 interface).

[0052] According to an active defense system based on an intelligent control heterogeneous network function provided by the present invention, the status data of the network function heterogeneous executor set includes at least one of the following: For each network function heterogeneous executor in the network function heterogeneous executor set, the total number of signaling requests executed in the network function heterogeneous executor service, and the total number of signaling requests waiting in the network function heterogeneous executor buffer; For each network function heterogeneous executor in the network function heterogeneous executor set, the average service time of the signaling requests in the current time slot; The waiting duration of each signaling request of each network function heterogeneous executor in the network function heterogeneous executor set.

[0053] In the embodiments of the present invention, the status data of any network function heterogeneous executor collected by the data collector includes: the total number of signaling requests currently in the network function heterogeneous executor service and waiting in the network function heterogeneous executor buffer; the average service time of the signaling requests in the current time slot; the waiting duration of each request in the network function heterogeneous executor (i.e., the time elapsed since the signaling request arrived at the network function heterogeneous executor).

[0054] In some embodiments, the current time slot average service time metric is used to quantify the real-time computing efficiency of each executor; the waiting duration of each signaling request is collected to achieve quality of service control. For example, a heat map of the waiting duration distribution is constructed to identify queuing anomalies for specific types of requests (such as VoIP signaling).

[0055] The data collector is responsible for collecting the input-output sequence information of the network function heterogeneous executors (services) from the communication agents of the network function heterogeneous executors (through the iF2 interface) and reporting it to the asynchronous multi-party voting decision sub-module in the perception computing module (through the F6 interface).

[0056] Through the embodiments of the present invention, by collecting the total number of current service requests and the number of waiting requests in the buffer of each executor, the system can perceive the load pressure distribution of each executor in real time. Combining with the reinforcement learning algorithm of the scheduling strategy learning center, the routing weight allocation can be dynamically adjusted.

[0057] In the embodiments of the present invention, heterogeneous executor perception and scheduling strategy learning include a perception computing module and a scheduling strategy optimization module.

[0058] According to an active defense system based on intelligent control heterogeneous network functions provided by the present invention, the perception computing module includes: an asynchronous multi-party voting decision sub-module, a trustworthiness calculation sub-module, a state observation sub-module, and a reward calculation sub-module; The asynchronous multi-party voting decision sub-module is used to perform asynchronous multi-party voting decision based on the current and historical input-output sequence information of the network function heterogeneous executors to obtain a decision result, where the input-output sequence information of the network function heterogeneous executors includes: the service type to which the input belongs, the service type to which the output belongs, and the service operation; and send the decision result to the trustworthiness calculation sub-module; The trustworthiness calculation sub-module is used to calculate the trustworthiness based on the decision result to obtain a trustworthiness result; and send the trustworthiness result to the state observation sub-module; The state observation sub-module is used to perform state perception based on the trustworthiness result to obtain the observation state of each network function heterogeneous executor in the network function heterogeneous executor set; and send the observation state of each network function heterogeneous executor to the reward calculation sub-module; The reward calculation sub-module is used to perform function calculation based on the observation state of each network function heterogeneous executor to obtain the reward function value of each network function heterogeneous executor.

[0059] In the embodiments of the present invention, the asynchronous multi-party voting decision sub-module reports the decision result to the NFX scheduling strategy optimization module based on deep reinforcement learning (DRL) (through the F8 interface), and sends the trust metric related parameters to the trustworthiness calculation sub-module (through the iF3 interface).

[0060] The trustworthiness calculation sub-module is responsible for sending the calculated trustworthiness to the reward calculation sub-module (through the iF4 interface).

[0061] The status observation sub-module is responsible for reporting the perceived status information to the reward calculation sub-module (through the iF5 interface).

[0062] The reward calculation sub-module comprehensively calculates the reward function value based on the trustworthiness and the perceived status information, and reports it to the DRL-based NFX scheduling policy optimization module (through the F7 interface).

[0063] Through the embodiments of the present invention, through multi-party voting judgment, single-point decision-making deviation is avoided, and the reliability and fairness of the judgment result are improved; based on the multi-party judgment result and real-time behavior data, the trustworthiness of NFX is dynamically calculated to reflect its reliability and security; the reward function comprehensively considers the trustworthiness and the system state to achieve balanced optimization of security, performance, and resource utilization rate.

[0064] According to an active defense system based on intelligent control heterogeneous network functions provided by the present invention, the scheduling policy optimization module is based on a deep reinforcement learning algorithm.

[0065] In the embodiments of the present invention, the deep reinforcement learning algorithm is specifically the Actor-Critic deep reinforcement learning algorithm. Actor: Responsible for generating a scheduling policy, that is, selecting a target network function heterogeneous execution body (NFX) according to the current state. Critic: Evaluates the policy generated by the Actor, quantifies the advantages and disadvantages of the current state and action through a value function, and provides an optimization direction for the Actor. Collaboration mechanism: The Actor and the Critic learn through interaction and continuously optimize the policy. The Actor adjusts the policy according to the feedback of the Critic, and the Critic updates the value function based on the action of the Actor to form a closed-loop optimization.

[0066] Through the embodiments of the present invention, the Actor-Critic algorithm supports online learning, can update the scheduling policy in real time, and adapt to the dynamically changing network environment and threat situation.

[0067] According to an active defense system based on intelligent control heterogeneous network functions provided by the present invention, the deep reinforcement learning algorithm is the Actor-Critic deep reinforcement learning algorithm, and the scheduling policy optimization module is further configured to, when the judgment result is abnormal or reaches a preset periodic time slot, perform learning optimization according to the Actor-Critic deep reinforcement learning algorithm based on the observed state of each network function heterogeneous execution body and the reward function value of the observed state of each network function heterogeneous execution body.

[0068] In the embodiment of the present invention, the scheduling policy optimization module adopts the Actor-Critic deep reinforcement learning method, updates the NFX scheduling probability policy based on the NFX observation state reported by the perception computing module, and distributes the latest scheduling policy to the network function heterogeneous execution entity communication agent (through the F9 interface); at the same time, it is responsible for performing Actor-Critic network learning optimization based on the historical NFX observation state and reward function value reported by the perception computing module, triggered by the abnormal judgment result or periodically triggered.

[0069] Reference Figure 4 , Figure 4 is a schematic diagram of the overall working process and mechanism of the active defense system based on intelligent control heterogeneous network functions provided by the present invention, including: a scheduling policy optimization module, a perception computing module, a data collector, a network function heterogeneous execution entity communication agent (NFX SCP), and a network function heterogeneous execution entity set (NFX set).

[0070] As Figure 4 shown, from the time dimension, the triggering execution method of the present invention is as follows: The network function heterogeneous execution entity scheduling decision of the network function heterogeneous execution entity communication agent is triggered by the received signaling request, and scheduling decisions and network function heterogeneous execution entity routing are performed in the order of signaling request arrival; In order to adapt to the changes in the signaling load and security state of the network function heterogeneous execution entity and effectively reduce system overhead, the scheduling policy update of the network function heterogeneous execution entity communication agent is triggered periodically in units of time slots, that is, let any time slot be denoted as t∈{0,1,2,...}, then before the end of each time slot t, the scheduling policy optimization module updates the scheduling policy and distributes it to the scheduling policy of the network function heterogeneous execution entity communication agent, while the network function heterogeneous execution entity communication agent keeps the scheduling policy it uses unchanged within each time slot.

[0071] The triggering of the deep reinforcement learning network and policy set optimization is divided into two cases: if the asynchronous multi-party voting judgment mechanism does not detect abnormal behavior, it is triggered and executed once every T time slots (where T is a positive integer), otherwise, it is triggered and executed in the current time slot, and the time slot is re-counted, with the current time slot as the first time slot of the cycle.

[0072] The following separately describes the NFX scheduling and signaling interaction mechanism, as well as the NFX perception and scheduling policy update and optimization mechanism, as Figure 4 shown.

[0073] Reference Figure 5 , Figure 5 is a schematic diagram of the communication NFX scheduling and signaling interaction working process and mechanism (communication model C) provided by the present invention.

[0074] Reference Figure 6 , Figure 6 is a schematic diagram of the communication NFX scheduling and signaling interaction workflow and mechanism (communication model D) provided by the present invention.

[0075] Based on the compatibility with the 3GPP R16 standard, the NFX scheduling and signaling interaction mechanism of ICH-NF active defense is as Figure 5 , Figure 6 shown.

[0076] Step 1, the conventional NF sends a signaling request message to the home conventional SCP. If service discovery is required, the signaling request message additionally includes service discovery query parameters, which include the required source NF type, target NF type, service name, and key information for matching the target NF.

[0077] Step 2, the conventional SCP performs a target NF service discovery process to the home NRF.

[0078] Step 3, the Network Repository Function (NRF) returns the NFX SCP information to which the target NF belongs to the conventional SCP.

[0079] Step 4, the conventional SCP forwards the signaling request message to the NFX SCP.

[0080] Step 5, the NFX SCP uses a probabilistic scheduling method to select an NFX for the arriving signaling request.

[0081] Specifically: the scheduling policy of the NFX SCP is updated before the end of each time slot t, and the scheduling policy used within time slot t + 1 remains unchanged. The scheduling policy is denoted as , where is the total number of NFXs; is the scheduling probability distribution of any NFX in the current time slot . The NFX SCP schedules and selects an NFX for the arriving signaling request according to in sequence.

[0082] Step 6, the NFX SCP forwards the signaling request message to the selected NFX.

[0083] Specifically: after the signaling request message is transmitted to each NFX, it is served one by one according to the principle of first come, first served; without loss of generality, the queues in the NFXs for caching and processing the signaling request messages have the same buffer capacity, that is, each NFX queue has the same maximum queue length and can accommodate at most C signaling request messages; when the cache queue of an NFX is full, any request scheduled to that NFX will be discarded.

[0084] Step 7: The target NFX sends a signaling response message to the NFX SCP.

[0085] Step 8: The NFX SCP forwards the signaling response message to the corresponding regular SCP, specifically following the 3GPP 23.501 specification.

[0086] Step 9: Finally, the regular SCP forwards this information to the regular NF, specifically following the 3GPP 23.501 specification.

[0087] For subsequent signaling interactions between the regular NF and the NFX, except for not performing Service Discovery Steps 2 and 3, other steps are completed in the same way.

[0088] Reference Figure 7 , Figure 7 is a schematic flowchart of the NFX perception and scheduling policy update and optimization mechanism provided by the present invention.

[0089] The NFX perception and scheduling policy update and optimization mechanism of the active defense system based on the intelligent control heterogeneous network function is as Figure 7 shown.

[0090] Step 1: NFX data collection and upload: The data collector obtains relevant information from the NFX set and the NFX SCP.

[0091] Specifically, for any NFX , the status data of the current time slot t directly collected and reported by the data collector from the NFX includes: The total number of signaling requests currently in the NFX service and waiting in the NFX buffer (denoted as ); The average service time of the signaling requests in the current time slot (denoted as ); The waiting duration of the th largest request in the NFX waiting duration (i.e., the time elapsed since the signaling request arrived at the NFX) (denoted as ); The data collector concurrently obtains signaling request-response log data from the NFX SCP (i.e., the NFX input-output sequence). Specifically: The data collector extracts the service type and service operation to which the input and output belong from the HTTP start line string in each NFX input-output sequence, and divides the sequences of the input and output with the same service operation into a group, denoted as: , where represents the ( ) service operation of the NFX in time slot t m 's An input-output sequence, For NFX The sequence set corresponding to the service operation m.

[0092] After collecting the data, the data collector forwards the relevant data to the perception computing module to complete the upload of NFX data.

[0093] Step 2, Asynchronous multi-party voting decision: After the perception computing module groups and aligns the input-output sequences according to the service operation type, it performs asynchronous multi-party voting decision, and the result is recorded as .

[0094] Step 3, NFX trustworthiness and reward parameter calculation: The perception computing module immediately calculates the NFX trustworthiness and reward parameters.

[0095] Step 4, Reporting of perception computing results: The perception computing module reports the decision result, the NFX observation state vector, and the reward parameter value to the scheduling policy optimization module.

[0096] Specifically: According to the calculation result of Step 3, the perception computing module reports the reward , the state observation to the scheduling policy optimization module; the scheduling policy optimization module then records the reward , the observation set , the action and other time slot trajectory information to provide a data basis for the optimization of the deep reinforcement learning network and the scheduling policy set.

[0097] Step 5, DRL neural network optimization condition judgment: The system makes a condition judgment based on whether there is an abnormal situation and whether the predetermined period T is reached. If neither an abnormality is detected nor the period T is reached, the system directly enters Step 7 for scheduling decision-making and update. If there is an abnormality or the period T has been reached (i.e., either of the two conditions is "yes"), the system enters Step 6 for optimization processing through the deep neural network of reinforcement learning.

[0098] Step 6, The scheduling policy optimization module optimizes the deep neural network of reinforcement learning.

[0099] Step 7, The scheduling policy optimization module selects the NFX scheduling probability policy through the executor deep neural network; and records the observation state vector and the corresponding reward parameter value.

[0100] Step 8, The scheduling policy optimization module generates an update of the scheduling probability policy and sends it to the NFX SCP.

[0101] Reference Figure 8 , Figure 8It is a schematic flowchart of the asynchronous multi-party voting and decision-making mechanism provided by the present invention.

[0102] The asynchronous multi-party voting and decision-making work process includes three steps: sequence grouping and padding based on service operation types, feature extraction and hashing, and voting and decision-making.

[0103] In the embodiment of the present invention, for the convenience of explaining the asynchronous multi-party voting and decision-making mechanism, let: within time slot t, any input-output sequence of any NFX ( ) is denoted as , where , is the set of all sequences of NFX within time slot t; Before time slot t, the input-output sequence of the NFX service operation m that has been voted and judged to be normal as required is denoted as .

[0104] The services and service operations between network functions are implemented by the HTTP / 2 protocol. That is, the service type and service operation information are transmitted in the form of URI through methods such as Get and Post in the HTTP start line characters. The general work process of NF service in 5GC is as follows: the service-consuming NF and the service-providing NF, the consuming NF sends an application -> the providing NF receives the request -> the providing NF generates a response and sends it (possibly an abnormal response code) -> the consuming NF receives the response / obtains the service.

[0105] Therefore, the present invention extracts the service types and service operations to which the input and output belong from the HTTP start line strings in the input-output sequences of each NFX, and divides the sequences of the input and output with the same service operation into a group, denoted as , indicating the ( )-th input-output sequence of the service operation m of NFX within time slot t, is the sequence set corresponding to the service operation m of NFX .

[0106] In the embodiment of the present invention, according to the situation of the NFX sequences in each group, group padding is performed. The specific padding method is as follows: If the number of NFX sources of the sequences in the group is greater than or equal to the set threshold ∂, no data padding is required. For example, the group of service operation 1 contains . When the threshold ∂ = 3, since the sequence sources in the group contain more than 3 different NFXs, no data padding is required.

[0107] If the number of NFX sources of the sequences in a group (denoted as ∂1) is less than the set threshold ∂, then ∂ - ∂1 historical sequence information from different NFXs that have been judged to be normal is added to complete it. For example, when the threshold ∂ = 3, the group of service operation 1 only contains , then can be used to complete it.

[0108] In the embodiment of the present invention, the extraction of group features and hash calculation include: Based on a pre-constructed group feature template (where represents the number of common features of group m, and the method for constructing the feature template will be described in detail later), the key fields of each sequence in the current time slot are extracted from each group as features. The features extracted from any group m can be denoted as , where . In addition, for the convenience of voting judgment, hash calculation processing is performed on each feature field of the sequence.

[0109] The feature template of each group is calculated using Algorithm 1. Specifically: First, collect the normal input-output sequence sets of each NFX offline and group them based on service operations, denoted as , where .

[0110] Then, based on the XOR operation of the corresponding key parameters of all sequences in any group , the common feature keyword field template is screened out and denoted as , represents the number of common features of group m.

[0111] In the embodiment of the present invention, the voting judgment includes: Set up a Boolean judgment sequence vector , where represents the abnormal judgment result of the NFX of group 1 at time slot t, where (where 0 represents normal and 1 represents abnormal). Based on for abnormal voting judgment, and the feature abnormal conditions of each group sequence are integrated to generate a corresponding judgment sequence vector. The specific method is as follows: Self-comparison and preliminary determination: For all sequences of any given NFX under the same group m, a distance-based algorithm is adopted for for feature anomaly detection. If no anomaly is found during the self-comparison process, any one Perform comprehensive voting comparison; if an abnormality is found during the self-comparison process, directly determine that the NFX is abnormal, do not perform comprehensive voting comparison, and output the abnormal result.

[0112] Comprehensive voting comparison: For the extracted features of all NFXs in the same group m, perform an exclusive OR algorithm for abnormal voting detection, and determine the minority type as abnormal and output the result.

[0113] Reference Figure 9 , Figure 9 is a schematic diagram of the voting decision method provided by the present invention.

[0114] Based on the above method, the determination results of various service operation groups of all NFXs can be obtained, and the comprehensive determination result of time slot t can be generated. The voting decision method is as Figure 9 shown.

[0115] To ensure the signaling request service quality and security of the scheduling strategy, for any time slot , define the reward function as: , where is the number of scheduling decisions that do not meet the signaling request service delay requirement within time slot , which is directly collected and reported by the data collector from the NFX set; is the total number of scheduling decisions within time slot ; is the predefined weight; is the sum of the trust values of all NFXs selected by the scheduling within time slot , defined as , where is the number of NFXs selected by the scheduling within time slot , is the trust value of NFX within time slot .

[0116] Based on the asynchronous voting decision result reported by the perception computing module, the trust value of NFX is defined as . Therefore, there is: Reference Figure 10 , Figure 10 is a schematic diagram of the deep reinforcement learning network and scheduling strategy optimization provided by the present invention.

[0117] If the state of the NFX can be directly perceived and is completely observable, that is, for all , , the decision problem of the above periodic adjustment signaling request scheduling probability distribution is a Markov Decision Process (MDP). Let , be the observation set, and rewrite the representation forms of the transition kernel function and the reward kernel function as and .

[0118] To adapt to the partial observability of the actual system, and even more severe situations, a parameterized memoryless policy is adopted, where the parameter , is a compact Euclidean parameter set. For the memoryless policy , the action value function for all , can be expressed as: , where is the conditional access distribution of the hidden MAP phase given the observation value. Further introduce the observation value function of the policy , , Furthermore, the advantage function of the policy can be expressed as: . As an approximation of the intractable surrogate objective, the following optimization objective is used to train the actor network: , where is an estimate of the unknown advantage , and is estimated using the Generalized Advantage Estimation (GAE) method as: , where is a predefined parameter used to adjust the trade-off between the bias and variance of the estimate; is a function of the parameter, used to learn an appropriate observation value function (unknown in the actual system). By minimizing the mean squared error loss, this approximation of the observation value function (which can be called the critic) can also be updated periodically.

[0119] A deterministic function is defined by a neural network , given the observations and the policy parameters , first is fed into the neural network parameterized by to compute ; then, the action is sampled by generating each , resulting in: , where (where ) are independent samples drawn from , where , denotes the rd element of

[0120] Based on the above analysis, during the pre-training stage before the first use of the scheduling policy optimization module, offline learning of the scheduling policy is performed using Algorithm 2 through scheduling requests, selections, and observed data. The scheduling policy optimization module also uses Algorithm 2 to update the deep learning network periodically or event-triggered during use (operation).

[0121] Refer to Figure 11 , Figure 11 which is a schematic diagram of the NFX scheduling policy update method provided by the present invention.

[0122] Define the set containing all possible NFX states as , where ; the sensing and computing module senses an observation of all current NFX states (denoted as ) reported by the data collector, where is the current phase of MAP; is the current state of NFX ; is the set of all system states. For , is the set of all possible NFX states when the queue length is , and any element in has the form , where is a non-negative real number satisfying the relationship.

[0123] As Figure 11 shows, in any time slot Before the end, the formal description of the NFX scheduling policy update method is as follows: The scheduling policy optimization module takes the observation as a condition, and through the Actor network, follows the current policy to select an action as the probability distribution for scheduling NFX in this time slot and sends it to the NFX SCP module in the eSCP. Among them, is a d-dimensional probability simplex, indicating the probability of selecting different NFXs for the scheduling request within the time slot ; ( ) is a probability density function conditional on the observation and on the behavior set .

[0124] Through the embodiments of the present invention, an intelligent control heterogeneous network function (ICH-NF) active defense architecture (hereinafter referred to as the "ICH-NF active defense architecture") and key technologies are proposed. The specific advantages include: First, on the basis of being compatible with the 3GPP R16 enhanced service-based architecture (eSBA), the ICH-NF active defense architecture is proposed. This architecture supports the NFX asynchronous multi-party sequence voting decision mechanism and realizes the online one-by-one scheduling control and policy optimization neural network learning of the trustworthiness and quality of service of the adaptive network function heterogeneous executor (NFX) based on deep reinforcement learning.

[0125] Second, in order to decouple the dependence of traditional voting decisions on the synchronous scheduling execution of multiple NFXs and solve the resource efficiency problem of multi-redundant NFX synchronous services, an NFX asynchronous multi-party sequence voting decision framework is proposed, which includes methods for NFX sequence grouping, complementing, feature extraction, and joint decision-making based on service operation behaviors, and realizes the voting decision of the input-output sequences of joint multi-NFX asynchronous homogeneous service operations.

[0126] Third, in order to achieve intelligent scheduling control of NFX, a reward function that takes into account the trustworthiness and quality of service of NFX is designed, and an NFX scheduling decision and policy update method based on deep reinforcement learning, as well as a deep reinforcement learning network optimization method, are proposed.

[0127] Next, the active defense method based on the intelligent control heterogeneous network function provided by the present invention will be described. The active defense method based on the intelligent control heterogeneous network function described below can be mutually referred to with the active defense method based on the intelligent control heterogeneous network function described above.

[0128] Through the network function heterogeneous execution body communication agent, when the source network function of the received signaling request is a conventional network function and the destination network function of the signaling request is a network function heterogeneous execution body, according to the dynamically updated probability scheduling policy, select a target network function heterogeneous execution body from the network function heterogeneous execution body set for the signaling request to route and forward the signaling request; Through the data collector, collect the status data of the network function heterogeneous execution body set and the network function heterogeneous execution body input / output sequence information of the network function heterogeneous execution body communication agent; Through the perception and calculation module, based on the status data and the network function heterogeneous execution body input / output sequence information, determine the comprehensive parameters of each network function heterogeneous execution body in the network function heterogeneous execution body set, where the comprehensive parameters include: decision result, trust degree result, observation status, and reward function value; Through the scheduling policy optimization module, update the current probability scheduling policy based on the comprehensive parameters of each network function heterogeneous execution body and issue the latest probability scheduling policy.

[0129] Specifically, the active defense method based on the intelligent control heterogeneous network function provided by the present invention can implement all the method steps implemented by the active defense method embodiment based on the intelligent control heterogeneous network function and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiment will not be specifically described herein.

[0130] Figure 12 It is a schematic diagram of the physical structure of the electronic device provided by the present invention, as Figure 12As shown in the figure, the electronic device may include: a processor 1210, a communications interface 1220, a memory 1230, and a communication bus 1240. Among them, the processor 1210, the communications interface 1220, and the memory 1230 complete communication with each other through the communication bus 1240. The processor 1210 may call the logical instructions in the memory 1230 to execute an active defense method based on the intelligent control heterogeneous network function. The method includes: through the network function heterogeneous execution body communication agent, when the source network function of the received signaling request is a conventional network function and the destination network function of the signaling request is a network function heterogeneous execution body, according to the dynamically updated probability scheduling strategy, select a target network function heterogeneous execution body from the network function heterogeneous execution body set for the signaling request to route and forward the signaling request; through the data collector, collect the status data of the network function heterogeneous execution body set and the network function heterogeneous execution body input-output sequence information of the network function heterogeneous execution body communication agent; through the perception calculation module, based on the status data and the network function heterogeneous execution body input-output sequence information, determine the comprehensive parameters of each network function heterogeneous execution body in the network function heterogeneous execution body set, where the comprehensive parameters include: judgment result, trust degree result, observation status, and reward function value; through the scheduling strategy optimization module, update the current probability scheduling strategy based on the comprehensive parameters of each network function heterogeneous execution body and send down the latest probability scheduling strategy.

[0131] In addition, when the logical instructions in the above-mentioned memory 1230 are implemented in the form of software function units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0132] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the active defense method based on the intelligent control heterogeneous network function provided by the above-mentioned various methods. The method includes: through the network function heterogeneous execution body communication agent, when the source network function of the received signaling request is a conventional network function and the destination network function of the signaling request is a network function heterogeneous execution body, according to the dynamically updated probability scheduling strategy, select a target network function heterogeneous execution body in the network function heterogeneous execution body set for the signaling request, so as to route and forward the signaling request; through the data collector, collect the status data of the network function heterogeneous execution body set and the network function heterogeneous execution body input / output sequence information of the network function heterogeneous execution body communication agent; through the perception calculation module, based on the status data and the network function heterogeneous execution body input / output sequence information, determine the comprehensive parameters of each network function heterogeneous execution body in the network function heterogeneous execution body set, wherein the comprehensive parameters include: judgment result, trust degree result, observation status and reward function value; through the scheduling strategy optimization module, update the current probability scheduling strategy based on the comprehensive parameters of each network function heterogeneous execution body, and issue the latest probability scheduling strategy.

[0133] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the active defense method based on the intelligent control heterogeneous network function provided by the above-mentioned various methods. The method includes: through the network function heterogeneous execution body communication agent, when the source network function of the received signaling request is a conventional network function and the destination network function of the signaling request is a network function heterogeneous execution body, according to the dynamically updated probability scheduling strategy, select a target network function heterogeneous execution body in the network function heterogeneous execution body set for the signaling request, so as to route and forward the signaling request; through the data collector, collect the status data of the network function heterogeneous execution body set and the network function heterogeneous execution body input / output sequence information of the network function heterogeneous execution body communication agent; through the perception calculation module, based on the status data and the network function heterogeneous execution body input / output sequence information, determine the comprehensive parameters of each network function heterogeneous execution body in the network function heterogeneous execution body set, wherein the comprehensive parameters include: judgment result, trust degree result, observation status and reward function value; through the scheduling strategy optimization module, update the current probability scheduling strategy based on the comprehensive parameters of each network function heterogeneous execution body, and issue the latest probability scheduling strategy.

[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An active defense system based on the intelligent control heterogeneous network function, characterized in that, Including: A set of network function heterogeneous executors, an enhanced service communication agent, and a heterogeneous executor perception and scheduling policy learning center; The enhanced service communication agent includes: a network function heterogeneous executor communication agent and a data collector; The network function heterogeneous executor communication agent is used to, when the source network function of the received signaling request is a conventional network function and the destination network function of the signaling request is a network function heterogeneous executor, select a target network function heterogeneous executor in the set of network function heterogeneous executors for the signaling request according to the dynamically updated probability scheduling policy, so as to route and forward the signaling request; The data collector is used to collect the status data of the set of network function heterogeneous executors and the network function heterogeneous executor input-output sequence information of the network function heterogeneous executor communication agent; and transmit the status data and the network function heterogeneous executor input-output sequence information to the heterogeneous executor perception and scheduling policy learning center; The heterogeneous executor perception and scheduling policy learning center includes: a perception computing module and a scheduling policy optimization module; The perception computing module is used to determine the comprehensive parameters of each network function heterogeneous executor in the set of network function heterogeneous executors based on the status data and the network function heterogeneous executor input-output sequence information, where the comprehensive parameters include: a judgment result, a trust degree result, an observation status, and a reward function value; The scheduling policy optimization module is used to update the current probability scheduling policy based on the comprehensive parameters of each network function heterogeneous executor, and send the latest probability scheduling policy to the enhanced service communication agent.

2. The active defense system based on the intelligent control heterogeneous network function according to claim 1, wherein The network function heterogeneous executor communication agent is further used for: When the source network function of the signaling request is a network function heterogeneous executor and the destination network function of the signaling request is a conventional network function, forwarding the signaling request to the conventional service communication agent to which the destination network function belongs.

3. The active defense system based on the intelligent control heterogeneous network function according to claim 1, characterized in that, The status data of the set of network function heterogeneous executors includes at least one of the following: The total number of signaling requests executed by each network function heterogeneous executor in the set of network function heterogeneous executors in the network function heterogeneous executor service, and the total number of signaling requests waiting in the network function heterogeneous executor buffer; The average service time of the signaling requests of each network function heterogeneous executor in the set of network function heterogeneous executors in the current time slot; The waiting duration of each signaling request of each network function heterogeneous executor in the set of network function heterogeneous executors.

4. The active defense system based on the intelligent control heterogeneous network function according to claim 1, wherein The perception computing module includes: an asynchronous multi-party voting judgment sub-module, a trust degree calculation sub-module, a status observation sub-module, and a reward calculation sub-module; The asynchronous multi-party voting judgment sub-module is used to perform asynchronous multi-party voting judgment based on the current and historical network function heterogeneous executor input-output sequence information to obtain a judgment result, where the network function heterogeneous executor input-output sequence information includes: the service type to which the input belongs, the service type to which the output belongs, and the service operation; and send the judgment result to the trust degree calculation sub-module; The trust degree calculation sub-module is used to calculate the trust degree based on the judgment result to obtain a trust degree result, and send the trust degree result to the state observation sub-module; The state observation sub-module is used to perform state perception based on the trust degree result to obtain the observation state of each network function heterogeneous executor in the network function heterogeneous executor set, and send the observation state of each network function heterogeneous executor to the reward calculation sub-module; The reward calculation sub-module is used to perform function calculation based on the observation state of each network function heterogeneous executor to obtain the reward function value of each network function heterogeneous executor.

5. The active defense system based on the intelligent control heterogeneous network function according to claim 4, characterized in that, The scheduling policy optimization module is based on a deep reinforcement learning algorithm.

6. The active defense system based on the intelligent control heterogeneous network function according to claim 5, characterized in that, The deep reinforcement learning algorithm is an actor-critic deep reinforcement learning algorithm. The scheduling policy optimization module is further used to, when the judgment result is abnormal or a preset cycle time slot is reached, perform learning optimization according to the actor-critic deep reinforcement learning algorithm based on the observation state of each network function heterogeneous executor and the reward function value of the observation state of each network function heterogeneous executor.

7. An active defense method based on the intelligent control heterogeneous network function, characterized in that, Including: Through the network function heterogeneous executor communication agent, when the source network function of the received signaling request is a conventional network function and the destination network function of the signaling request is a network function heterogeneous executor, a target network function heterogeneous executor in the network function heterogeneous executor set is selected for the signaling request according to the dynamically updated probability scheduling policy to route and forward the signaling request; Through the data collector, the state data of the network function heterogeneous executor set and the network function heterogeneous executor input-output sequence information of the network function heterogeneous executor communication agent are collected; Through the perception calculation module, based on the state data and the network function heterogeneous executor input-output sequence information, the comprehensive parameters of each network function heterogeneous executor in the network function heterogeneous executor set are determined, where the comprehensive parameters include: judgment result, trust degree result, observation state, and reward function value; Through the scheduling policy optimization module, the current probability scheduling policy is updated based on the comprehensive parameters of each network function heterogeneous executor, and the latest probability scheduling policy is sent down.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the active defense method based on the intelligent control heterogeneous network function as claimed in claim 7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the active defense method based on the intelligent control heterogeneous network function as claimed in claim 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the active defense method based on the intelligent control heterogeneous network function as claimed in claim 7.