Model training method, security policy optimization method, related device, equipment, storage medium and computer program product

By using AI-trained models to generate and optimize security strategies, the problem of low efficiency in security strategy optimization in existing technologies is solved, enabling the security protection system to be adaptive and have a rapid response capability, thus ensuring network stability.

CN121125138APending Publication Date: 2025-12-12CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411977184.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

The lack of automated security policy optimization schemes in existing security protection systems leads to low efficiency in manual security policy optimization and an inability to respond quickly to cybersecurity threats.

Method used

The system utilizes AI technology to train a first model to generate new security policies, and a second model to determine the availability of policy templates, thereby enabling automatic optimization and rapid response of security policies.

Benefits of technology

It improves the adaptability and response speed of the security protection system, enabling it to quickly respond to cybersecurity threats, prevent security vulnerabilities and attacks, and ensure the continuous and stable operation of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125138A_ABST
    Figure CN121125138A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, a security policy optimization method, a model training device, a security policy optimization device, first equipment, second equipment, a storage medium and a computer program product. The model training method comprises the steps that a first training data set and a second training data set are determined, and each sample in the first training data set comprises a security protection scheme for network security threats, a security policy corresponding to the security protection scheme and a policy template corresponding to the security policy; each sample in the second training data set comprises a security policy and a policy template corresponding to the security policy; a first model is trained by using an artificial intelligence (AI) technology and a first training data set, a second model is trained by using the AI technology and a second training data set, and the first model is used for generating a new security policy based on a security protection scheme for a network security threat and a security policy corresponding to the security protection scheme. The availability of the new security policy can be determined based on the second model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular to a model training method, a security strategy optimization method, related devices, equipment, storage media, and computer program products. Background Technology

[0002] In related technologies, such as Figure 1 As shown, the modular framework of a security protection system can include five layers: a resource layer (also known as a security resource layer), a decision layer, a control layer, a management layer, and a business layer. The resource layer can contain sensors and data sources for various security tools, and can provide raw security tool data such as network traffic, logs, and endpoint status. The decision layer can perform in-depth analysis of security tool data, discover and confirm security threats, generate security protection plans, and make corresponding response decisions. The control layer can uniformly control the resource layer based on resource orchestration and scheduling capabilities, realizing security capability scheduling. The management layer is responsible for managing modules, data, and tasks within the security protection system. The business layer can correspond to the customer's business platform, covering various customer businesses; furthermore, the business layer can interface with the security tools in the resource layer and is protected by these security tools.

[0003] However, in the aforementioned security protection systems, there is still no effective solution in terms of how to automatically optimize the security policies corresponding to the security protection schemes. Summary of the Invention

[0004] To address the related technical issues, embodiments of this application provide a model training method, a security strategy optimization method, related devices, equipment, storage media, and computer program products.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides a model training method, including:

[0007] A first training dataset and a second training dataset are determined. Each sample in the first training dataset includes a security protection scheme for network security threats, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy. Each sample in the second training dataset includes a security policy and a policy template corresponding to the security policy. The policy template can reflect one or more dimensions of the features of the corresponding security policy.

[0008] A first model is trained by using an artificial intelligence (AI) technology and the first training data set, and a second model is trained by using the AI technology and the second training data set, the first model being used for generating a new security policy based on a security protection scheme for a network security threat and a security policy corresponding to the security protection scheme, and the second model being used for generating a policy template corresponding to the security policy based on the security policy, and the availability of the new security policy being able to be determined based on the second model.

[0009] In the above scheme, the policy template comprises a security rule and a matching condition corresponding to the security rule, and the matching condition comprises one or more of the following:

[0010] a source address;

[0011] a destination address;

[0012] a protocol type;

[0013] a port number;

[0014] an action;

[0015] a priority.

[0016] The embodiments of the present application also provide a security policy optimization method, comprising:

[0017] obtaining a security protection scheme for a network security threat and a first security policy corresponding to the security protection scheme, inputting the security protection scheme and the first security policy into a first model, and obtaining a second security policy output by the first model;

[0018] inputting the first security policy into a second model, obtaining a first policy template output by the second model, inputting the second security policy into the second model, and obtaining a second policy template output by the second model;

[0019] in a case where the second policy template is consistent with the first policy template, determining that security policy optimization is completed; and

[0020] the first model and the second model are models trained by using any of the above model training methods.

[0021] In the above scheme, the method further comprises:

[0022] in a case where the second policy template is inconsistent with the first policy template, obtaining a third security policy corresponding to the security protection scheme, and optimizing the first model based on the third security policy.

[0023] The embodiments of the present application also provide a model training device, comprising:

[0024] a first processing unit configured to determine a first training data set and a second training data set, each sample in the first training data set comprising a security protection scheme for a network security threat, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy, each sample in the second training data set comprising a security policy and a policy template corresponding to the security policy, the policy template being capable of reflecting one or more dimensions of features of the corresponding security policy;

[0025] a second processing unit configured to train a first model using an AI technology and the first training data set, and train a second model using the AI technology and the second training data set, the first model being configured to generate a new security policy based on a security protection scheme for a network security threat and a security policy corresponding to the security protection scheme, the second model being configured to generate a policy template corresponding to a security policy based on the security policy, and the availability of the new security policy being capable of being determined based on the second model.

[0026] Embodiments of the present application further provide a security policy optimization apparatus, comprising:

[0027] a third processing unit configured to obtain a security protection scheme for a network security threat and a first security policy corresponding to the security protection scheme, input the security protection scheme and the first security policy into the first model, and obtain a second security policy output by the first model;

[0028] a fourth processing unit configured to input the first security policy into the second model, obtain a first policy template output by the second model, input the second security policy into the second model, and obtain a second policy template output by the second model;

[0029] a fifth processing unit configured to determine that security policy optimization is completed in a case where the second policy template is consistent with the first policy template; and

[0030] the first model and the second model are models trained using any of the model training methods described above.

[0031] Embodiments of the present application further provide a first device, comprising: a first communication interface and a first processor; wherein

[0032] the first processor is configured to:

[0033] determine a first training data set and a second training data set, each sample in the first training data set comprising a security protection scheme for a network security threat, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy, each sample in the second training data set comprising a security policy and a policy template corresponding to the security policy, the policy template being capable of reflecting one or more dimensions of features of the corresponding security policy;

[0034] train a first model using AI technology and the first training data set, the first model being used to generate a new security policy based on a security protection scheme for a network security threat and a security policy corresponding to the security protection scheme, and train a second model using AI technology and the second training data set, the second model being used to generate a policy template corresponding to a security policy based on the security policy, the availability of the new security policy being capable of being determined based on the second model.

[0035] The embodiments of the present application also provide a second device, comprising: a second communication interface and a second processor; wherein,

[0036] The second processor is configured to:

[0037] obtain a security protection scheme for a network security threat and a first security policy corresponding to the security protection scheme, input the security protection scheme and the first security policy into a first model, and obtain a second security policy output by the first model;

[0038] input the first security policy into a second model, obtain a first policy template output by the second model, input the second security policy into the second model, and obtain a second policy template output by the second model;

[0039] in a case where the second policy template is consistent with the first policy template, determine that security policy optimization is completed; wherein,

[0040] The first model and the second model are models trained by using any of the model training methods.

[0041] The embodiments of the present application also provide a first device, comprising: a first processor and a first memory for storing a computer program capable of running on the processor,

[0042] When the first processor runs the computer program, the first processor is configured to perform the steps of any of the model training methods.

[0043] The embodiments of the present application also provide a second device, comprising: a second processor and a second memory for storing a computer program capable of running on the processor,

[0044] The second processor is configured to execute the computer program, and perform the steps of any of the security policy optimization methods.

[0045] The embodiments of the present application further provide a storage medium having a computer program stored thereon, where the computer program, when executed by a processor, implements the steps of any of the model training methods or the steps of any of the security policy optimization methods.

[0046] The embodiments of the present application further provide a computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the steps of any of the model training methods or the steps of any of the security policy optimization methods.

[0047] The model training method, security policy optimization method, related apparatus, devices, storage media, and computer program products provided in this application embodiment include: determining a first training dataset and a second training dataset, wherein each sample in the first training dataset includes a security protection scheme for network security threats, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy; each sample in the second training dataset includes a security policy and a policy template corresponding to the security policy, wherein the policy template can reflect one or more dimensions of features of the corresponding security policy; training a first model using AI technology and the first training dataset, and training a second model using AI technology and the second training dataset; the first model is used to generate a new security policy based on the security protection scheme for network security threats and the security policy corresponding to the security protection scheme; the second model is used to generate a policy template corresponding to the security policy based on the security policy; and the availability of the new security policy can be determined based on the second model. The solution provided in this application utilizes AI technology and a specific dataset (i.e., the first training dataset) to train a first model for generating new security policies based on security protection schemes against network security threats and the corresponding security policies. This allows the first model to automatically optimize the security policies corresponding to the security protection schemes. Furthermore, it utilizes AI technology and another specific dataset (i.e., the second training dataset) to train a second model for generating policy templates corresponding to the security policies. These policy templates reflect one or more dimensions of the corresponding security policies, and the second model can determine the availability of the new security policies. This can also be understood as determining the executability of the new security policies; that is, the second model can determine whether the security policy optimization is complete or successful. Thus, by using both the first and second models, the adaptability and response speed of the security protection system can be significantly improved, enabling rapid response and self-repair against network security threats. This effectively prevents security vulnerabilities and attacks, ensuring the continuous and stable operation of the network. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of the security protection system architecture for related technologies;

[0049] Figure 2 This is a schematic flowchart of the model training method in an embodiment of this application;

[0050] Figure 3 This is a flowchart illustrating the security policy optimization method according to an embodiment of this application;

[0051] Figure 4This is a schematic diagram of the improved security protection system architecture used in this application example;

[0052] Figure 5 A flowchart illustrating the application of intelligent optimization security policy in this application;

[0053] Figure 6 This is a schematic diagram of the model training device structure according to an embodiment of this application;

[0054] Figure 7 This is a schematic diagram of the security strategy optimization device structure according to an embodiment of this application;

[0055] Figure 8 This is a schematic diagram of the structure of the first device according to an embodiment of this application;

[0056] Figure 9 This is a schematic diagram of the structure of the second device in the embodiment of this application. Detailed Implementation

[0057] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.

[0058] In related technologies, for Figure 1 The network security protection scheme of the security protection system shown (i.e., the security protection scheme against network security threats) may require manual threat handling after analyzing security data such as security threats; or, when it is necessary to use the generated security policy to handle security incidents, manual intervention may be required to optimize the security policy. Considering the complexity of the network environment and the large number of connected devices, manual intervention may be overwhelmed by security risks, leading to delays in security protection, increasing the likelihood of attacks on business systems, and slowing down the recovery of business systems to normal operation. Therefore, it is evident that the security protection system of related technologies has limitations in terms of automation and adaptability, and there is an urgent need to achieve automatic optimization of security policies.

[0059] Based on this, in various embodiments of this application, a first model is trained using AI technology and a specific dataset to generate new security policies based on security protection schemes against cybersecurity threats and the corresponding security policies. This allows the first model to be used to automatically optimize the security policies corresponding to the security protection schemes. Furthermore, a second model is trained using AI technology and another specific dataset to generate policy templates corresponding to the security policies. These policy templates reflect one or more dimensions of the corresponding security policies, and the second model can determine the availability of the new security policies. This can also be understood as determining the executability of the new security policies; that is, the second model can determine whether the security policy optimization is complete or successful. Thus, by using both the first and second models, the adaptability and response speed of the security protection system can be significantly improved, enabling rapid response and self-repair against cybersecurity threats. This effectively prevents security vulnerabilities and attacks, ensuring the continuous and stable operation of the network.

[0060] It should be noted that in the various embodiments of this application, "one or more" means at least one or more items, and "multiple" means at least two or more items.

[0061] This application provides a model training method, such as... Figure 2 As shown, the method includes:

[0062] Step 201: Determine the first training dataset and the second training dataset. Each sample in the first training dataset includes a security protection scheme for network security threats, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy. Each sample in the second training dataset includes a security policy and a policy template corresponding to the security policy. The policy template can reflect one or more dimensions of the features of the corresponding security policy.

[0063] Step 202: Train a first model using AI technology and the first training dataset, and train a second model using AI technology and the second training dataset. The first model is used to generate a new security policy based on the security protection scheme for network security threats and the security policy corresponding to the security protection scheme (which can be understood as the original security policy). The second model is used to generate a policy template corresponding to the security policy based on the security policy. The availability of the new security policy can be determined based on the second model.

[0064] In practical applications, the model training method provided in this application embodiment can be applied to... Figure 1The security protection system shown can be specifically applied to the control layer of the system; in other words, the first and second models after training can be deployed in the orchestration engine of the control layer, which can schedule and optimize the first and second models as needed. Furthermore, it is understood that the model training method provided in this application embodiment can be implemented by the orchestration engine (i.e., the orchestration engine trains the first and second models), or it can be implemented by other modules inside or outside the security protection system; this application embodiment does not limit this.

[0065] In practical applications, from a hardware implementation perspective, the model training method provided in this application embodiment can be implemented by a first device, which may include a server, etc. Furthermore, after the first and second models are trained, they can be deployed to a second device, whereby the second device utilizes the first and second models to optimize security policies; that is, the second device schedules and optimizes the first and second models as needed. Here, it can be understood that the second device also implements the functionality of the orchestration engine; the second device may also include a server, etc., and the second device may be the same as or different from the first device.

[0066] In practical applications, the decision engine in the decision layer of the security protection system can determine security protection schemes against network security threats. Specifically, when a user initiates an access request to a business system in the business layer of the security protection system through a terminal, all network traffic entering and leaving the business system must be detected by various security tools (also known as security protection devices, etc.) in the resource layer of the security protection system. These security tools can detect abnormal behavior in network traffic, discover malicious or suspected security threat traffic, and issue abnormal alerts (i.e., security alerts). The security management engine in the management layer of the security protection system can transmit security data from the resource layer to the decision engine. The security data can at least include security alert information from the security tools, and the security alert information can at least include relevant information about the network security threats detected by the corresponding security tools, such as the type of network security threat, attack behavior, attack attempts, and related risk model data. The decision engine can deploy one or more pre-trained deep learning networks, i.e., one or more specific intelligent large models. These models can classify the corresponding network security threats based on the security data and generate (i.e., determine) corresponding security protection schemes.

[0067] In practical applications, a security protection scheme can correspond to one or more security tools, meaning that the security protection scheme can be executed or implemented by one or more security tools. Accordingly, the security policy (i.e., the original security policy) corresponding to the security protection scheme can be understood as the current security policy of the security tool corresponding to the security protection scheme, or it can be understood as the pre-configured or already configured security policy in the security tool corresponding to the security protection scheme.

[0068] Specifically, the security protection scheme may include one or more of the following:

[0069] Identification of the corresponding security tool, such as its name and / or ID;

[0070] The target Internet Protocol (IP) address for the corresponding security policy, such as the IP addresses that need to be blocked / intercepted;

[0071] Operations corresponding to security policies, such as blocking / intercepting specific ports, blocking / intercepting specific IPs, configuring specific firewalls, etc.

[0072] The execution time of the corresponding security policy.

[0073] In practical applications, the policy template corresponding to the security policy can also be understood as the policy template of the corresponding security tool. The policy template can reflect one or more dimensions of the characteristics of the corresponding security policy, meaning it can reflect one or more security rules of the corresponding security tool. Specifically, one dimension of the security policy can be understood as a security rule of the corresponding security tool. Specifically, the policy template can include security rules and corresponding matching conditions. The security rules can be understood as operations used to protect against network security threats, such as blocking / intercepting specific ports, blocking / intercepting specific IPs, configuring specific firewalls, etc. The matching conditions can include one or more of the following: source address (e.g., source IP), destination address (e.g., destination IP), protocol type, port number, action (i.e., security protection action), and priority (i.e., security protection priority). For example, if the matching conditions include a source IP, the corresponding security rule can include blocking that source IP; that is, the policy template can include the IP to be blocked. If the matching conditions include a port number, the corresponding security rule can include blocking that port number; that is, the policy template can include the port number to be blocked.

[0074] In practical applications, policy templates can affect the normal operation of security tools. That is, a security policy can only operate normally in a security tool if the security rules contained in the policy template corresponding to the security policy are consistent with the security rules of the security tool. In other words, the security policy corresponds to (or belongs to) the security tool, meaning that the security policy is a security policy that the security tool can use.

[0075] Therefore, the availability of the new security policy can be determined based on the second model, that is, the executability of the new security policy can be determined based on the second model; in other words, the availability / executability of the new security policy can be determined based on the second model, or it can be understood as the completion / success of security policy optimization based on the second model. Specifically, for the original security policy and the new security policy, the policy templates corresponding to these two security policies can be determined using the second model respectively. If the two policy templates are consistent, it can be determined that the new security policy is a security policy that can be used by the security tools corresponding to the original security policy, that is, the availability / executability of the new security policy can be determined. In this way, through the second model, the problem of security tools stopping working due to fields in the new security policy that cannot be recognized by the security tools corresponding to the original security policy can be avoided, ensuring the normal operation of the security tools.

[0076] In practical applications, the specific number of samples included in the first training dataset and the specific method of collecting these samples can be set as needed, and this application embodiment does not limit this. The first model may include a generative model (such as a Transformer model). During the process of training the first model using AI technology and the first training dataset, for each sample in the first training dataset, a security protection scheme and the security policy corresponding to the security protection scheme can be used as input data, and the policy template corresponding to the security policy can be used as standard output data. The first model can generate output data based on the input data and pre-configured random model parameters, calculate the difference between the generated output data and the standard output data, calculate the gradient of model parameter changes through backpropagation to obtain new model parameters, and then generate new output data based on the input data and the new model parameters, thereby repeatedly updating the model parameters until the difference between the generated output data and the standard output data reaches a pre-configured target threshold (such as 0.01). In addition, the specific number of samples included in the second training dataset and the specific method of collecting these samples can also be set as needed, and this application embodiment does not limit this. The second model may also include a generative model (such as a Transformer model). In the process of training the second model using AI technology and the second training dataset, for each sample in the second training dataset, the security policy can be used as input data and the policy template corresponding to the security policy can be used as output data to train the second model. The specific training process of the second model can be similar to the training process of the first model, and will not be described in detail here.

[0077] In practical applications, as can be seen from the above description, after the first device completes the training of the first model and the second model, it can deploy the first model and the second model to the second device, and the second device can use the first model and the second model to optimize the security policy.

[0078] Based on this, embodiments of this application also provide a security policy optimization method, applied to a second device, such as... Figure 3 As shown, the method includes:

[0079] Step 301: Obtain a security protection scheme for network security threats and a first security policy corresponding to the security protection scheme; input the security protection scheme and the first security policy into a first model to obtain a second security policy output by the first model.

[0080] Step 302: Input the first security policy into the second model to obtain the first policy template output by the second model; input the second security policy into the second model to obtain the second policy template output by the second model;

[0081] Step 303: If the second policy template is consistent with the first policy template, determine that the security policy optimization is complete; wherein,

[0082] The first model and the second model are models trained using the model training methods provided by one or more of the above-mentioned technical solutions.

[0083] Here, it can be understood that the security policy optimization method provided in the embodiments of this application is applied to Figure 1 The security protection system shown can be specifically applied to the orchestration engine of the system's control layer, that is, the function of the orchestration engine can be implemented by the second device.

[0084] Furthermore, the first security policy can be understood as the original security policy, and the second security policy can be understood as the new security policy; the determination that the security policy optimization is complete can be understood as determining that the new security policy (i.e., the second security policy) has availability / executability, that is, determining that the security policy optimization is successful.

[0085] In practical applications, if the second policy template is inconsistent with the first policy template, it can be determined that the security policy optimization has failed, and the second security policy lacks usability / executability. In this case, manual optimization of the security policy can be performed.

[0086] Based on this, in one embodiment, the method may further include:

[0087] If the second strategy template is inconsistent with the first strategy template, obtain the third security strategy corresponding to the security protection scheme, and optimize the first model based on the third security strategy.

[0088] In practical applications, the third security strategy may include manually optimized security strategies.

[0089] The model training method provided in this application embodiment determines a first training dataset and a second training dataset. Each sample in the first training dataset includes a security protection scheme for network security threats, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy. Each sample in the second training dataset includes a security policy and a policy template corresponding to the security policy. The policy template can reflect one or more dimensions of features of the corresponding security policy. A first model is trained using AI technology and the first training dataset, and a second model is trained using AI technology and the second training dataset. The first model is used to generate a new security policy based on the security protection scheme for network security threats and the security policy corresponding to the security protection scheme. The second model is used to generate a policy template corresponding to the security policy based on the security policy. The availability of the new security policy can be determined based on the second model. The solution provided in this application utilizes AI technology and a specific dataset (i.e., the first training dataset) to train a first model for generating new security policies based on security protection schemes against network security threats and the corresponding security policies. This allows the first model to automatically optimize the security policies corresponding to the security protection schemes. Furthermore, it utilizes AI technology and another specific dataset (i.e., the second training dataset) to train a second model for generating policy templates corresponding to the security policies. These policy templates reflect one or more dimensions of the corresponding security policies, and the second model can determine the availability of the new security policies. This can also be understood as determining the executability of the new security policies; that is, the second model can determine whether the security policy optimization is complete or successful. Thus, by using both the first and second models, the adaptability and response speed of the security protection system can be significantly improved, enabling rapid response and self-repair against network security threats. This effectively prevents security vulnerabilities and attacks, ensuring the continuous and stable operation of the network.

[0090] The following section provides a more detailed description of this application with reference to application examples.

[0091] This application example provides a security policy optimization scheme based on security alerts. Figure 1 The security protection system shown (also known as the autonomous security protection system) has been improved, specifically the orchestration engine of the control layer. The improved architecture of the autonomous security protection system is as follows: Figure 4As shown, the improved orchestration engine (which can be referred to as the intelligent orchestration engine in the following description) deploys a policy generation model (i.e., the first model mentioned above) and a policy detection model (i.e., the second model mentioned above). The policy generation model can generate new security policies for security tools in response to security threats, enabling flexible and rapid automatic responses; the policy detection model can detect the feasibility of the generated new security policies, thereby ensuring effective responses.

[0092] The following is combined Figure 4 Describe the workflow between the modules in the improved security protection system framework of this application example.

[0093] The first stage involves users accessing various scenarios and services within the communication network, where traffic filtering and security assessments are performed. Specifically, when an end user initiates an access request to a communication network service or AI system service through a client, all network traffic entering and leaving the service system must be inspected by various security tools in the resource layer. These security tools can detect abnormal behavior in network traffic, identify malicious or suspected security threat traffic, record it as detailed access logs and monitoring scan results, and report abnormal alarms (i.e., security alarms), attack attempts, and risk model data (i.e., information related to network security threats) in real time to the traditional security tool engine and AI security tool engine in the resource layer. These engines can perform preliminary cleaning, deduplication, and aggregation of relevant traffic data and security information. The processed security data (which can at least contain information related to network security threats) can be reported to the security management engine. The security management engine can simultaneously perform data storage and other operations while transmitting security data to the intelligent decision engine.

[0094] Following this is the centralized analysis and intelligent distribution phase. Specifically, the intelligent decision engine can perform fine-grained classification of the characteristics of cybersecurity threats. This finely classified security data can be sent to the protection scheme generation model deployed within the intelligent decision engine. The advanced deep learning large-scale model network within this model further analyzes attack behaviors, identifying complex intrusion patterns and concealed threat sources. This intelligent large-scale model can learn and extract features from historical data, identifying complex attack patterns that are difficult to detect using conventional rules, thus improving the accuracy and coverage of threat detection. Afterwards, based on the analysis results, the intelligent decision engine can generate targeted security protection schemes and report them to the security management engine, providing guidance for subsequent policy decision-making, execution, and optimization.

[0095] Following this is the strategy integration and report compilation phase. Specifically, after receiving the protection strategy plan (i.e., security protection plan) from the intelligent big model at the decision-making layer, the security management engine can combine it with the current policy configuration (i.e., the original security policy) of the security protection devices (i.e., the relevant security tools) to generate a comprehensive protection analysis report. This report can at least include the security protection plan for cybersecurity threats, which may include the security tool name, policy target IP, policy operation, policy execution time, etc. The report may also include relevant information about cybersecurity threats. In addition, these reports can include not only threat overview (i.e., relevant information about cybersecurity threats) and protection recommendations (i.e., security protection plans), but also device performance statistics and system health assessments. This report information can then be sent to the intelligent orchestration engine at the control layer. The intelligent orchestration engine can combine the current policy (i.e., the original security policy) of the security tools to orchestrate and optimize the policy, ensuring the feasibility and compatibility of the new policy (i.e., the new security policy).

[0096] In this application example, the strategy optimization process of the intelligent orchestration engine is performed automatically. Specifically, as... Figure 5 As shown, the security protection scheme and the current security policy of the corresponding security tool (i.e., the configured security policy, or the original security policy) can be grouped together. The security protection scheme and the original security policy (i.e., the first security policy mentioned above) are then input into the policy generation model M1 (i.e., the first model mentioned above) to obtain the new policy (i.e., the new security policy, or the second security policy mentioned above) output by the policy generation model M1. Then, the templates of the original security policy and the new policy (i.e., the first policy template and the second policy template mentioned above) can be extracted using the policy detection model M2 (i.e., the second model mentioned above). It can then be determined whether the two policy templates are consistent. If the two policy templates are consistent, the policy optimization is considered complete. If the two policy templates are inconsistent, a log can be recorded, indicating that a new policy needs to be manually generated, i.e., the policy needs to be manually optimized. The manually generated policy (i.e., the third security policy mentioned above), the security protection scheme, and the corresponding policy template are then used as training data to train the policy generation model M1.

[0097] The training process of the policy generation model M1 can include: using known security protection schemes and corresponding security policies, and policy templates as training data, taking the security protection schemes and corresponding security policies as inputs, and the policy templates as standard output data, to train a generative model (such as a Transformer model); during training, the generative model can generate output data based on the input data and random model parameters, calculate the difference between the model's output data and the standard output data, calculate the gradient of model parameter changes through backpropagation to obtain new model parameters, and then generate new output data based on the input data and the new model parameters, thereby repeatedly updating the model parameters until the difference between the model's output data and the standard output data reaches a target threshold (such as 0.01). The training process of the policy detection model M2 can include: extracting known security policies and policy templates as training data, taking the security policies as input data, and the policy templates as output data, and training a generative model (such as a Transformer model) in a similar way to training the policy generation model M1, to obtain the policy detection model M2.

[0098] In this application example, the final protection policy (i.e., the new security policy) generated by the intelligent orchestration engine can be distributed to the traditional and AI tool engines at the resource layer, and further forwarded to various security protection devices to complete policy optimization. Specifically, the intelligent orchestration engine can return the generated orchestration scheme (which can at least contain the new security policy) to the security management engine, which can then issue instructions to the security tool engines to carry out protection according to the orchestration scheme. This enables seamless integration of security policy design and execution, thereby ensuring continuous protection and dynamic adaptability of business systems.

[0099] In this application example, after receiving security analysis results and security protection plans, the security management engine can combine information such as the current security policies (i.e., the original security policies) of the security protection devices to generate a protection policy plan that includes a threat overview and protection recommendations, as well as a performance statistics and health assessment report. Simultaneously, the management engine can integrate the protection policies into a unified security remediation command. This remediation command can be sent by the management engine to the intelligent orchestration engine in the control layer. The intelligent orchestration engine can then orchestrate and optimize the remediation command according to established security policies and business requirements, and send the remediation command containing the new security policy back to the security tool engine in the resource layer. This allows the various security protection devices to perform security repair and protection operations on the business systems, while simultaneously optimizing the security policies of the security protection devices.

[0100] In this application example, as described above, AI technology can intelligently extract policy templates for security tools. This allows for the timely extraction of policy templates for newly integrated security tools, increasing the compatibility and timeliness of security protection. For example, the format of a firewall security policy template can be shown in Table 1.

[0101]

[0102] Table 1

[0103] The solution provided in this application example has the following advantages:

[0104] 1) The strategy generation model M1, based on AI technology, has greater flexibility than the traditional rule-based strategy generation method and can adapt to diverse security protection scenarios.

[0105] 2) By introducing the policy detection model M2, the availability of the new security policy is improved, thereby enhancing the overall stability of the system;

[0106] 3) By using a generative model (i.e., policy generation model M1) to automate the optimization of security tool rules and policies, and then combining it with a generative model (i.e., policy detection model M2) to automatically extract templates of new and old security policies to confirm the executability of new security policies, the system's adaptability and response speed are significantly improved, enabling rapid response and self-repair. This effectively prevents security vulnerabilities and attacks, ensures the continuous and stable operation of the network, and at least solves the problem that the need for manual threat handling in related technologies leads to limitations in the overall automation and adaptability of the system.

[0107] To implement the model training method of this application embodiment, this application embodiment also provides a model training apparatus, such as... Figure 6 As shown, the device includes:

[0108] The first processing unit 601 is used to determine a first training dataset and a second training dataset. Each sample in the first training dataset includes a security protection scheme for network security threats, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy. Each sample in the second training dataset includes a security policy and a policy template corresponding to the security policy. The policy template can reflect one or more dimensions of the features of the corresponding security policy.

[0109] The second processing unit 602 is used to train a first model using AI technology and the first training dataset, and to train a second model using AI technology and the second training dataset. The first model is used to generate a new security policy based on a security protection scheme for network security threats and the security policy corresponding to the security protection scheme. The second model is used to generate a policy template corresponding to the security policy based on the security policy. The availability of the new security policy can be determined based on the second model.

[0110] In practical applications, the functions of the first processing unit 601 and the second processing unit 602 can be equivalent to the functions of the intelligent orchestration engine in the application example described above. Furthermore, the first processing unit 601 and the second processing unit 602 can be implemented by a processor in a model training device.

[0111] It should be noted that the model training device provided in the above embodiments is only illustrated by the division of the above program modules during model training. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules (such as the intelligent orchestration engine in the above application example) to complete all or part of the processing described above. In addition, the model training device and the model training method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0112] To implement the security policy optimization method of this application embodiment, this application embodiment also provides a security policy optimization device, such as... Figure 7 As shown, the device includes:

[0113] The third processing unit 701 is used to obtain a security protection scheme against network security threats and a first security policy corresponding to the security protection scheme, input the security protection scheme and the first security policy into a first model, and obtain a second security policy output by the first model.

[0114] The fourth processing unit 702 is used to input the first security policy into the second model to obtain the first policy template output by the second model; and to input the second security policy into the second model to obtain the second policy template output by the second model.

[0115] The fifth processing unit 703 is used to determine that security policy optimization is complete when the second policy template is consistent with the first policy template; wherein,

[0116] The first model and the second model are models trained using the model training methods provided by one or more of the above-mentioned technical solutions.

[0117] In one embodiment, the fifth processing unit 703 is further configured to obtain a third security policy corresponding to the security protection scheme when the second policy template is inconsistent with the first policy template, and optimize the first model based on the third security policy.

[0118] In practical applications, the functions of the third processing unit 701, the fourth processing unit 702, and the fifth processing unit 703 can be equivalent to the functions of the intelligent orchestration engine in the application example described above. Furthermore, the third processing unit 701, the fourth processing unit 702, and the fifth processing unit 703 can be implemented by the processor in the security policy optimization device.

[0119] It should be noted that the security policy optimization device provided in the above embodiments is only illustrated by the division of the above program modules when performing security policy optimization. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules (such as the intelligent orchestration engine in the above application example) to complete all or part of the processing described above. In addition, the security policy optimization device and the security policy optimization method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0120] Based on the hardware implementation of the above program modules, and in order to implement the model training method of the embodiments of this application, the embodiments of this application also provide a first device, such as... Figure 8 As shown, the first device 800 includes:

[0121] The first communication interface 801 is capable of exchanging information with other electronic devices;

[0122] The first processor 802 is connected to the first communication interface 801 to enable information interaction with other electronic devices and to execute the model training method provided by one or more of the above technical solutions when running a computer program.

[0123] The computer program is stored in the first memory 803.

[0124] Specifically, the first processor 802 is used for:

[0125] A first training dataset and a second training dataset are determined. Each sample in the first training dataset includes a security protection scheme for network security threats, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy. Each sample in the second training dataset includes a security policy and a policy template corresponding to the security policy. The policy template can reflect one or more dimensions of the features of the corresponding security policy.

[0126] A first model is trained using AI technology and the first training dataset, and a second model is trained using AI technology and the second training dataset. The first model is used to generate a new security policy based on a security protection scheme for cybersecurity threats and the security policy corresponding to the security protection scheme. The second model is used to generate a policy template corresponding to the security policy based on the security policy. The availability of the new security policy can be determined based on the second model.

[0127] It should be noted that the specific processing procedure of the first processor 802 can be understood by referring to the above method, and will not be repeated here.

[0128] Of course, in practical applications, the various components in the first device 800 are coupled together through the bus system 804. It can be understood that the bus system 804 is used to implement communication between these components. In addition to the data bus, the bus system 804 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 8 The general labeled all buses as Bus System 804.

[0129] The first memory 803 in this embodiment is used to store various types of data to support the operation of the first device 800. Examples of such data include any computer program used to operate on the first device 800.

[0130] The methods disclosed in the embodiments of this application can be applied to the first processor 802, or implemented by the first processor 802. The first processor 802 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in the first processor 802. The first processor 802 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The first processor 802 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the first memory 803. The first processor 802 reads the information in the first memory 803 and completes the steps of the aforementioned method in combination with its hardware.

[0131] In an exemplary embodiment, the first device 800 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.

[0132] Based on the hardware implementation of the above program modules, and in order to implement the security policy optimization method of this application embodiment, this application embodiment also provides a second device, such as... Figure 9 As shown, the second device 900 includes:

[0133] The second communication interface 901 is capable of exchanging information with other electronic devices;

[0134] The second processor 902 is connected to the second communication interface 901 to enable information interaction with other electronic devices and to execute the security policy optimization method provided by one or more of the above-mentioned technical solutions when running computer programs.

[0135] The computer program is stored in the second memory 903.

[0136] Specifically, the second processor 902 is used for:

[0137] Obtain a security protection scheme for network security threats and a first security policy corresponding to the security protection scheme; input the security protection scheme and the first security policy into a first model to obtain a second security policy output by the first model.

[0138] The first security policy is input into the second model to obtain the first policy template output by the second model; the second security policy is input into the second model to obtain the second policy template output by the second model.

[0139] If the second policy template is identical to the first policy template, the security policy optimization is determined to be complete; wherein,

[0140] The first model and the second model are models trained using the model training methods provided by one or more of the above-mentioned technical solutions.

[0141] In one embodiment, the second processor 902 is further configured to obtain a third security policy corresponding to the security protection scheme when the second policy template is inconsistent with the first policy template, and optimize the first model based on the third security policy.

[0142] It should be noted that the specific processing procedure of the second processor 902 can be understood by referring to the above method, and will not be repeated here.

[0143] Of course, in practical applications, the various components in the second device 900 are coupled together via a bus system 904. It can be understood that the bus system 904 is used to implement communication between these components. In addition to a data bus, the bus system 904 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 9 The general designated all buses as Bus System 904.

[0144] The second memory 903 in this embodiment is used to store various types of data to support the operation of the second device 900. Examples of such data include any computer program used to operate on the second device 900.

[0145] The methods disclosed in the above embodiments of this application can be applied to, or implemented by, the second processor 902. The second processor 902 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry of the hardware or by instructions in the form of software within the second processor 902. The second processor 902 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The second processor 902 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, specifically a second memory 903. The second processor 902 reads information from the second memory 903 and, in conjunction with its hardware, completes the steps of the aforementioned method.

[0146] In an exemplary embodiment, the second device 900 may be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components to perform the aforementioned method.

[0147] It is understood that the memories (first memory 803, second memory 903) in the embodiments of this application can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.

[0148] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a first memory 803 storing a computer program, which can be executed by a first processor 802 of a first device 800 to complete the steps of any of the aforementioned model training methods. Another example is a second memory 903 storing a computer program, which can be executed by a second processor 902 of a second device 900 to complete the steps of any of the aforementioned security policy optimization methods. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.

[0149] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by a first processor 802 of a first device 800 to complete the steps of any of the aforementioned model training methods; or, the computer program can be executed by a second processor 902 of a second device 900 to complete the steps of any of the aforementioned security policy optimization methods.

[0150] It should be noted that terms such as "first" and "second" are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0151] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.

[0152] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.

Claims

1. A model training method, characterized in that, include: A first training dataset and a second training dataset are determined. Each sample in the first training dataset includes a security protection scheme for network security threats, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy. Each sample in the second training dataset includes a security policy and a policy template corresponding to the security policy. The policy template can reflect one or more dimensions of the features of the corresponding security policy. A first model is trained using artificial intelligence (AI) technology and the first training dataset, and a second model is trained using AI technology and the second training dataset. The first model is used to generate a new security policy based on a security protection scheme for cybersecurity threats and the security policy corresponding to the security protection scheme. The second model is used to generate a policy template corresponding to the security policy based on the security policy. The availability of the new security policy can be determined based on the second model.

2. The method according to claim 1, characterized in that, The policy template includes security rules and corresponding matching conditions for the security rules. The matching conditions include one or more of the following: Source address; Destination address; Protocol type; Port number; action; Priority.

3. A security policy optimization method, characterized in that, include: Obtain a security protection scheme for network security threats and a first security policy corresponding to the security protection scheme; input the security protection scheme and the first security policy into a first model to obtain a second security policy output by the first model. Input the first security policy into the second model to obtain the first policy template output by the second model; Input the second security policy into the second model to obtain the second policy template output by the second model; If the second policy template is identical to the first policy template, the security policy optimization is determined to be complete; wherein, The first model and the second model are models trained using the model training method described in claim 1 or 2.

4. The method according to claim 3, characterized in that, The method further includes: If the second strategy template is inconsistent with the first strategy template, obtain the third security strategy corresponding to the security protection scheme, and optimize the first model based on the third security strategy.

5. A model training device, characterized in that, include: The first processing unit is configured to determine a first training dataset and a second training dataset. Each sample in the first training dataset includes a security protection scheme for network security threats, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy. Each sample in the second training dataset includes a security policy and a policy template corresponding to the security policy. The policy template can reflect one or more dimensions of the features of the corresponding security policy. The second processing unit is used to train a first model using AI technology and the first training dataset, and to train a second model using AI technology and the second training dataset. The first model is used to generate a new security policy based on a security protection scheme for network security threats and the security policy corresponding to the security protection scheme. The second model is used to generate a policy template corresponding to the security policy based on the security policy. The availability of the new security policy can be determined based on the second model.

6. A security strategy optimization device, characterized in that, include: The third processing unit is used to obtain a security protection scheme for network security threats and a first security policy corresponding to the security protection scheme, input the security protection scheme and the first security policy into a first model, and obtain a second security policy output by the first model. The fourth processing unit is used to input the first security policy into the second model to obtain the first policy template output by the second model; Input the second security policy into the second model to obtain the second policy template output by the second model; The fifth processing unit is used to determine that security policy optimization is complete when the second policy template is consistent with the first policy template; wherein, The first model and the second model are models trained using the model training method described in claim 1 or 2.

7. A first device, characterized in that, include: A first communication interface and a first processor; wherein... The first processor is configured to: A first training dataset and a second training dataset are determined. Each sample in the first training dataset includes a security protection scheme for network security threats, a security policy corresponding to the security protection scheme, and a policy template corresponding to the security policy. Each sample in the second training dataset includes a security policy and a policy template corresponding to the security policy. The policy template can reflect one or more dimensions of the features of the corresponding security policy. A first model is trained using AI technology and the first training dataset, and a second model is trained using AI technology and the second training dataset. The first model is used to generate a new security policy based on a security protection scheme for cybersecurity threats and the security policy corresponding to the security protection scheme. The second model is used to generate a policy template corresponding to the security policy based on the security policy. The availability of the new security policy can be determined based on the second model.

8. A second device, characterized in that, include: The second communication interface and the second processor; wherein... The second processor is used for: Obtain a security protection scheme for network security threats and a first security policy corresponding to the security protection scheme; input the security protection scheme and the first security policy into a first model to obtain a second security policy output by the first model. The first security policy is input into the second model to obtain the first policy template output by the second model; the second security policy is input into the second model to obtain the second policy template output by the second model. If the second policy template is identical to the first policy template, the security policy optimization is determined to be complete; wherein, The first model and the second model are models trained using the model training method described in claim 1 or 2.

9. A first device, characterized in that, include: A first processor and a first memory for storing computer programs capable of running on the processor. Wherein, when the first processor is used to run the computer program, it performs the steps of the method described in claim 1 or 2.

10. A second device, characterized in that, include: A second processor and a second memory for storing computer programs that can run on the processor. Wherein, when the second processor is used to run the computer program, it performs the steps of the method described in claim 3 or 4.

11. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 1 or 2, or the steps of the method of claim 3 or 4.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 1 or 2, or the steps of the method of claim 3 or 4.