Industrial host control method and system based on artificial intelligence

By implementing static and dynamic confidence models on industrial hosts combined with environmental assessment, combining eBPF and AI decision engine to analyze syscall flow data, the problem of insufficient monitoring of unknown attacks and dynamic behaviors in traditional security mechanisms is solved, and more efficient abnormal process detection and system protection are achieved.

CN120337219AActive Publication Date: 2025-07-18SHENZHEN INNOVATIVE CLOUD COMPUTER CO LTD

Patent Information

Application Number
CN202510516626.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-18
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Traditional security mechanisms are difficult to deal with new unknown attacks. Once the whitelist mechanism is tampered with and invalid, it lacks effective monitoring of the running behavior of authorized programs. Static security ignores dynamic changes, resulting in security vulnerabilities that continue to exist.

Method used

The industrial host control method based on artificial intelligence is adopted to monitor the process through a static trustworthiness model, a dynamic trustworthiness model and an operating environment evaluation model. The syscall flow data is analyzed in combination with the eBPF behavior collector and the AI decision engine, the process behavior entropy value is obtained and the moving average method is used for smoothing processing, and the circuit breaker instruction is triggered to prevent abnormal behavior.

Benefits of technology

It enhances system security, detects abnormal processes through multi-level monitoring, reduces the rate of error judgment, provides fast response and protection, adapts to complex security threats, and improves the detection ability of abnormal behavior and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337219A_ABST
    Figure CN120337219A_ABST
Patent Text Reader

Abstract

The invention provides an industrial host control method and system based on artificial intelligence. The method comprises the steps that a white list process is set, and when a process accesses the industrial host, the industrial host monitors whether the process accessing the industrial host is abnormal or not through application layer data before and during running of the process; then, real-time syscale flow data is collected through an eBPF behavior collector of an inner kernel layer of the industrial control host, and the syscale flow data is analyzed through an AI decision engine to judge whether the syscale flow data is abnormal or not; and acquiring a process behavior entropy value, smoothing the behavior entropy value by adopting a moving average method, and controlling triggering of a fusing instruction based on monitoring of an application layer and a kernel layer. Through the method and the corresponding system, the capability of detecting the abnormal process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention proposes an industrial host control method and system based on artificial intelligence, which relates to the technical field of software security protection. Background Art

[0002] Traditional security mechanisms often rely on predefined rules and signatures and are difficult to cope with new and unknown attacks. Although the whitelist mechanism can effectively prevent the running of unauthorized programs, once the programs in the whitelist are invaded or tampered with, traditional methods are helpless and lack effective monitoring of the runtime behavior of authorized programs. Most existing solutions only focus on static security and ignore the dynamic changes in runtime behavior, resulting in the continuous existence of security vulnerabilities. Summary of the Invention

[0003] The present invention provides an industrial host control method based on artificial intelligence to solve the above-mentioned problems:

[0004] An industrial host control method based on artificial intelligence proposed by the present invention, the method includes:

[0005] Set a whitelist process. When a process accesses the industrial host, the industrial host monitors the process through a static credibility model, a dynamic credibility model, and a runtime environment evaluation model;

[0006] Collect real-time syscall stream data through an eBPF behavior collector, and analyze the syscall stream data through an AI decision engine;

[0007] Obtain the process behavior entropy value, smooth the behavior entropy value by using the moving average method, and control the triggering of the fuse instruction based on the monitoring of the model and the analysis of the AI decision engine.

[0008] Further, set a whitelist process. When a process accesses the industrial host, the industrial host monitors the process through a static credibility model, a dynamic credibility model, and a runtime environment evaluation model, including:

[0009] Sort out all the processes running on the industrial host, screen out the whitelist processes, and store the whitelist process information in a database. The whitelist process information includes: process name, process ID, executable file path, file hash value, digital signature, startup parameters, running time range, required system permissions, user identity;

[0010] When a process accesses the industrial host, the industrial host first obtains the static credibility of the process accessing the industrial host through the static credibility model by comparing the SHA3-512 hash values of the first preset number of files in the process with the golden hash values in the database and comparing the metadata of the second preset number of files with the standard metadata in the database;

[0011] On the basis of considering the natural attenuation of trust over time, combining the comparison result of the system call sequence within a preset time with the set of normal system call sequences generated by historical behaviors, the dynamic credibility of the runtime behavior security of the process is obtained through the dynamic credibility model;

[0012] The weighted sum of the ratios of the current values of each hardware-level and network-level monitoring index to the baseline value is calculated, and after being mapped by the activation function, the quantitative evaluation value of the security of the environment in which the process runs is obtained through the operating environment evaluation model. The hardware-level and network-level monitoring indexes include: CPU instruction cycle abnormality rate, memory access entropy, DMA request frequency, network traffic rate, network traffic distribution, and system log information.

[0013] Further, specifically, the static credibility model is:

[0014]

[0015] Among them, T static represents the static credibility, n represents the number of files participating in the static credibility hash check, SHA3-51h(f i ) represents the SHA3-512 hash value of the i-th file, I(·) represents the indicator function, if the condition in the parentheses holds, the function value is 1, otherwise it is 0, DB g (f i ) represents the standard hash value of the i-th file, DB m (f j ) represents the standard metadata of the j-th file, m represents the number of files participating in the static credibility metadata check, MData(f j ) represents the metadata of the j-th file. The metadata includes file size, file creation time, file modification time, file access time, file owner, file access permission, software version number, compilation time, signature status, signature certificate information, basic attributes of dependent files, permission information, version information, digital signature information, and associated information, DB m (f j ) represents the standard metadata of the j-th file, stored in the database for comparison,

[0016] The dynamic credibility model is:

[0017]

[0018] Among them, T dynamic (t) represents the dynamic credibility at time t, λ(t) represents the trust attenuation coefficient, t represents time, L represents the number of system call sequences, and b k represents the k-th system call sequence, and B normal represents the set of normal system call sequences generated from historical behaviors;

[0019] The environmental assessment model is

[0020]

[0021] Among them, T c (t) represents the environmental security assessment value, σ(·) represents the activation function, and w j represents the weight of the j-th monitoring index, represents the current value of the j-th monitoring index, represents the baseline value of the j-th monitoring index.

[0022] Furthermore, real-time syscall stream data is collected through an eBPF behavior collector, and the syscall stream data is analyzed through an AI decision engine, including:

[0023] Deploy customized eBPF probes at the industrial control host kernel layer to capture all syscall events and obtain the key field data in all syscall events. The key field data includes: process PID, syscall number, timestamp, and specific parameters passed during the execution of the syscall;

[0024] Perform data cleaning and formatting on the key field data. The data cleaning and formatting include: filtering invalid calls, i.e., filtering syscall = 0 and abnormal PIDs, desensitizing the parameters, i.e., hashing sensitive parameters such as file paths and network addresses, and time series alignment, i.e., using Lamport logical clocks to ensure the order of cross-core events;

[0025] Extract the features from the cleaned and formatted key field data. The features include basic statistical features and context-related features. The basic statistical features include: call frequency, call type entropy, parameter diversity, and IO operation ratio. The context-related features include: process tree blood relationship and resource access chain;

[0026] Label the extracted feature data as normal and abnormal, divide the labeled feature data into a training set, a validation set, and a test set, construct an initial LSTM model, and train the initial LSTM model based on the training set.

[0027] Further, obtain the process behavior entropy value, smooth the behavior entropy value using the moving average method, and control the triggering of the fusing instruction based on the monitoring of the application layer and the kernel layer, including:

[0028] Obtain the process behavior entropy value and smooth the behavior entropy value using the moving average method;

[0029]

[0030] Among them, H(t) represents the behavior entropy value at time t, and the state space U includes: system call types, file access modes (i.e., the ratio of read times to write times), network connection topologies, and the resource usage of the process;

[0031]

[0032] When trigger the security fuse;

[0033] When trigger the security fuse;

[0034] Input the real-time detected data into the trained model to determine whether there is an anomaly. If there is an anomaly, trigger the security fuse.

[0035] An industrial host control system based on artificial intelligence proposed by the present invention, the system includes:

[0036] An application layer monitoring module, used to set the whitelist process. When the process accesses the industrial host, the industrial host monitors the process through a static credibility model, a dynamic credibility model, and a running environment evaluation model;

[0037] A kernel layer monitoring module, which collects real-time syscall stream data through an eBPF behavior collector and analyzes the syscall stream data through an AI decision engine;

[0038] A trigger fusing instruction module, which obtains the process behavior entropy value, smooths the behavior entropy value using the moving average method, and controls the triggering of the fusing instruction based on the monitoring of the model and the analysis of the AI decision engine.

[0039] Further, the application layer monitoring module includes:

[0040] A module for storing whitelist process information, used to sort out all the processes running on the industrial host, screen out the whitelist processes, and store the whitelist process information through a database. The whitelist process information includes: process name, process ID, executable file path, file hash value, digital signature, startup parameters, running time range, required system permissions, user identity;

[0041] A static credibility acquisition module, which is used to, when a process accesses the industrial host, first compare the SHA3-512 hash values of the first preset number of files in the process with the golden hash values in the database and compare the metadata of the second preset number of files with the standard metadata in the database, that is, obtain the static credibility of the process accessing the industrial host through a static credibility model;

[0042] A dynamic credibility acquisition module, which is used to, on the basis of considering the natural attenuation of trust over time, combine the comparison result between the system call sequence within a preset time and the set of normal system call sequences generated from historical behaviors, and obtain the dynamic credibility of the security of the process runtime behavior through a dynamic credibility model;

[0043] An environment evaluation module, which is used to perform weighted summation on the ratios of the current values of each hardware-level and network-level monitoring index to the baseline values, and then, after being mapped by an activation function, obtain a quantitative evaluation value of the security of the environment where the process runs through an operating environment evaluation model, and the hardware-level and network-level monitoring indexes include: CPU instruction cycle anomaly rate, memory access entropy, DMA request frequency, network traffic rate, network traffic distribution, and system log information.

[0044] Further, the kernel layer monitoring module further includes:

[0045] A static credibility model storage module. Specifically, the static credibility model is:

[0046]

[0047] Among them, T static represents the static credibility, n represents the number of files participating in the static credibility hash check, SHA3-51h(f i ) represents the SHA3-512 hash value of the i-th file, I(·) represents an indicator function, if the condition in the parentheses holds, the function value is 1, otherwise it is 0, DB g (f i ) represents the standard hash value of the i-th file, DB m (f j ) represents the standard metadata of the j-th file, m represents the number of files participating in the static credibility metadata check, MData(f j ) represents the metadata of the j-th file, and the metadata includes file size, file creation time, file modification time, file access time, file owner, file access permission, software version number, compilation time, signature status, signature certificate information, basic attributes of dependent files, permission information, version information, digital signature information, and associated information, DB m (f j) Represents the standard metadata of the j-th file, stored in the database for comparison,

[0048] A module for storing a dynamic credibility model, and the dynamic credibility model is:

[0049]

[0050] Among them, T dynamic (t) represents the dynamic credibility at time t, λ(t) represents the trust decay coefficient, t represents time, L represents the number of system call sequences, and b k represents the k-th system call sequence, and B normal represents the set of normal system call sequences generated from historical behaviors;

[0051] A module for storing an environment assessment model, and the environment assessment model is

[0052]

[0053] Among them, T c (t) represents the environmental security assessment value, σ(·) represents the activation function, and w j represents the weight of the j-th monitoring index, represents the current value of the j-th monitoring index, represents the baseline value of the j-th monitoring index.

[0054] Furthermore, the abnormal judgment module includes:

[0055] A module for capturing keyword field data, which is used to deploy customized eBPF probes in the kernel layer of the industrial control host, capture all syscall events, and obtain the keyword field data in all syscall events. The keyword field data includes: process PID, syscall number, timestamp, and specific parameters passed during the execution of the syscall;

[0056] A preprocessing module, which is used to perform data cleaning and formatting on the keyword field data. The data cleaning and formatting include: filtering invalid calls, that is, filtering syscall = 0 and abnormal PIDs, desensitizing the parameters, that is, performing hash processing on sensitive parameters such as file paths and network addresses, and time series alignment, that is, using Lamport logical clocks to ensure the order of cross-core events;

[0057] A feature extraction module, which is used to extract features from the keyword field data after cleaning and formatting. The features include basic statistical features and context correlation features. The basic statistical features include: call frequency, call type entropy, parameter diversity, and IO operation ratio. The context correlation features include: process tree blood relationship and resource access chain;

[0058] The division training module is used to label the extracted feature data as normal and abnormal, divide the labeled feature data into a training set, a validation set, and a test set, construct an initial LSTM model, and train the initial LSTM model based on the training set.

[0059] Further, the trigger fuse instruction module includes:

[0060] The process behavior entropy value acquisition module is used to acquire the process behavior entropy value and smooth the behavior entropy value by using the moving average method;

[0061]

[0062] Among them, H(t) represents the behavior entropy value at time t, and the state space U includes: system call type, file access mode (i.e., read / write ratio), network connection topology, and resource usage of the process;

[0063]

[0064] The judgment trigger module is used to when, trigger the safety fuse;

[0065] When when, trigger the safety fuse;

[0066] Input the real-time detected data into the trained model to judge whether there is an abnormality. If there is an abnormality, trigger the safety fuse.

[0067] The beneficial effects of the present invention are as follows: enhancing system security, multi-level monitoring, combining application layer and kernel layer monitoring, comprehensively checking processes from different angles. Application layer monitoring can discover some obvious abnormal behaviors, such as illegal file access, abnormal network connections, etc.; eBPF collection and AI analysis at the kernel layer can penetrate into the underlying operations of processes and detect some hidden attack behaviors, greatly improving the detection ability of abnormal processes; the behavior entropy value assists in judgment. The process behavior entropy value provides a quantitative index for evaluating the complexity and uncertainty of process behaviors. By monitoring and analyzing the behavior entropy value, processes with sudden changes in behavior patterns can be discovered in a timely manner, even if these changes are difficult to be detected in traditional rule detection, thereby further enhancing system security; reducing the misjudgment rate, smoothing processing by the moving average method. The smoothing of the behavior entropy value can reduce the interference of short-term fluctuations and make the judgment results more stable and reliable. Brief Description of the Drawings

[0068] Figure 1 It is a schematic diagram of an industrial host control method based on artificial intelligence according to the present invention. Detailed Embodiments

[0069] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.

[0070] In the following description, many specific details are set forth in order to fully understand the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0072] An embodiment of the present invention, an industrial host control method based on artificial intelligence, the method includes:

[0073] Set a whitelist process. When a process accesses the industrial host, the industrial host monitors the process through a static credibility model, a dynamic credibility model and a running environment evaluation model; collect real-time syscall stream data through an eBPF behavior collector, and analyze the syscall stream data through an AI decision engine; obtain the process behavior entropy value, smooth the behavior entropy value by using the moving average method, and control the triggering of the fusing instruction based on the monitoring of the model and the analysis of the AI decision engine, that is, set a whitelist process. When a process accesses the industrial host, the industrial host monitors whether there is an abnormality in the process accessing the industrial control host through the application layer data before and during the running of the process; then collect real-time syscall stream data through the eBPF behavior collector in the kernel layer of the industrial control host, analyze the syscall stream data through the AI decision engine, and judge whether there is an abnormality; obtain the process behavior entropy value, smooth the behavior entropy value by using the moving average method, and control the triggering of the fusing instruction based on the monitoring of the application layer and the kernel layer.

[0074] The working principle and effects of the above technical solution are as follows: Whitelist process setting and application layer monitoring. The administrator pre-determines a list of processes that can be trusted and adds these processes to the whitelist. The processes in the whitelist are considered legal and secure, and they are allowed to access the industrial host normally; Application layer data monitoring. When a process attempts to access the industrial host, the industrial host monitors the application layer data before and during the process operation. The application layer data contains various behavior information of the process, such as the resources requested by the process, the files operated on, the target addresses of network communications, etc. By analyzing this data, the industrial host can determine whether the process conforms to the normal behavior pattern. If the process is not in the whitelist or its behavior significantly differs from the normal behavior of the whitelist processes, it may be determined as an abnormal process; eBPF (Extended Berkeley Packet Filter) is a program that can run in the kernel and has the characteristics of high efficiency and flexibility. An eBPF behavior collector is deployed in the kernel layer of the industrial control host, which can capture the real-time syscall (system call) stream data of the process. Syscall is the interface for user programs to interact with the operating system kernel, and all operations of the process will ultimately be implemented through syscall. Therefore, the syscall stream data can reflect the underlying behavior of the process; The collected syscall stream data is transmitted to the AI decision engine for analysis. The AI decision engine usually uses machine learning or deep learning algorithms to learn and identify normal and abnormal syscall patterns. Through training with a large amount of historical data, the engine can establish a model of normal behavior. When receiving real-time syscall stream data, the engine will compare it with the normal model. If a significant deviation is found, it is determined that the process is abnormal; The process behavior entropy value is used to measure the uncertainty and complexity of the process behavior. The state space includes multiple dimensions such as system call types, file access modes, network connection topologies, and process resource usage. By statistically analyzing these dimensions, the entropy value of the process behavior can be calculated. The higher the entropy value, the more complex and unpredictable the process behavior is, and the greater the possibility of abnormality; To reduce the impact of short-term fluctuations in the behavior entropy value on the judgment result, the moving average method is used to smooth the behavior entropy value. The moving average method calculates the average value of the behavior entropy value within a certain time window, making the change of the entropy value more stable and facilitating the observation and analysis of long-term behavior trends; By comprehensively applying the monitoring results of the application layer and the kernel layer, as well as the smoothed process behavior entropy value, to control the triggering of the fuse instruction. If the application layer monitoring finds that the process is abnormal, or the AI decision engine analyzes the syscall stream data and determines that the process is abnormal, or the process behavior entropy value exceeds the preset threshold, then the fuse instruction is triggered to cut off the connection between the process and the industrial host to prevent it from causing further damage to the system. Enhance system security,

[0075] Multi-level monitoring combines application-layer and kernel-layer monitoring to comprehensively inspect processes from different perspectives. Application-layer monitoring can detect some obvious abnormal behaviors, such as illegal file access and abnormal network connections. The eBPF collection and AI analysis at the kernel layer can delve into the underlying operations of processes and detect some hidden attack behaviors, greatly improving the detection ability of abnormal processes. The behavioral entropy value assists in judgment. The process behavioral entropy value provides a quantitative indicator for evaluating the complexity and uncertainty of process behaviors. By monitoring and analyzing the behavioral entropy value, processes with sudden changes in behavioral patterns can be detected in a timely manner, even if these changes are difficult to be discovered in traditional rule detection, thus further enhancing the system's security. Reducing the false positive rate. The moving average method is used for smoothing processing. The smoothing of the behavioral entropy value can reduce the interference of short-term fluctuations, making the judgment results more stable and reliable. For example, a process may experience a short-term increase in the behavioral entropy value due to normal business requirements at certain specific moments, but the moving average method can filter out these short-term fluctuations and avoid misjudging as abnormal behaviors. The learning ability of the AI decision engine. The AI decision engine can continuously optimize the recognition ability of normal and abnormal behaviors through learning and training on a large amount of historical data. It can adapt to different system environments and business scenarios and reduce false positives caused by environmental changes or business requirement changes. Quick response and protection. Real-time monitoring and analysis. The entire monitoring and analysis process is carried out in real time, and abnormal behaviors of processes can be detected in a timely manner. Once an anomaly is detected, the system can quickly trigger a fuse instruction to cut off the connection between the abnormal process and the industrial host, prevent the spread of attacks and further damage to the system, and ensure the stable operation of the industrial host. Whitelist mechanism: The setting of whitelist processes provides a certain degree of flexibility. Administrators can dynamically adjust the process list in the whitelist according to actual business requirements and security policies. At the same time, for new processes not in the whitelist, the system can also determine their security through real-time monitoring and analysis, with good scalability. The trainability of the AI decision engine: The AI decision engine can adapt to new attack patterns and behavioral characteristics by continuously updating training data. As the system runs and data accumulates, the performance of the engine can be continuously improved to better cope with increasingly complex security threats.

[0076] In one embodiment of the present invention, a whitelist process is set. When a process accesses the industrial host, the industrial host monitors the process through a static credibility model, a dynamic credibility model, and a running environment evaluation model, including:

[0077] Set the whitelist process of the industrial host, and store the information of the whitelist process in the database; when a process accesses the industrial host, the industrial host obtains the static credibility of the process through the comparison of hash values and metadata in the static credibility model; based on the comparison result of the set of normal system call sequences constructed by the system call sequence and historical behavior within a preset time, obtain the dynamic credibility of the security of the process runtime behavior through the dynamic credibility model; obtain the quantitative evaluation value of the security of the environment where the process runs through the environment evaluation model, that is, sort out all the processes running on the industrial host, filter out the whitelist processes, and store the whitelist process information in the database. The whitelist process information includes: process name, process ID, executable file path, file hash value, digital signature, startup parameters, running time range, required system permissions, user identity; when a process accesses the industrial host, the industrial host first compares the SHA3-512 hash value of the first preset number of files in the process with the golden hash value in the database and compares the metadata of the second preset number of files with the standard metadata in the database, that is, obtain the static credibility of the process accessing the industrial host through the static credibility model; on the basis of considering the natural attenuation of trust over time, combine the comparison result of the system call sequence within a preset time and the set of normal system call sequences generated by historical behavior, and obtain the dynamic credibility of the security of the process runtime behavior through the dynamic credibility model; perform weighted summation on the ratio of the current value of each hardware-level and network-level monitoring index to the baseline value, and then through the activation function mapping, obtain the quantitative evaluation value of the security of the environment where the process runs through the running environment evaluation model. The hardware-level and network-level monitoring indexes include: CPU instruction cycle anomaly rate, memory access entropy, DMA request frequency, network traffic rate, network traffic distribution, and system log information.

[0078] The working principle and effect of the above technical solution are as follows: whitelist process combing and storage, process combing and screening, comprehensive combing of all processes running on the industrial host, screening out processes that are considered safe and reliable according to the system's security policies and business requirements, and including them in the whitelist; storing detailed information of the whitelist process, such as process name, process ID, executable file path, file hash value, digital signature, startup parameters, running time range, required system permissions, user identity, etc., in the database. This information serves as an important basis for subsequent process legitimacy judgment; when a process accesses the industrial host, select a first preset number of files from the process, calculate their SHA3-512 hash values, and compare them with the golden hash values of the corresponding files stored in the database. If the hash values are consistent, it means that the file content has not been tampered with; if they are inconsistent, there may be a situation where the file has been maliciously modified; at the same time, select a second preset number of files, and compare their metadata (such as file size, creation time, modification time, etc.) with the standard metadata stored in the database. Abnormal changes in metadata may also indicate that there are security risks in the file; the static credibility of the process is calculated through a static credibility model based on the results of hash value comparison and metadata comparison, and this credibility reflects the security of the process at the file level; trust decays naturally, taking into account that trust decays naturally over time, that is, over time, even if the process was secure before, it may become unsafe due to changes in the system environment or attacks. Therefore, a trust decay mechanism is introduced when evaluating the dynamic credibility of the process; system call sequence comparison, combining the system call sequence of the process within a preset time, and comparing it with the normal system call sequence set generated by the Markov chain model constructed from historical behavior. The Markov chain model can describe the normal system call behavior pattern of the process. If the system call sequence of the current process is significantly different from the normal sequence, it means that the running behavior of the process may be abnormal; comprehensive trust The results of the attenuation and system call sequence comparison are used to calculate the dynamic credibility of the runtime behavior of the process through a dynamic credibility model. This credibility reflects the behavioral security of the process during operation; the current values of various hardware-level and network-level monitoring indicators are collected, including CPU instruction cycle exception rate, memory access entropy, DMA request frequency, network traffic rate, network traffic distribution and system log information, etc. These indicators can reflect the state of the environment in which the process is running; the ratio of the current value of each monitoring indicator to the baseline value is weighted and summed, and the weight of each monitoring indicator is set according to its importance to the system security. Then, the result of the weighted summation is mapped through an activation function to convert it into a quantitative evaluation value within a specific range; through the operating environment evaluation model, the quantitative evaluation value is used as the evaluation result of the security of the environment in which the process is running, which reflects the impact of the process running environment on its security.Enhance system security. By comprehensively evaluating processes from three dimensions: the static file level, the dynamic running behavior level, and the running environment level, it can more accurately identify the security risks of processes. For example, the static credibility model can detect whether files have been tampered with, the dynamic credibility model can discover abnormalities in the running behavior of processes, and the running environment evaluation model can evaluate the impact of environmental factors on process security, thus effectively preventing various types of attacks; by comparing hash values, metadata, system call sequences, and monitoring indicators in real time, it can timely detect abnormal changes in processes. Once an abnormality is detected, the system can take corresponding measures, such as restricting process permissions, terminating processes, etc., to prevent the occurrence and spread of security incidents; improve the accuracy of trust evaluation. Considering the time factor, the dynamic credibility model takes into account the natural attenuation of trust over time, making the trust evaluation more in line with the actual situation. As time goes by, the security of processes may change. By introducing a trust attenuation mechanism, the trust level of processes can be adjusted in a timely manner to avoid ignoring potential security risks due to long-term trust; model based on historical behavior. Use the Markov chain model to generate a set of normal system call sequences as the reference standard for dynamic credibility evaluation. This model is constructed based on the historical behavior of processes and can accurately describe the normal behavior patterns of processes, improving the ability to identify abnormal behaviors; comprehensive monitoring indicators. The running environment evaluation model comprehensively considers multiple hardware-level and network-level monitoring indicators and can comprehensively reflect the state of the environment in which the process is running. In a complex industrial environment, various factors may affect the security of processes. By monitoring and evaluating these indicators, the impact of environmental changes on process security can be detected in a timely manner and corresponding measures can be taken for prevention; scalability. This technical solution has a certain degree of scalability and can add or adjust monitoring indicators and evaluation models according to actual needs to adapt to the security requirements of different industrial scenarios; provide quantitative evaluation results, including static, dynamic credibility, and environmental evaluation values: Through the static credibility model, dynamic credibility model, and running environment evaluation model, quantitative evaluation values of the static credibility, dynamic credibility, and running environment security of processes are obtained respectively, providing an intuitive reference for system administrators to help them better understand the security status of processes and make more accurate decisions.

[0079] In one embodiment of the present invention, the static credibility model obtains the static credibility of the process by comparing the SHA3-512 hash value of the first preset number of files in the process with the golden hash value in the database and comparing the metadata of the second preset number of files with the standard metadata in the database. Specifically, the static credibility model is:

[0080]

[0081] Wherein, T static represents the static credibility, n represents the number of files participating in the static credibility hash check, SHA3-51h(fi ) represents the SHA3-512 hash value of the i-th file, I(·) represents the indicator function, if the condition in the parentheses holds, the function value is 1, otherwise it is 0, DB g (f i ) represents the standard hash value of the i-th file, DB m (F j ) represents the standard metadata of the j-th file, m represents the number of files participating in the static credibility metadata verification, MData(f j ) represents the metadata of the j-th file, the metadata includes file size, file creation time, file modification time, file access time, file owner, file access permission, software version number, compilation time, signature status, signature certificate information, basic attributes of dependent files, permission information, version information, digital signature information and associated information, DB m (f j ) represents the standard metadata of the j-th file, stored in the database for comparison,

[0082] The dynamic credibility model obtains the dynamic credibility by combining the comparison result between the system call sequence within a preset time and the set of normal system call sequences generated from historical behaviors. Specifically, the dynamic credibility model is as follows:

[0083]

[0084] Among them, T dynamic (t) represents the dynamic credibility at time t, λ(t) represents the trust decay coefficient, t represents time, L represents the number of system call sequences, b k represents the k-th system call sequence, B normal represents the set of normal system call sequences generated from historical behaviors;

[0085] The construction of the normal system call sequence is essentially to model the historical system call sequence, and its mathematical formal definition and calculation process are as follows:

[0086] Define the state space, let the set of system call types be where each s i represents a specific system call (such as open, read, execve), by statistically analyzing the system call sequences generated by all legal whitelist applications in the historical data, constructing the state transition probability matrix P ∈ R n×n ,

[0087] Calculate the transition probability,

[0088] (1) Construct the original count matrix, and statistically analyze the first-order transition frequencies from the historical data:

[0089] N(si →s j ) = count(the number of times system call s i is immediately followed by s j in the historical data)

[0090] For each state s i , calculate the probability distribution of its transition to other states:

[0091]

[0092] where ∈ represents the smoothing factor (usually taking Laplace smoothing ε = 1×10^{-6}) to avoid the zero - probability problem, represents the sum of all transition times starting from s i ;

[0093] For a given system call sequence b = (b1, b h ,..., b K ), the probability that it belongs to normal behavior is:

[0094]

[0095] In actual calculation, to avoid numerical underflow, the logarithmic probability form is adopted:

[0096]

[0097] For the real - time captured syscall sequence, b current = (b1,..., b K ), the rule for judging whether it is abnormal is: Then it is judged as abnormal;

[0098] where the threshold τ threshold is calculated through historical normal data:

[0099] τ threshold = μ normal - 3σ normal

[0100] μ normal represents the mean of the logarithmic probabilities of historical normal sequences, and σ normal represents the standard deviation of the logarithmic probabilities of historical normal sequences;

[0101] To adapt to the changes in system behavior, B normal can be incrementally updated periodically. Add the syscall sequences that have not triggered the fuse into the historical data set and update the sliding window:

[0102] N new (s i →s j) = αN old (s i →s j ) + (1 - α)N recent (s i →s j )

[0103] where α ∈ [0, 1] is the forgetting factor;

[0104] The environment assessment model performs a weighted sum of the ratios of the current values of each hardware - level and network - level monitoring index to the baseline values, and then obtains a quantitative assessment value of the security of the environment in which the process runs after mapping through an activation function. The hardware - level and network - level monitoring indexes include: CPU instruction cycle anomaly rate, memory access entropy, DMA request frequency, network traffic rate, network traffic distribution, and system log information. Specifically, the environment assessment model is

[0105]

[0106] where, T c (t) represents the environmental security assessment value, σ(·) represents the activation function, w j represents the weight of the j - th monitoring index, represents the current value of the j - th monitoring index, represents the baseline value of the j - th monitoring index.

[0107] The working principle and effects of the above technical solution are as follows: The static credibility model mainly evaluates the security of a process at the file level by verifying the hash value and metadata of a file. The specific steps are as follows: Hash value verification: For the n files participating in the static credibility hash verification, calculate the SHA3-512 hash value of each file, and compare it with the standard hash value of the corresponding file stored in the database. Use an indicator function to determine whether they are equal. If they are equal, the function value is 1, otherwise it is 0. Sum up the comparison results of all files and divide by the number of files to obtain the score of the hash value verification; For the m files participating in the static credibility metadata verification, obtain the metadata of each file and compare it with the standard metadata of the corresponding file stored in the database. Similarly, use an indicator function to determine whether they are equal. Sum up the comparison results of all files and divide by the number of files to obtain the score of the metadata verification. The static credibility calculation adds the hash value verification score and the metadata verification score to obtain the final static credibility. The value range of the static credibility is between 0 and 2. The closer the value is to 2, the higher the static security of the file; The dynamic credibility model mainly considers the natural decay of trust over time and the matching degree between the system call sequence of a process and the normal sequence to evaluate the security of the process's runtime behavior. The specific steps are as follows: Introduce a trust decay coefficient, which is dynamically adjusted over time according to the system's historical attack situation and the current security posture. As time goes by, trust will naturally decay, and this decay process is described by an exponential function; Count the number of system call sequences within a preset time. For each system call sequence, determine whether it belongs to the set of normal system call sequences generated by the Markov chain model constructed from historical behaviors. Use an indicator function to judge. If it is not in the set of normal system call sequences B normalAmong them, the function value is 1, otherwise it is 0. Sum up the comparison results of all system call sequences and divide by the number of sequences to obtain the proportion of abnormal system call sequences; The dynamic credibility calculation subtracts the proportion of abnormal system call sequences from 1, and then multiplies by the coefficient after trust attenuation to obtain the dynamic credibility at time t. The value range of the dynamic credibility is between 0 and 1, and the value closer to 1 indicates that the dynamic behavior of the process is safer; The environment assessment model mainly assesses the security of the environment in which the process runs by comparing the current values of hardware-level and network-level monitoring indicators with the baseline values, and performing weighted summation and activation function mapping. The specific steps are as follows: Monitoring indicator collection and comparison: Collect the current values of hardware-level and network-level monitoring indicators, compare them with the corresponding baseline values, and calculate the ratio between the two; Assign a weight to each monitoring indicator, multiply the ratio of each monitoring indicator by its weight, and then add up all the results to obtain the result of weighted summation; Input the result of weighted summation into the activation function for mapping to obtain the final environmental security assessment value. The activation function usually maps the input value to a specific interval, and the value closer to 1 indicates that the environment in which the process runs is safer. Enhance security, multi-dimensional assessment, comprehensively evaluate the process through three dimensions: static credibility, dynamic credibility, and environmental security assessment, and potential security risks can be discovered from different perspectives. The static credibility model can detect whether the file has been tampered with, the dynamic credibility model can discover abnormalities in the running behavior of the process, and the environment assessment model can evaluate the impact of the running environment on the process security, thus effectively preventing various types of attacks; Detect abnormalities in a timely manner: Compare the hash values, metadata, system call sequences, and monitoring indicators in real time, and abnormalities in the process and environment can be discovered in a timely manner. Once an abnormality is discovered, the system can take corresponding measures, such as restricting the process permissions, terminating the process, etc., to prevent the occurrence and spread of security incidents; Improve the accuracy of trust assessment, considering the time factor, the dynamic credibility model takes into account the natural attenuation of trust over time, making the trust assessment more in line with the actual situation. As time goes by, the security of the process may change. By introducing a trust attenuation mechanism, the trust level of the process can be adjusted in a timely manner to avoid ignoring potential security risks due to long-term trust; Model based on historical behavior: Use the Markov chain model to generate a set of normal system call sequences as the reference standard for dynamic credibility assessment. This model is constructed based on the historical behavior of the process, can accurately describe the normal behavior pattern of the process, and improves the ability to identify abnormal behavior; The environment assessment model comprehensively considers multiple hardware-level and network-level monitoring indicators and can comprehensively reflect the state of the environment in which the process runs; In a complex industrial environment, various factors may affect the security of the process. By monitoring and evaluating these indicators, the impact of environmental changes on the process security can be discovered in a timely manner, and corresponding measures can be taken for prevention.

[0108] In one embodiment of the present invention, real-time syscall stream data is collected through the eBPF behavior collector in the kernel layer of the industrial control host, and the AI decision engine analyzes the syscall stream data to determine whether there is an abnormality, including:

[0109] The eBPF probe captures all syscall events and obtains the key field data in all syscall events;

[0110] Clean and format the key field data, extract the features in the cleaned and formatted key field data, divide the extracted feature data into a training set, a validation set, and a test set, construct an initial LSTM model, and train the initial LSTM model based on the training set. That is, a customized eBPF probe is deployed in the kernel layer of the industrial control host to capture all syscall events and obtain the key field data in all syscall events. The key field data includes: process PID, syscall number, timestamp, and specific parameters passed when the syscall is executed;

[0111] Clean and format the key field data. The data cleaning and formatting include: filtering invalid calls, i.e., filtering syscall = 0 and abnormal PIDs, desensitizing the parameters, i.e., hashing sensitive parameters such as file paths and network addresses, and time series alignment, i.e., using Lamport logical clocks to ensure the order of cross-core events;

[0112] Extract the features in the cleaned and formatted key field data. The features include basic statistical features and context-related features. The basic statistical features include: call frequency, call type entropy, parameter diversity, and IO operation ratio. The context-related features include: process tree blood relationship and resource access chain;

[0113] Label the extracted feature data as normal and abnormal, divide the labeled feature data into a training set, a validation set, and a test set, construct an initial LSTM model, and train the initial LSTM model based on the training set.

[0114] The working principle and effects of the above technical solution are as follows: eBPF (Extended Berkeley Packet Filter) is a powerful technology in the Linux kernel that allows users to dynamically load and execute custom programs in the kernel without modifying the kernel code. Customized eBPF probes are deployed at the kernel layer of the industrial control host, and these probes are mounted at key positions related to the system calls (syscall) in the kernel; when a process in the system initiates a system call, the probes are triggered, thereby capturing all syscall events; then, key field data is extracted from these events. For example, the process PID can be used to identify which process initiated the call, the syscall number can clarify which specific system function was called, the timestamp records the time when the call occurred, and the specific parameters contain the detailed information passed during the call. These information are crucial for subsequent analysis of process behavior; data cleaning and formatting, filtering invalid calls: syscall = 0 usually represents an invalid system call, and abnormal PIDs (such as negative PIDs or out of the reasonable range) may be caused by data errors or abnormal situations. Filtering these invalid data can reduce noise interference and enable subsequent analysis to be based on more accurate and effective data; parameter desensitization, file paths, network addresses, etc. are sensitive parameters that contain important information of the system. To protect the security and privacy of the system, these parameters are hashed. Hashing will convert sensitive information into a hash value of a fixed length, which not only retains the characteristics of the data but also avoids the direct leakage of sensitive information; in a multi-core system, the occurrence times of events on different cores may be inconsistent in order due to hardware and scheduling reasons. Lamport logical clock is an algorithm used to assign logical timestamps to events in a distributed system. By using it, the orderliness of cross-core events can be guaranteed, enabling subsequent feature extraction and analysis to be based on the correct time order; feature extraction, basic statistical features: call frequency: count the number of system calls within a unit time, which reflects the activity intensity of the process. If the call frequency of a certain process suddenly increases or decreases, it may imply that the behavior of the process has become abnormal; call type entropy: used to measure the diversity of system call types. The higher the entropy value, the more dispersed the call types; the lower the entropy value, the more concentrated the call types. Abnormal processes may exhibit a different distribution of call types from normal processes; parameter diversity: analyze the changes in system call parameters, reflecting the flexibility of the process in calling system functions.An abnormal change in parameter diversity may imply abnormal process behavior. For example, a malicious process may use some uncommon parameters to perform specific operations; IO operation ratio: Calculate the ratio of system calls related to input / output (IO) in the total number of calls, which reflects the degree of IO activity of the process. An abnormal IO operation ratio may be related to behaviors such as data leakage and malicious file reading and writing; Contextual association features: Process tree lineage: By analyzing the parent-child and ancestor relationships of processes, understand the creation and inheritance of processes. An abnormal process tree structure may indicate the creation or disguise of malicious processes. For example, an unknown process suddenly creates a large number of child processes; Resource access chain: Record the access order and relationships of processes to system resources (such as files, network ports, devices, etc.), and can discover abnormal resource access patterns, such as illegal file access and abnormal network connections; Data annotation and dataset division: Annotate the extracted feature data and divide it into two categories: normal and abnormal. The basis for annotation can be security rules, expert experience, or known normal and abnormal behavior patterns in history. Then divide the annotated feature data into a training set, a validation set, and a test set; The training set is used to train the initial LSTM model to let the model learn the feature patterns of normal and abnormal behaviors; The validation set is used to evaluate the performance of the model during training and adjust the hyperparameters of the model (such as learning rate, number of neurons in the hidden layer, etc.) to prevent the model from overfitting; The test set is used to finally evaluate the generalization ability of the model to ensure that the model can also perform well on unseen data; LSTM (Long Short-Term Memory Network) is a special type of Recurrent Neural Network (RNN) that can handle long-term dependencies in sequential data. In this solution, the feature sequence of system calls is used as input, and the LSTM model will learn the patterns and regularities in the sequence. During training, the model will continuously adjust its own parameters (such as weights and biases) according to the input feature data and the corresponding annotations (normal or abnormal) to minimize the error between the prediction result and the true label. Through multiple iterations of training, the model gradually learns the feature representations of normal and abnormal behaviors, thus acquiring the ability to predict anomalies for new data. Enhance the security of industrial control systems, anomaly behavior detection: By comprehensively monitoring and analyzing system call events, it can timely detect abnormal process behaviors, such as the activities of malware and illegal data access. The LSTM model can learn normal system call patterns and issue an alarm when behaviors that do not conform to the normal pattern occur, helping administrators take timely measures to prevent the system from being attacked; Privacy protection: Parameter desensitization processing ensures that sensitive information in the system will not be leaked, protecting the privacy and security of industrial control hosts and avoiding potential risks caused by information leakage; Improve system stability, performance optimization: By analyzing features such as the frequency of system calls and the IO operation ratio, the load situation and performance bottlenecks of the system can be understood.Administrators can optimize the system based on this information, allocate resources reasonably, and improve the operation efficiency and stability of the system. For example, if it is found that the high call frequency of a certain process causes a decline in system performance, the process can be optimized or restricted; fault prediction, abnormal system call patterns may be early signs of system failures. By continuously monitoring and analyzing system call events, potential faults can be detected in advance, preventive maintenance can be carried out, system downtime can be reduced, and losses caused by system failures can be minimized; providing interpretability and traceability, behavior analysis, the extracted basic statistical features and context-related features provide administrators with detailed process behavior information. By analyzing these features, the activities of the process can be deeply understood, the root causes of abnormal behaviors can be found, and strong support can be provided for security audits and fault troubleshooting. For example, the origin of an abnormal process can be traced through the process tree blood relationship; event tracing, information such as timestamps and process tree blood relationships enables administrators to trace the occurrence process of system call events, understand the sequence and correlation of events, helps to restore the full picture of the events, and provides a basis for subsequent investigations and handling; this solution is based on eBPF technology and has good adaptability and scalability. eBPF probes can be customized according to the requirements of different industrial control systems and can flexibly capture different types of system call events. At the same time, the LSTM model can continuously learn by updating training data, adapt to new abnormal behavior patterns and attack methods, and ensure the security and stability of the system.

[0115] In one embodiment of the present invention, the process behavior entropy value is obtained, the behavior entropy value is smoothed by using the moving average method, and the triggering of the fusing instruction is controlled based on the monitoring of the application layer and the kernel layer, including:

[0116] Obtain the process behavior entropy value and smooth the behavior entropy value by using the moving average method;

[0117]

[0118] Among them, H(t) represents the behavior entropy value at time t, and the state space U includes: system call types, file access modes (i.e., the ratio of read times to write times), network connection topologies, and the resource usage of the process;

[0119]

[0120] When trigger the security fuse;

[0121] When trigger the security fuse;

[0122] Input the real-time detected data into the trained model to determine whether there is an abnormality. If there is an abnormality, trigger the security fuse.

[0123] Divide the historical time into multiple time windows of equal length, and count the number of attacks in each window respectively; assign different weights to different time windows, with the weight of the window closer to the current time being larger, to reflect that the recent attack situation has a greater impact on the current risk assessment; different attacks have different severities, and a severity coefficient can be assigned to each attack and taken into account when calculating the attack frequency.

[0124]

[0125] Among them, n represents the number of divided time windows, t i represents the i-th time window, N i represents the number of attacks detected in the i-th time window, S ij represents the severity coefficient of the j-th attack in the i-th time window, w i represents the weight of the i-th time window, satisfying Δt i represents the duration of the i-th time window, where where A ∈ (0, 1) is the attenuation coefficient.

[0126] The working principle and effects of the above technical solution are as follows: effectively identify abnormal behaviors, conduct multi-dimensional analysis, calculate the behavior entropy value by comprehensively considering multiple dimensions such as system call types, file access patterns, network connection topologies, and process resource usage, which can comprehensively describe the behavior characteristics of processes, thereby more accurately detecting abnormal behaviors. For example, when the network connection topology of a process suddenly becomes complex (the Shannon entropy increases), and at the same time, the file access pattern also shows abnormal changes, the behavior entropy value will change accordingly, which helps to timely detect potential security threats; smooth processing and statistical judgment, the moving average method is used to smooth the behavior entropy value, reducing the influence of noise and short-term fluctuations, making the abnormal judgment based on the entropy value more reliable. Combining statistical methods (such as comparison with the mean and standard deviation) can effectively identify abnormal situations deviating from the normal behavior pattern, improving the accuracy of abnormal detection; dynamically adapt to the system state, considering the fusing conditions of multiple factors: the second fusing trigger condition comprehensively considers factors such as process-related thresholds, the change rate of the behavior entropy value, system load, and historical attack frequency. This enables the fusing mechanism to dynamically adjust the trigger threshold according to the real-time state and historical situation of the system. For example, when the system load is high, the fusing threshold is appropriately increased to avoid false triggering of fusing due to normal business fluctuations; when the historical attack frequency is low, the threshold is also appropriately increased to reduce unnecessary fusing operations; the adaptability of the training model, the trained model is used to make abnormal judgments on real-time data. The model can learn the normal and abnormal behavior patterns in different situations and can adapt to the changes in the system state as the data is continuously updated and learned, improving the detection ability for new types of abnormal behaviors; timely trigger the security fusing mechanism, when abnormal process behaviors are detected, relevant processes or connections are quickly cut off to prevent abnormal behaviors from causing further damage to the system, ensuring the security and stability of industrial hosts or systems. For example, when malware attempts to attack through abnormal system calls or network connections, the fusing can be triggered in a timely manner to prevent the spread of the attack and the expansion of the harm; through accurate abnormal detection and reasonable fusing mechanisms, the system failures and downtime caused by abnormal behaviors are reduced, improving the reliability and availability of the system. At the same time, the dynamic adaptation ability to the system state also enables the system to operate stably under different workloads and security environments. When calculating the historical average attack frequency, consider the time decay factor. In the security field, recent attack events often reflect the current security risks faced by the system more than long-term attack events. As time goes by, the security status of the system, protection measures, and the means of attackers may all change, and the reference value of long-term attack events for current risk assessment gradually decreases. Therefore, by setting exponentially decaying weights, a larger weight is given to the closer time window, and a smaller weight is given to the farther time window, which can more reasonably reflect the importance of attack data in different time windows; different attack events cause different degrees of harm to the system.For example, a data leakage attack may be much more serious than a simple network scanning attack. Therefore, when calculating the attack frequency, introducing an attack severity coefficient can more comprehensively measure the impact of the attack on the system. By quantifying the severity of each attack event and taking it into account in the calculation, it is possible to avoid evaluating risks solely based on the number of attacks and ignoring the actual harm of the attacks; Different time windows may have different durations. For example, in some cases, in order to analyze recent attack situations in more detail, the recent time may be divided into shorter time windows, while the long-term time may use longer time windows. Therefore, when calculating the historical average attack frequency, it is necessary to consider the duration of each time window. By performing a weighted sum of the durations, it is ensured that the contributions of time windows of different durations to the final result are reasonable. More accurately reflecting the current security risk, due to considering the time decay factor, this formula can focus more on recent attack events, thus more accurately reflecting the security risk faced by the system currently. This helps security managers promptly detect changes in the security situation and take corresponding protection measures. For example, if the number of recent attacks increases or the attack severity improves, the historical average attack frequency calculated by the formula will increase correspondingly, alerting managers to strengthen security protection; Combining the attack severity coefficient, the formula can comprehensively consider the number and severity of attacks, providing a more comprehensive attack frequency evaluation indicator. This enables security assessment to not only focus on the number of attacks but also consider the actual harm caused by the attacks to the system, helping to more accurately evaluate the security status of the system. For example, if the number of attacks is small within a period of time but the severity of each attack is high, the calculated attack frequency will also be high, indicating that the system faces a greater security risk; Considering the differences in the durations of time windows makes the formula more flexible and adaptable. The historical time can be divided into different numbers and durations of time windows according to actual needs, and the formula can reasonably handle these different division methods to ensure the accuracy and reliability of the calculation results. For example, when conducting short-term security analysis, the time window can be divided more finely; when conducting long-term trend analysis, longer time windows can be used, and the formula can effectively calculate the corresponding historical average attack frequency; The historical average attack frequency calculated by this formula can be used as an important security indicator to provide a decision-making basis for security decision management.

[0127] An embodiment of the present invention, an industrial host control system based on artificial intelligence, the system includes:

[0128] An application layer monitoring module, used to set whitelist processes. When a process accesses the industrial host, the industrial host monitors whether there are abnormalities in the process accessing the industrial control host through the application layer data before and during the process operation;

[0129] The kernel layer monitoring module then collects real-time syscall stream data through the eBPF behavior collector in the kernel layer of the industrial control host, and analyzes the syscall stream data through the AI decision engine to determine whether there is an abnormality;

[0130] The fuse trigger instruction module is used to obtain the process behavior entropy value, smooth the behavior entropy value by using the moving average method, and trigger the fuse instruction based on the monitoring of the application layer and the kernel layer.

[0131] In one embodiment of the present invention, the application layer monitoring module includes:

[0132] The white list process information storage module is used to sort out all the processes running on the industrial host, screen out the white list processes, and store the white list process information in a database. The white list process information includes: process name, process ID, executable file path, file hash value, digital signature, startup parameters, running time range, required system permissions, user identity;

[0133] The static credibility acquisition module is used to, when a process accesses the industrial host, the industrial host first compares the SHA3-512 hash values of the first preset number of files in the process with the golden hash values in the database and compares the metadata of the second preset number of files with the standard metadata in the database, that is, obtains the static credibility of the process accessing the industrial host through the static credibility model;

[0134] The dynamic credibility acquisition module is used to, on the basis of considering the natural attenuation of trust over time, combine the comparison result between the system call sequence within a preset time and the set of normal system call sequences generated by historical behaviors, and obtain the dynamic credibility of the security of the process runtime behavior through the dynamic credibility model;

[0135] The environment evaluation module is used to perform weighted summation on the ratios of the current values of each hardware-level and network-level monitoring index to the baseline value, and then through the activation function mapping, obtain the quantitative evaluation value of the security of the environment in which the process runs through the runtime environment evaluation model. The hardware-level and network-level monitoring indexes include: CPU instruction cycle abnormality rate, memory access entropy, DMA request frequency, network traffic rate, network traffic distribution, and system log information.

[0136] In one embodiment of the present invention, the kernel layer monitoring module further includes:

[0137] The static credibility model storage module. Specifically, the static credibility model is:

[0138]

[0139] Where T staticRepresents the static credibility, n represents the number of files participating in the static credibility hash check, SHA3 - 51h(f i ) represents the SHA3 - 512 hash value of the i-th file, I(·) represents the indicator function, if the condition in the parentheses holds, the function value is 1, otherwise it is 0, DB g (f i ) represents the standard hash value of the i-th file, DB m (f j ) represents the standard metadata of the j-th file, m represents the number of files participating in the static credibility metadata check, MData(f j ) represents the metadata of the j-th file, the metadata includes file size, file creation time, file modification time, file access time, file owner, file access permissions, software version number, compilation time, signature status, signature certificate information, basic attributes of dependent files, permission information, version information, digital signature information and associated information, DB m (f j ) represents the standard metadata of the j-th file, stored in the database for comparison,

[0140] Store the dynamic credibility model module, the dynamic credibility model is:

[0141]

[0142] Among them, T dynamic (t) represents the dynamic credibility at time t, λ(t) represents the trust decay coefficient, t represents time, L represents the number of system call sequences, b k represents the k-th system call sequence, B normal represents the set of normal system call sequences generated from historical behaviors;

[0143] Store the environment assessment model module, the environment assessment model is

[0144]

[0145] Among them, T c (t) represents the environmental security assessment value, σ(·) represents the activation function, w j represents the weight of the j-th monitoring index, represents the current value of the j-th monitoring index, represents the baseline value of the j-th monitoring index.

[0146] In an embodiment of the present invention, the abnormal judgment module includes:

[0147] The keyword field data capture module is used to deploy customized eBPF probes at the kernel layer of the industrial control host, capture all syscall events, and obtain the keyword field data in all syscall events. The keyword field data includes: process PID, syscall number, timestamp, and specific parameters passed during the execution of the syscall.

[0148] The preprocessing module is used to clean and format the keyword field data. The data cleaning and formatting include: filtering invalid calls, i.e., filtering syscall = 0 and abnormal PIDs, parameter desensitization, i.e., hashing sensitive parameters such as file paths and network addresses, and time series alignment, i.e., using Lamport logical clocks to ensure the order of cross-core events.

[0149] The feature extraction module is used to extract features from the cleaned and formatted keyword field data. The features include basic statistical features and context-related features. The basic statistical features include: call frequency, call type entropy, parameter diversity, and IO operation ratio. The context-related features include: process tree blood relationship and resource access chain.

[0150] The division and training module is used to label the extracted feature data as normal and abnormal, divide the labeled feature data into a training set, a validation set, and a test set, construct an initial LSTM model, and train the initial LSTM model based on the training set.

[0151] In one embodiment of the present invention, the trigger fuse instruction module includes:

[0152] The process behavior entropy value acquisition module is used to acquire the process behavior entropy value and smooth the behavior entropy value using the moving average method.

[0153]

[0154] Among them, H(t) represents the behavior entropy value at time t, and the state space U includes: system call type, file access mode, i.e., the ratio of read times to write times, network connection topology, and resource usage of the process.

[0155]

[0156] The judgment and trigger module is used to trigger a safety fuse when

[0157] When trigger a safety fuse;

[0158] Input the real-time detected data into the trained model to judge whether there is an abnormality. If there is an abnormality, trigger a safety fuse.

[0159] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An industrial host control method based on artificial intelligence, characterized in that, The method includes: Setting a whitelist process. When a process accesses the industrial host, the industrial host monitors the process through a static credibility model, a dynamic credibility model, and a running environment evaluation model; Collecting real-time syscall stream data through an eBPF behavior collector and analyzing the syscall stream data through an AI decision engine; Obtaining a process behavior entropy value, smoothing the behavior entropy value using the moving average method, and controlling the triggering of a fusing instruction based on the monitoring of the model and the analysis of the AI decision engine.

2. The industrial host control method based on artificial intelligence according to claim 1, wherein, Setting a whitelist process. When a process accesses the industrial host, the industrial host monitors the process through a static credibility model, a dynamic credibility model, and a running environment evaluation model, including: Sorting out all processes running on the industrial host, screening out the whitelist processes, storing the whitelist process information in a database, and the whitelist process information includes: process name, process ID, executable file path, file hash value, digital signature, startup parameters, running time range, required system permissions, user identity; When a process accesses the industrial host, the industrial host first compares the SHA3-512 hash value of the first preset number of files in the process with the golden hash value in the database and compares the metadata of the second preset number of files with the standard metadata in the database, that is, obtaining the static credibility of the process accessing the industrial host through the static credibility model; On the basis of considering the natural decay of trust over time, combining the comparison result between the system call sequence within a preset time and the set of normal system call sequences generated by historical behaviors, obtaining the dynamic credibility of the security of the process runtime behavior through the dynamic credibility model; Performing a weighted sum of the ratios of the current values of each hardware-level and network-level monitoring index to the baseline value, and then through activation function mapping, obtaining a quantitative evaluation value of the security of the environment where the process runs through the running environment evaluation model. The hardware-level and network-level monitoring indexes include: CPU instruction cycle anomaly rate, memory access entropy, DMA request frequency, network traffic rate, network traffic distribution, and system log information.

3. The method for controlling an industrial host based on artificial intelligence according to claim 2, characterized in that Specifically, the static credibility model is: Among them, T static represents the static credibility, n represents the number of files participating in the static credibility hash check, SHA3-51h(f i ) represents the SHA3-512 hash value of the i-th file, I(·) represents the indicator function, if the condition in the parentheses holds, the function value is 1, otherwise it is 0, DB g (f i ) represents the standard hash value of the i-th file, DB m (f j ) represents the standard metadata of the j-th file, m represents the number of files participating in the static credibility metadata check, MData(f j ) represents the metadata of the j-th file, the metadata includes file size, file creation time, file modification time, file access time, file owner, file access permission, software version number, compilation time, signature status, signature certificate information, basic attributes of dependent files, permission information, version information, digital signature information and associated information, DB m (f j ) represents the standard metadata of the j-th file, stored in the database for comparison, The dynamic credibility model is: Among them, T dynamic (t) represents the dynamic credibility at time t, λ(t) represents the trust decay coefficient, t represents time, L represents the number of system call sequences, and b k represents the k-th system call sequence, and B normal represents the set of normal system call sequences generated from historical behaviors; The environment evaluation model is Among them, T c (t) represents the environmental safety assessment value, σ(·) represents the activation function, w j represents the weight of the j-th monitoring index, represents the current value of the j-th monitoring index, represents the baseline value of the j-th monitoring index.

4. The industrial host control method based on artificial intelligence according to claim 1, wherein Collecting real-time syscall stream data through an eBPF behavior collector and analyzing the syscall stream data through an AI decision engine, including: Deploying a customized eBPF probe in the kernel layer of the industrial control host, capturing all syscall events, and obtaining the key field data in all syscall events. The key field data includes: process PID, syscall number, timestamp, and specific parameters passed when the syscall is called and executed. Perform data cleaning and formatting on the keyword field data. The data cleaning and formatting include: invalid call filtering, i.e., filtering syscall = 0 and abnormal PID; parameter desensitization, i.e., performing hash processing on sensitive parameters such as file paths and network addresses; and time series alignment, i.e., using Lamport logical clocks to ensure the orderliness of cross-core events. Extract features from the keyword field data after cleaning and formatting. The features include basic statistical features and context-related features. The basic statistical features include: call frequency, call type entropy, parameter diversity, and IO operation ratio. The context-related features include: process tree blood relationship and resource access chain. Label the extracted feature data as normal and abnormal, divide the labeled feature data into a training set, a validation set, and a test set, construct an initial LSTM model, and train the initial LSTM model based on the training set.

5. The industrial host control method based on artificial intelligence according to claim 1, characterized in that, Obtain the process behavior entropy value, perform smoothing processing on the behavior entropy value using the moving average method, and control the triggering of the fuse instruction based on the monitoring of the model and the analysis of the AI decision engine, including: Obtain the process behavior entropy value and perform smoothing processing on the behavior entropy value using the moving average method. Among them, H(t) represents the behavior entropy value at time t, the state space U includes: system call type, file access mode, i.e., the ratio of read times to write times, network connection topology, and resource usage of the process, and p(s|t) represents the probability that the system call type s appears within the time window t. Among them, H smoothed (t) represents the behavioral entropy at time t, and M represents the smoothing window size; When , a safety fuse is triggered, where μ H represents the historical normal entropy mean, and σ H represents the historical normal entropy standard deviation, and Δt represents the time window length; When occurs, a safety fuse is triggered, where η represents the entropy change sensitivity coefficient, represents the time derivative of the behavioral entropy, τ d represents the preset dynamic threshold reference value, and f(A) represents the historical average attack frequency; T(t) = α(t)·T static + β(t)·T dynamic (t) + γ(t)·T c (t) Among them, α(t), β(t), and γ(t) represent weight coefficients. Input the real-time detected data into the trained model to determine whether there is an abnormality. If there is an abnormality, trigger a safety fuse.

6. An industrial host control system based on artificial intelligence, characterized in that, The system includes: An application layer monitoring module for setting whitelist processes. When a process accesses the industrial host, the industrial host monitors the process through a static credibility model, a dynamic credibility model, and a running environment evaluation model. A kernel layer monitoring module that collects real-time syscall stream data through an eBPF behavior collector and analyzes the syscall stream data through an AI decision engine. A trigger fuse instruction module that obtains the process behavior entropy value, performs smoothing processing on the behavior entropy value using the moving average method, and controls the triggering of the fuse instruction based on the monitoring of the model and the analysis of the AI decision engine.

7. The industrial host control system based on artificial intelligence according to claim 6, characterized in that, The application layer monitoring module includes: A module for storing whitelist process information, which is used to sort out all processes running on the industrial host, filter out whitelist processes, and store the whitelist process information in a database. The whitelist process information includes: process name, process ID, executable file path, file hash value, digital signature, startup parameters, running time range, required system permissions, and user identity. The static credibility acquisition module is used to, when a process accesses the industrial host, the industrial host first obtains the static credibility of the process accessing the industrial host through a static credibility model by comparing the SHA3-512 hash values of the first preset number of files in the process with the golden hash values in the database and comparing the metadata of the second preset number of files with the standard metadata in the database; The dynamic credibility acquisition module is used to, on the basis of considering the natural attenuation of trust over time, combine the comparison results between the system call sequence within a preset time and the set of normal system call sequences generated from historical behaviors, and obtain the dynamic credibility of the security of the process runtime behavior through a dynamic credibility model; The environment assessment module is used to perform weighted summation on the ratios of the current values of each hardware-level and network-level monitoring index to the baseline values, and then, after being mapped by an activation function, obtain a quantitative assessment value of the security of the environment in which the process runs through an operating environment assessment model. The hardware-level and network-level monitoring indexes include: CPU instruction cycle anomaly rate, memory access entropy, DMA request frequency, network traffic rate, network traffic distribution, and system log information.

8. The industrial host control system based on artificial intelligence according to claim 7, characterized in that, The kernel layer monitoring module further includes: The static credibility model storage module. Specifically, the static credibility model is: Among them, T static represents the static credibility, n represents the number of files participating in the static credibility hash check, SHA3-51h(f i ) represents the SHA3-512 hash value of the i-th file, I(·) represents the indicator function, if the condition in the parentheses holds, the function value is 1, otherwise it is 0, DB g (f i ) represents the standard hash value of the i-th file, DB m (f j ) represents the standard metadata of the j-th file, m represents the number of files participating in the static credibility metadata check, MData(f j ) represents the metadata of the j-th file, and the metadata includes file size, file creation time, file modification time, file access time, file owner, file access permission, software version number, compilation time, signature status, signature certificate information, basic attributes of dependent files, permission information, version information, digital signature information and associated information, DB m (f j ) represents the standard metadata of the j-th file, stored in the database for comparison, The dynamic credibility model storage module. The dynamic credibility model is: Among them, T dynamic (t) represents the dynamic credibility at time t, λ(t) represents the trust decay coefficient, t represents time, L represents the number of system call sequences, b k represents the k-th system call sequence, B normal represents the set of normal system call sequences generated from historical behaviors; The environment assessment model storage module. The environment assessment model is Among them, T c (t) represents the environmental safety assessment value, σ(·) represents the activation function, and w j represents the weight of the j-th monitoring index, represents the current value of the j-th monitoring index, represents the baseline value of the j-th monitoring index.

9. The industrial host control system based on artificial intelligence according to claim 6, characterized in that, The anomaly judgment module includes: Deploy customized eBPF probes in the industrial control host kernel layer to capture all syscall events and obtain the key field data in all syscall events. The key field data includes: process PID, syscall number, timestamp, and specific parameters passed during the execution of the syscall; Perform data cleaning and formatting on the key field data. The data cleaning and formatting include: invalid call filtering, i.e., filtering syscall = 0 and abnormal PID, parameter desensitization, i.e., performing hash processing on sensitive parameters such as file paths and network addresses, and timing alignment, i.e., using Lamport logical clocks to ensure the orderliness of cross-core events; Extract the features in the key field data after cleaning and formatting. The features include basic statistical features and context correlation features. The basic statistical features include: call frequency, call type entropy, parameter diversity, and IO operation ratio. The context correlation features include: process tree blood relationship and resource access chain; Label the extracted feature data as normal and abnormal, divide the labeled feature data into a training set, a validation set, and a test set, construct an initial LSTM model, and train the initial LSTM model based on the training set.

10. The industrial host control system based on artificial intelligence according to claim 6, characterized in that, The trigger fuse instruction module includes: The process behavior entropy value acquisition module is used to obtain the process behavior entropy value and perform smoothing processing on the behavior entropy value using the moving average method; Among them, H(t) represents the behavior entropy value at time t, and the state space U includes: system call type, file access mode, i.e., the ratio of read times to write times, network connection topology, and resource usage of the process; Determination of whether to trigger the module, which is used to trigger the safety fuse when occurs; When is triggered, a safety fuse blows; Input the real-time detected data into the trained model to determine whether there is an anomaly. If there is an anomaly, trigger a safety fuse.

Citation Information

Patent Citations

  • Computer security management system and method based on artificial intelligence

    CN117574361A

  • Container behavior monitoring method and system based on eBPF technology

    CN117763545A

  • Network security intelligent protection method and system based on endogenous security mechanism

    CN118972157A

  • Intrusion detection method, apparatus and system, electronic device and computer-readable medium

    US20240283805A1

  • Industrial control system communication network anomaly classification method

    WO2022057260A1

Cited By

  • Host security early warning method, device and equipment in data center and readable medium

    CN121543088A

  • Host security alert methods, devices, equipment, and readable media in data centers

    CN121543088B

  • Network control method and system based on trusted state of application process, and electronic equipment

    CN121842077A

  • A method for monitoring abnormal behavior of enterprise terminals based on eBPF dynamic probes

    CN122571586A