Process behavior log generation method, vehicle and computer readable storage medium

By generating process behavior logs through eBPF programs and random forest models, the problems of high storage pressure and poor detection timeliness in existing technologies are solved, and efficient process behavior monitoring and traceability analysis are achieved.

CN120803860APending Publication Date: 2025-10-17ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510747641.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In existing technologies, monitoring solutions rely on post-log analysis, which leads to high storage pressure and easy deletion of log data. Traditional logs do not contain advanced behavioral semantics, resulting in high system overhead and poor detection timeliness.

Method used

The eBPF program is used to monitor process behavior, generate target feature vectors and input them into the random forest model for classification, generate process behavior logs, reduce storage space and improve monitoring performance.

Benefits of technology

It achieves high real-time and high-accuracy security monitoring, and improves log quality and traceability analysis capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803860A_ABST
    Figure CN120803860A_ABST
Patent Text Reader

Abstract

The invention provides a process behavior log generation method and a computer readable storage medium, the method comprises: monitoring a process behavior of a target file operation, and obtaining target data, the target data comprising process behavior data; performing statistics on the target data to generate statistical data; vectorizing the statistical data to generate a target feature vector; and inputting the target feature vector into a random forest model for classification so as to identify the process behavior, and generating a process behavior log. According to the method, statistics is performed on data such as process behaviors related to target file operation, the obtained statistical data is vectorized and then input into the random forest model for classification, so that behavior semantics of the process are identified, high-real-time and high-accuracy safety monitoring is realized, and log quality and traceability analysis capability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer security monitoring, and in particular to a process behavior log generation method, a vehicle and a computer readable storage medium. BACKGROUND

[0002] With the continuous improvement of system complexity, audit logs, data tracing and real-time monitoring technology have become an important means to ensure system security.

[0003] In the prior art, most monitoring solutions rely on post-log analysis, but due to limited storage space, log data is often deleted before analysis, resulting in timeliness and data integrity problems in detecting abnormal behavior. The prior art has made certain achievements in log abstraction, behavior analysis and data tracing, but at the same time, there are the following shortcomings: most methods rely on complete audit logs, which need to bear immeasurable storage pressure; traditional logs do not contain advanced behavior semantics, which need to spend a lot of time to extract the semantics of the logs; a large number of logs need to be generated, which brings huge system overhead, resulting in a delay of disk-intensive services (such as MySQL) increasing by tens of times at most.

[0004] Therefore, there is an urgent need for a process behavior log generation method, a vehicle and a computer readable storage medium to solve the above problems. SUMMARY

[0005] The technical problem solved by the present application is to provide a process behavior log generation method, a vehicle and a computer readable storage medium, which can output the advanced semantics of the process as a log, reducing the storage space of the log, reducing the overhead of the system and improving the monitoring performance of the system.

[0006] The technical problem solved by the present application is solved by the following technical solution: A process behavior log generation method, comprising: monitoring the process behavior of a target file operation and obtaining target data, the target data comprising process behavior data; performing statistics on the target data to generate statistical data; vectorizing the statistical data to generate a target feature vector; inputting the target feature vector into a random forest model for classification to identify the process behavior and generate a process behavior log.

[0007] In a preferred embodiment of the present application, the step of monitoring the process behavior of a target file operation and obtaining target data comprises: defining an eBPF program through a Python interface of a BCC tool, and hooking the eBPF program to a target system call entry related to the target file operation to collect target call data; filtering the target call data based on a preset filtering condition to obtain the target data.

[0008] In the preferred embodiment of the present application, the preset filtering condition includes process name and command line parameters.

[0009] In the preferred embodiment of the present application, the step of generating statistical data by counting the target data includes aggregating and counting the collected target data through an eBPF Map structure to generate statistical data of the target data, and the statistical data includes read-write operation count of processes, system call interval, and cumulative statistics of error return code.

[0010] In the preferred embodiment of the present application, before the step of generating a target feature vector by vectorizing the statistical data, the method includes cleaning and formatting the statistical data of all processes, and respectively counting read-write ratio, system call interval, and error call frequency of each process; and generating a target feature vector by vectorizing the read-write ratio, system call interval, and error call frequency of each process.

[0011] In the preferred embodiment of the present application, the dimension of the target feature vector at least includes one of the following: process read operation count, process write operation count, system call time interval mean and variance, and error return code ratio.

[0012] In the preferred embodiment of the present application, the step of inputting the target feature vector into a random forest model for classification to identify the process behavior and generate a process behavior log includes inputting the target feature vector into a random forest model to classify the process behavior and identify the process behavior; generating the process behavior log based on the process behavior and the statistical data, and storing the process behavior log in a classified manner.

[0013] In the preferred embodiment of the present application, after the step of storing the identified process behavior as a process behavior log, the method includes obtaining the process behavior corresponding to each process and the target file operated based on the association between the process behavior log and the context PID.

[0014] In the preferred embodiment of the present application, the training method of the random forest model includes obtaining basic training data of the random forest model; automatically labeling the basic training data according to a preset process-behavior corresponding rule, associating the basic training data with process behavior, and constructing a preliminary feature vector through the basic training data; inputting the preliminary feature vector into an initial random forest model for training to generate the random forest model.

[0015] A vehicle applying the steps of the process behavior log generation method of any one of the above.

[0016] A computer readable storage medium, a computer program is stored on the computer readable storage medium, the computer program is executed by the processor to realize the steps of the process behavior log generation method in any one of the above.

[0017] The technical effects achieved by the technical scheme are as follows: the system call information / process behavior data related to the operation of the target file is counted, the statistical data is vectorized, the obtained vector is input into a random forest model for classification, and the behavior semantics of the process is identified, thereby realizing high real-time and high-accuracy security monitoring, improving log quality and traceability analysis capability.

[0018] The above description is only a summary of the technical scheme of the present application, in order to more clearly understand the technical means of the present application, the content of the specification can be implemented, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the following preferred embodiments are described in detail, and the accompanying drawings are described in detail. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 A step flow chart of a process behavior log generation method according to the present application.

[0020] Figure 2 A system architecture diagram of a process behavior log generation method according to the present application.

[0021] Figure 3 A system architecture diagram of a model training method according to an embodiment of the present application.

[0022] Figure 4 A detailed architecture diagram of a process behavior log generation method according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined purposes, the embodiments of the present application are described in detail below, and the examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below are only a part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the embodiments of the present application. Through the description of the specific embodiments, the technical means and effects taken by the present application to achieve the predetermined purposes can be more deeply and specifically understood, and the accompanying drawings are only provided for reference and description, and are not used to limit the present application.

[0024] It should be noted that the present application mainly monitors the process behavior of Linux file operation and outputs the process behavior log.

[0025] Please refer to Figure 1 , Figure 1 A step flow chart of a process behavior log generation method shown by the present application.

[0026] As Figure 1 shown, the process behavior log generation method provided by an embodiment of the present application includes the following steps: S11: monitoring the process behavior of target file operation and obtaining target data.

[0027] Among them, the target data includes process behavior data and file operation data and other information. The target file can be any file stored in the system.

[0028] Optionally, the step of monitoring the process behavior of target file operation and obtaining target data includes: defining an eBPF program and attaching the eBPF program to a target system call entry related to the target file operation to collect target call data; filtering the target call data based on a preset filtering condition to obtain the target data.

[0029] Optionally, the preset filtering condition includes process name and command line parameter.

[0030] Figure 2 A system architecture diagram of a process behavior log generation method shown by the present application. As Figure 2 shown, the system architecture of the present application is divided into kernel state (User Space) and user state (Kernel Space).

[0031] In the kernel state, the process behavior related to the target file operation (or the related process behavior when operating the target file) in the system is monitored by defining an eBPF program (extended Berkeley Packet Filter, a revolutionary kernel technology that allows dynamic extension of kernel functions without modifying kernel source code or loading kernel modules) through a BCC tool (BPF Compiler Collection, an open source tool set based on eBPF (extended Berkeley Packet Filter) technology, widely used in Linux system performance analysis, dynamic tracking and kernel programming.) The tool can inject C language code in the kernel state into the kernel space. The kernel state program will monitor the system calls related to the target file, such as open, openat, read, write, close, unlink, renameat2, etc.

[0032] In particular, the present embodiment is generally used on Linux 5.4 and above versions of the operating system, and requires that the kernel has enabled eBPF support. First, install the BCC toolset on the target system (i.e. the above operating system) to ensure that the kernel symbols are available for debugging and monitoring. It is recommended to use Ubuntu 20.04 or CentOS 8 or later when deploying, and attention should also be paid to the adjustment of kernel parameters when configuring the environment to support large-scale event capture and transmission.

[0033] S12: Statistics of target data, generating statistical data.

[0034] In particular, the statistical data includes: read-write operation count of the process, system call interval, and cumulative statistics of error return code.

[0035] In order to reduce the storage space and improve the detection performance of the system, after using the eBPF program to monitor the process behavior related to the target file and collecting the target data in real time, the collected target data needs to be further statistically analyzed to obtain statistical data, which includes the high-level behavior semantics of the process (i.e. process behavior, such as read-write operation count of the process as described above).

[0036] Optionally, the step of generating statistical data by statistically analyzing the target data includes: aggregating and statistically analyzing the collected target data through an eBPF Map (hash table) structure to generate statistical data of the target data.

[0037] In the present embodiment, it is necessary to statistically analyze the file operation data and related process behavior data, i.e. to statistically analyze the target file operation and related process behavior.

[0038] In particular, the statistical data includes: read-write operation count of the process, system call interval, and cumulative statistics of error return code.

[0039] Illustratively, when the target file is operated, the related system call is executed, and the monitoring program (i.e. eBPF program) will statistically analyze the specified system call information (filtering conditions are set in the eBPF program, so that the system call information meeting the filtering conditions is monitored), such as: parameters, return values, pid and timestamps, etc. In addition to obtaining statistical information of process behavior, the kernel program also statistically analyzes the file operation information, such as: file name, read-write byte number, etc.

[0040] Exemplarily, when the target file is closed, the statistics of the monitored file will be transmitted from the kernel mode to the user mode through the Ring Buffer; however, if the relevant process is not closed / stopped at this time, the statistics of the process behavior has not been completed, and thus the behavior of the process cannot be judged and can only be recorded as the process name.

[0041] Exemplarily, when the process related to the target file is closed, the statistics of the process behavior is completed, and the statistics are transmitted to the user mode through the Ring Buffer for processing. The contents of the statistical process behavior data mainly include: the number of system calls of the process, the number of system call errors, the running time, the number of read / write system calls, and the like.

[0042] Exemplarily, in the kernel mode, a monitoring program (i.e., the eBPF program described above) is written and loaded by using the eBPF technology, and the program is attached to the key system calls by using the tracepoint technology. For example, the read, write, open, close, and the like related to the file operation. The monitoring program mainly includes the following steps: System call interception: the eBPF program is defined by using the Python interface of BCC, and is attached to the entrance of the target system call to capture the call parameters, return values, timestamps, and the like to obtain the target call data. For example, the call start time can be recorded at the beginning of each system call, and the call delay can be calculated after the end.

[0043] Data filtering and sampling: in order to reduce the system overhead, conditional filtering is adopted, and only the target file operation commands of the non-graphical terminal are monitored. The filtering conditions are set in the eBPF program, such as the process name, command line parameters, and the like, to filter out the relevant processes that need to be monitored.

[0044] Event data aggregation: the eBPF Map (hash table) structure is used to aggregate and count the collected target data.

[0045] Batch data transmission: the transmission strategy of the target data is "delayed transmission". That is, after the process is executed, all the statistical data of the target data in the cache are transmitted to the user mode in batches through the ring buffer, so as to reduce the performance loss caused by frequent cross-state interaction.

[0046] Based on the above manner, the application will not directly transmit the detected system call condition to the user state for logging as a log, but will record the statistical data of the target file and / or process behavior as a log, thereby greatly reducing the number of generated logs and the performance loss caused by cross-state data replication (kernel state sending to user state).

[0047] S13: Vectorize the statistical data to generate a target feature vector.

[0048] Specifically, the dimensions of the target feature vector at least include one of the following: process read operation count, process write operation count, system call time interval mean and variance, and error return code ratio.

[0049] Optionally, before the step of vectorizing the statistical data to generate a target feature vector, it includes: cleaning and formatting the statistical data of all processes, and respectively counting the read-write ratio, system call interval and error call frequency of each process; according to the read-write ratio, system call interval and error call frequency of each process, vectorize to generate a target feature vector.

[0050] As shown in Figure 2 A data processing module is developed in the user state, which is written in Python language, and its main tasks include: Data reception: batch receive the statistical data of the target data obtained by monitoring from the kernel state through the ring buffer.

[0051] Preprocessing: clean and format the received statistical data, and then count the read-write ratio, system call interval and error call frequency of each process in the received statistical data (the statistical data can include process behavior data of multiple processes related to target file operation).

[0052] Feature construction: convert the processed statistical data into a multi-dimensional feature vector, and each dimension corresponds to one of the above process behavior indicators.

[0053] Exemplarily, the pre-designed target feature vector includes four dimensions. For example, dimension 1: process read operation count; dimension 2: process write operation count; dimension 3: system call time interval mean and variance; and dimension 4: error return code ratio, etc.

[0054] It should be noted that an optimized data transmission mechanism is designed in the present application: the target data collected in the kernel state is transmitted in batches to the user state through RingBuffer after the target file is closed and / or the process execution is completed. Through this data transmission mechanism, the frequent interaction between the kernel state and the user state is greatly reduced, thereby reducing the performance overhead and delay risk caused by data transmission; at the same time, through the centralized transmission mode, the system load in the real-time transmission process is also reduced.

[0055] It should be noted that the user state of the BCC tool used in the present application will be written in Python language, so the user state program has very high flexibility, which can complete machine learning and vectorization through third-party libraries.

[0056] Optionally, when the user state program monitors the file statistical data of the target file operation, a log will be generated and stored in a log file, which will be associated with the subsequent generated process behavior log.

[0057] Illustratively, when the user state program monitors the statistical data of the process behavior, the statistical data will be vectorized, that is, the statistical data will be calculated and stored using the data structure of numpy Array (in NumPy, Array is a data structure for storing a collection of elements, which are usually numerical types. NumPy arrays are efficient and powerful, and they support multi-dimensional data, which can perform fast mathematical and matrix operations). The key indicators for describing the process behavior include: the ratio of the number of read and write system calls (note that the zero division problem needs to be considered), the system call error frequency, the system call frequency, etc.

[0058] It should be noted that in the present embodiment, the kernel state monitoring module and the user state data processing module interact with each other through a standardized interface. The user state mainly consists of the following parts: data receiver: through the interface provided by BPF, the data packets (statistical data of target data) transmitted in batches by the kernel state are received. Data parser: parse the data packet content to construct the target feature vector in the predetermined format. Semantic detector: call the random forest model to classify and judge the feature vector. The overall architecture is as shown in Figure 2 As shown in the figure, the modules are loosely coupled through data queues to ensure that each link can still operate stably under high load.

[0059] S14: input the target feature vector into the random forest model for classification to identify the process behavior and generate the process behavior log.

[0060] Optionally, the step of inputting the target feature vector into the random forest model for classification to identify the process behavior and generate a process behavior log includes: inputting the target feature vector into the random forest model to classify the process behavior and identify the process behavior; and generating the process behavior log based on the process behavior and the statistical data and storing the process behavior log in a classified manner.

[0061] Optionally, the step of storing the identified process behavior in a classified manner as a process behavior log is followed by: obtaining the process behavior corresponding to each process and the target file operated based on the association between the process behavior log and the context PID.

[0062] By way of example, after the vectorization of the statistical data of the process behavior is completed, the target feature vector obtained is input into the random forest model for classification, so that the behavior semantics of the process can be identified. The user state handler stores the process behavior obtained in a classified manner as a process behavior log, and the user can determine the process behavior corresponding to each process and the target file operated based on the association between the context PID (process name) and the process behavior log and the target file log.

[0063] Specifically, in order to realize real-time semantic detection of process behavior, a random forest model is integrated in the user state. The specific implementation steps are as follows: pre-training data set construction: the common file operation commands (about 20) are manually annotated and classified in advance to form the basic training data of the model. A preliminary feature vector is constructed based on these basic training data. Model parameter setting: the number of trees of the random forest model can be set to 100, and the maximum depth can be limited to 10 layers to achieve high classification accuracy and real-time response requirements. Real-time prediction: when the random forest model is actually running, the feature vector is input into the pre-trained random forest model, and the model outputs the corresponding behavior classification result. The user state program records the behavior semantics obtained by classification as detailed logs for subsequent traceability analysis when needed.

[0064] Optionally, the training method of the random forest model includes: obtaining basic training data of the random forest model; automatically annotating the basic training data according to a preset process and behavior corresponding rule, associating the basic training data with process behavior, and constructing a preliminary feature vector based on the basic training data; inputting the preliminary feature vector into an initial random forest model for training to generate the random forest model.

[0065] It should be noted that for real-time monitoring scenarios, while deep learning models and emerging large language models (LLMs) offer high detection accuracy, they are computationally complex, require high computing power, and have long response times, making them unsuitable for real-time scenarios. Therefore, this application uses the random forest model from traditional machine learning as the primary process behavior recognition algorithm.

[0066] Random forests have the following advantages: High efficiency: They can quickly classify when processing multi-dimensional feature vectors and are suitable for real-time data processing; Prevent overfitting: Random forests integrate multiple decision trees to avoid overfitting of a single decision tree due to insufficient data samples or noise interference, so they have higher accuracy; Low resource consumption: Compared with deep learning models, random forest algorithms have lower requirements for computing resources and can achieve efficient detection while ensuring real-time performance.

[0067] The training data of the random forest model also needs to be collected by the corresponding model training program. The system architecture diagram of the model training method is as follows: Figure 3 As shown, it is also divided into two parts: kernel mode and user mode.

[0068] The kernel-mode program receives user-mode information and maintains a process filtering configuration. This filtering removes unwanted process behavior information from the random forest model training. In this implementation, we focus on the 20 most commonly used file manipulation commands in Linux, such as cat, cp, mv, rm, unzip, zip, and gzip.

[0069] Similar to the aforementioned implementation of the process behavior log generation method, the kernel-mode program used to collect model training data also needs to statistically store process behavior information. Only when the process ends is the collected process behavior data statistically transferred to user mode via the RingBuffer for analysis. Unlike the aforementioned process behavior log generation method, the model training program used to collect training data does not need to pay attention to the target file's statistical information. That is, the kernel-mode program only collects process behavior statistics for training the random forest model used in machine learning.

[0070] The user-mode program automatically labels the training data according to a predefined set of process-action mapping rules. For example, if the kernel-mode data is {.pid=21536, .comm=”zip”, …}, and the command-action mapping rules are {“zip”: “compress”, “gzip”: “compress”, “cat”: “view”, …}, then the label of the current data item is “compress”, and the training data is the vector [21536, …, “compress”] (the last dimension of the vector is the label of the data item). The predefined process-action mapping rules described above are manually written by the user.

[0071] The user-mode program converts the kernel-mode data into vectors for model training and then feeds them into the random forest model for training. After training, the user-mode program saves the trained random forest model in binary format, which can be loaded directly using the Python machine learning model library.

[0072] Once the initial random forest model is trained, it will generate process behavior characteristics related to file operations. The machine learning algorithm will classify all processes in the system based on the trained process behavior characteristics (i.e., the random forest model) and determine the process behavior corresponding to the statistical data of the process behavior data input into the random forest model.

[0073] Figure 4 The following is a detailed architecture diagram of a process behavior log generation method according to an embodiment of the present invention. Figure 4 As shown in the architecture diagram, Kernel Space: Utilizing eBPF technology to capture system call information in kernel space, and combined with BCC tools for data filtering and aggregation, system call-related data (such as call parameters, timestamps, and error codes) is cached in the kernel. This data is aggregated through structures such as the Ring Buffer to generate statistical data for subsequent batch transmission. User Space: The user space receives statistical data of target data transmitted from kernel space in batches through its data receiving module. It performs data preprocessing and constructs feature vectors for the received data. This data is then used to implement real-time monitoring of process behavior semantics using a random forest model (which learns and updates based on the data during use). The process behavior classification results obtained by the random forest model are output by the logging module, which records detailed logs (i.e., process behavior logs). The data persistence layer (Data Persistence) stores historical logs, random forest model parameters, and detection records, ensuring data persistence, continuous model optimization, and security incident tracing capabilities. This architecture diagram demonstrates the close integration of kernel and user space in this application and the data flow between modules.

[0074] In order to achieve the requirement of real-time monitoring of processes, some key parameters need to be strictly regulated: in the kernel state, the sampling interval and data cache size of data collection can be dynamically adjusted according to the actual situation, so as to ensure that the key events (such as target file operations and related processes) can still be collected in real time under the condition of high load of the system. The data processing module in the user state needs to be designed with multiple threads to ensure that data reception and processing can be performed in parallel, thereby reducing the overall response time. The prediction time of the random forest model needs to be controlled within milliseconds to meet the real-time requirement, and the model update frequency also needs to be dynamically set according to the actual monitoring data to maintain the detection accuracy.

[0075] Through the above embodiments, a full-process scheme of kernel state real-time monitoring based on eBPF and BCC tools is realized, and the random forest model is combined to record process behavior logs in the user state. The organic cooperation between the modules enables the system to achieve high real-time and high-accuracy security monitoring under the premise of low resource consumption, thereby improving the log quality and traceability analysis capability of the Linux system.

[0076] The application also provides a vehicle, which can apply the steps of the process behavior log generation method according to any one of the above embodiments.

[0077] The application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps of the process behavior log generation method according to any one of the above embodiments.

[0078] It should be understood that although each step in the flowchart of the accompanying drawings is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other orders. Moreover, at least part of the steps in the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or sub-steps or stages of other steps.

[0079] Those skilled in the art can clearly understand the technical solutions of the embodiments of the present application through the above description of the embodiments of the present application, which can be implemented by hardware or by means of software and necessary universal hardware platform. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the embodiments of the present application.

[0080] The preferred embodiments of the present application are described in detail above in combination with the drawings, but the present application is not limited to the specific details in the above-described embodiments, the above-described embodiments and the drawings are exemplary, and the modules or flows in the drawings are not necessarily required to implement the embodiments of the present application and cannot be understood as limiting the present application. Within the technical concept of the present application, various simple modifications and combinations of the technical solutions of the present application can be made, and these simple modifications and combinations all belong to the protection scope of the present application.

Claims

1. A method for generating a process behavior log, characterized in that: include: Monitor the process behavior of the target file operation and obtain target data, wherein the target data includes process behavior data; Performing statistics on the target data to generate statistical data; Vectorizing the statistical data to generate a target feature vector; The target feature vector is input into a random forest model for classification to identify the process behavior and generate a process behavior log.

2. The process behavior log generation method according to claim 1, characterized in that: The steps of monitoring the process behavior of the target file operation and obtaining the target data include: Defining an eBPF program, and attaching the eBPF program to a target system call entry related to the target file operation to collect target call data; The target call data is filtered based on preset filtering conditions to obtain the target data, wherein the preset filtering conditions include: process name and command line parameters.

3. The process behavior log generation method according to claim 1, characterized in that: The step of performing statistics on the target data to generate statistical data includes: The collected target data is aggregated and counted through the eBPF Map structure to generate statistical data of the target data.

4. The process behavior log generation method according to claim 1, characterized in that: Before the step of vectorizing the statistical data to generate a target feature vector, the method includes: Clean and format the statistical data of all processes, and count the read and write ratio, system call interval, and error call frequency of each process respectively. The statistical data includes the read and write operation counts, system call intervals, and cumulative statistics of error return codes of the process; Vectorization is performed based on the read-write ratio, system call interval, and error call frequency of each process to generate a target feature vector.

5. The process behavior log generation method according to claim 1, characterized in that: The dimension of the target feature vector includes at least one of the following: process read operation count, process write operation count, system call time interval mean and variance, and error return code ratio.

6. The process behavior log generation method according to claim 1, characterized in that: The step of inputting the target feature vector into a random forest model for classification to identify the process behavior and generate a process behavior log includes: Inputting the target feature vector into a random forest model to classify the process behavior and identify the process behavior; The process behavior log is generated based on the process behavior and the statistical data, and the process behavior log is stored in a classified manner.

7. The process behavior log generation method according to claim 1, characterized in that: After the step of classifying and storing the identified process behaviors as process behavior logs, the method includes: The process behavior corresponding to each process and the target file operated are obtained based on the association relationship between the process behavior log and the context PID.

8. The process behavior log generation method according to claim 1, characterized in that: The training method of the random forest model includes: Obtaining basic training data for the random forest model; Automatically labeling the basic training data according to preset process-behavior correspondence rules, associating the basic training data with process behaviors, and constructing a preliminary feature vector using the basic training data; The preliminary feature vector is input into an initial random forest model for training to generate the random forest model.

9. A vehicle, characterized in that: The vehicle applies the process behavior log generation method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the process behavior log generation method according to any one of claims 1 to 8.