A method for identifying and controlling violation of bastion machine operation instruction

CN122845211APending Publication Date: 2026-09-29CHINA SOUTHERN POWER GRID DIGITAL GRID GROUP (GUANGDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610992536.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

这就导致审计日志在关键的传递环节发生断裂,完整数据流的重建变得异常困难

Benefits of technology

本发明公开了一种堡垒机作业指令违规行为识别与管控方法,针对运维人员通过管道符号与重定向符号组合构造长命令,将敏感文件读取与网络外发操作拆分至多个子shell隔离执行空间,从而绕过传统单进程审计机制的数据泄露场景。本发明通过词法解析提取管道传递节点,追踪各子shell创建时触发的系统调用,识别跨越隔离壁垒的文件描述符传递环节,提取输入输出文件描述符映射的进程上下文节点,构建进程调用链与文件访问时序的路径重建图,将原本在隔离壁垒处断链的局部审计日志按传递顺序串联,还原出完整的跨子shell数据流转轨迹,最终通过评估轨迹中的敏感文件读取与网络外发载荷实现违规行为的精准识别,有效解决了复杂管道命令场景下的数据泄露监控盲区问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845211A_ABST
    Figure CN122845211A_ABST
Patent Text Reader

Abstract

The application provides a method for identifying and controlling violation of bastion machine operation instructions, comprising: obtaining surface characters input by a bastion machine terminal, extracting redirection symbol and pipe symbol combination structure composed of each instruction in the surface characters through lexical analysis; extracting input and output file descriptors in a transmission link, comparing input and output file descriptors mapped to a subshell isolated execution space entrance and exit to obtain process context nodes associated with the entrance and exit; traversing a path reconstruction graph in the order of process call chain transmission, concatenating branches in mutual communication to obtain a complete data flow conversion track; evaluating a sensitive file reading path in the complete data flow conversion track, extracting a network external action at the end of the track, classifying the load size of the network external action to obtain a violation behavior identification conclusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a method for identifying and controlling violations of bastion host operation instructions. Background Technology

[0002] As a security gateway for a company's core assets, the bastion host bears the responsibility of auditing and controlling all operational and maintenance operations. In organizations with extremely high data security requirements, such as financial institutions, internet companies, and government agencies, operations and maintenance personnel need to access sensitive resources such as database servers, application servers, and file storage systems daily through the bastion host. These operations harbor significant risks of data leakage; for example, operations and maintenance personnel may unauthorizedly export customer information, download financial data, or transfer business logs. Traditional auditing methods often rely on direct matching and interception of surface-level command characters. This approach ignores the spatial distribution differences of complex command combinations during internal system execution, making it easy to lose track of targets when faced with sophisticated spoofing operations. Currently, mainstream bastion host products generally adopt a keyword-based interception strategy, which involves setting up a blacklist to list prohibited command keywords, such as prohibiting the use of certain sensitive commands. When the commands entered by operations and maintenance personnel contain these keywords, the system immediately blocks the operation and logs it. The essence of this method is to directly compare the surface characters of the command text; as long as the character sequence is not in the blacklist, it is considered a compliant operation. This surface-level character matching method completely fails when facing complex scenarios. In real-world operations and maintenance (O&M) environments, O&M personnel often need to use pipe symbols to chain multiple commands together, forming complex data processing flows. For example, an O&M personnel might need to extract error information for a specific time period from a log file. They might first read the file content, then use a text filtering tool to filter lines containing error keywords, sort the results by time, and finally count the number of errors. Alternatively, an O&M personnel might input a long command containing multiple pipe symbols, intending to read a sensitive file, filter it multiple times, and then transfer it. In the system backend, this long command is not executed in a unified area but is distributed across multiple isolated independent spaces. When data is transferred from one space to another, the monitoring system can only record local data inflows and outflows, but cannot cross the spatial barriers to chain these actions into a complete trajectory. This leads to breaks in the audit logs at critical transmission points, making the reconstruction of the complete data flow extremely difficult. Therefore, in complex environments with multi-level pipe distributed operation, accurately obtaining the command combination structure and process isolation status, and then completely reconstructing the data flow path and dynamically determining violations, becomes a key issue in effectively controlling sensitive data leakage in bastion hosts. Summary of the Invention

[0003] This invention provides a method for identifying and controlling violations of bastion host operation instructions, the method comprising: Obtain the surface characters input from the bastion host terminal, and extract the combination structure of redirection symbols and pipe symbols composed of each instruction through lexical analysis; Based on the symbol combination structure, identify the pipe passing nodes formed by pipe symbol connections, trace the process creation system calls triggered when the pipe passing nodes are executed, identify multiple mutually isolated subshell isolated execution spaces created by them and carrying the split execution of long commands, as well as file read and write operations within each subshell isolated execution space; Analyze the data entry and exit actions generated by file read and write operations in each subshell isolated execution space, determine the local audit logs formed by the monitoring system recording each item, and analyze the transmission links that cross the isolation barrier of the subshell isolated execution space after connecting them in time sequence. Extract the input and output file descriptors in the transmission process, compare the entry and exit points of the subshell isolated execution space mapped by the input and output file descriptors, and obtain the process context nodes associated with the entry and exit points; Extract the process call chain and file access sequence between the context nodes of each process, determine the transmission order of the process call chain, and construct a path reconstruction graph. Connect the transmission links where the local audit logs are broken at the isolation barrier in the path reconstruction graph. Traverse the path reconstruction graph according to the process call chain, and connect the interconnected branches to obtain the complete data flow trajectory. Evaluate the sensitive file reading path in the complete data flow trajectory, extract the network outgoing actions at the trajectory endpoint, classify the payload scale of the network outgoing actions, and obtain the conclusion of violation identification.

[0004] Preferably, the step of obtaining the surface characters input by the bastion host terminal and extracting the combination structure of redirection symbols and pipe symbols composed of each instruction through lexical analysis includes: Obtain the original character sequence entered by the operation and maintenance personnel in the bastion host terminal session channel, extract the echo control character and cursor movement character from the original character sequence, mark the boundaries of the paragraphs enclosed in quotation marks, and distinguish the literal text inside the quotation marks from the command text outside the quotation marks. The command text is scanned character by character using a lexical analysis method and matched against a preset operator vocabulary. Two types of symbolic words, namely redirection symbols and pipe symbols, are identified and recorded along with their character offset positions into the lexical unit sequence. Based on the order of appearance of symbolic lexical units in the lexical unit sequence, the command words and parameter items between two adjacent symbolic lexical units are merged into instruction fragments. Each instruction fragment is then concatenated with the redirection symbols and pipe symbols preceding and following it according to their appearance positions. The connection relationship between the output end of the left instruction fragment and the input end of the right instruction fragment is marked in the concatenation result to obtain the symbol combination structure.

[0005] Preferably, the step of identifying the pipe transfer node formed by the pipe symbol connection based on the symbol combination structure, tracing the process creation system call triggered when the pipe transfer node is executed, identifying multiple mutually isolated subshell execution spaces created by it and carrying the split execution of long commands, and the file read and write operations within each subshell execution space, includes: The positions marked as pipeline symbols by word categories are selected from the symbol combination structure and arranged in order of appearance to obtain the pipeline transmission node sequence; The kernel audit subsystem is used to mount the hook point of the process creation system call, capture the child process identifier and parent process identifier returned by each process creation, organize the capture results into a process derivation tree, and identify several child processes continuously derived by the same parent process under the execution trigger of the pipe passing node as the subshell isolated execution space; For each of the subshell isolated execution spaces, corresponding hook points are attached to the file open, read, and write system calls issued during operation. The file descriptor and the bound file path are captured, and the file descriptor, along with its process identifier, read / write direction, and operation time, are registered as file read / write operation records.

[0006] Preferably, the step of parsing the data entry and exit actions generated by file read and write operations within each subshell isolated execution space, determining the local audit logs formed by the monitoring system recording each item, and analyzing the transmission links that cross the isolation barrier of the subshell isolated execution space after sequential connection includes: The file descriptor integer number, absolute path, operation direction identifier, and operation time are parsed from the read and write operation records of each file, and the local audit log is obtained by grouping them according to the process identifier. The actions in the local audit log are arranged in ascending order of operation time. Adjacent action pairs belonging to different subshell isolated execution spaces are traversed. It is determined whether the difference in the order of operation time between the data outgoing action in the preceding space and the data entering action in the following space falls into a preset connection time window. Those that fall into the window are marked as belonging to the same transmission link that crosses the isolation barrier. The outgoing direction block and the incoming direction block are connected in series to obtain the transmission link sequence.

[0007] Preferably, the step of extracting input and output file descriptors in the transmission process, comparing the entry and exit points of the subshell isolated execution space mapped by the input and output file descriptors, and obtaining the process context node associated with the entry and exit points includes: Traverse the sequence of transmission links, read the integer number of the file descriptor and the process identifier of the process from the outgoing and incoming blocks, and obtain the input and output file descriptor tuples; In the kernel-side file descriptor table held by the subshell isolated execution space corresponding to the process identifier of the outgoing block, the write end of the kernel pipe object bound to the output file descriptor is looked up as the space exit, and the read end of the kernel pipe object bound to the input file descriptor is looked up as the space entry in the corresponding table of the incoming block. Compare the kernel pipe object index numbers pointed to by the space exit and space entrance. If they point to the same index number, it is determined that a docking has been formed. Extract the executable image path, command line argument string, effective user identifier, and parent process identifier corresponding to the process identifiers of both ends of the docking, and assemble them into process context nodes associated with the entry and exit points according to the index number of the docked kernel pipe object.

[0008] Preferably, the step of extracting the process call chain and file access sequence between the context nodes of each process, determining the transmission order of the process call chain, and constructing a path reconstruction graph, connecting the transmission links where the local audit log is broken at the isolation barrier in the path reconstruction graph, includes: For each process context node, trace the process identifier to the parent process identifier and then back to the command interpreter main process identifier to obtain the parent-child derivation chain, and then merge the common process identifier nodes into a process call chain. Arrange the access actions in ascending order of access time to obtain the access sequence string, and assign them a transmission sequence number from small to large. A path reconstruction graph is constructed using process context nodes as vertices and kernel pipe object index numbers that connect adjacent vertices by transmission sequence numbers as edges. For positions lacking edge connections, matching index numbers are retrieved from the self-transmission link sequence and added.

[0009] Preferably, the step of traversing the path reconstruction graph according to the process call chain order and connecting the interconnected branches to obtain the complete data flow trajectory includes: From the set of vertices in the path reconstruction graph, select the vertex with the smallest transmission sequence number as the starting vertex, and read the directed edges in ascending order of transmission sequence number to form a trajectory segment. For positions with multiple parallel branches, read the kernel pipeline object index number from the directed edge attribute field, determine branches with the same index number as common source branches and aggregate them to the same vertex, and repeatedly perform the head-to-tail connection until all trajectory segments are concatenated into a single sequence.

[0010] Preferably, the process of evaluating sensitive file reading paths in the complete data flow trajectory, extracting network outbound actions at the trajectory endpoints, classifying the payload scale of network outbound actions, and obtaining conclusions on violation identification includes: Traverse the file read and write operation records of each vertex in the complete data flow trajectory, extract the operation direction identifier as the absolute path of the read, and perform character prefix matching with the pre-established sensitive file path library. Those that fall within the coverage area are registered as sensitive read hit points. The vertex with the largest sequence number is used as the trajectory endpoint. Entries with the binding object category of socket and the operation direction identifier of write are filtered. The destination Internet Protocol address is compared with the list of internal address ranges. Those not included are identified as network outbound actions and the number of bytes is accumulated to obtain the payload scale. By comparing the low-level and high-level thresholds in the load scale classification threshold table, low-level, medium-level, or high-level outward identification labels are assigned respectively, and the results are summarized as the conclusion of violation identification.

[0011] Preferably, the redirection symbols cover input redirection, output redirection, append redirection, and file descriptor copy variants, and the pipe symbols cover unidirectional pipes and error stream merging pipes.

[0012] Preferably, the step of mapping each of the subshell isolated execution spaces to the position points in the pipeline transmission node sequence one by one in the order of derivation is used to determine the instruction fragment carried by each of the subshell isolated execution spaces.

[0013] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses a method for identifying and controlling violations of bastion host operation commands. It addresses data leakage scenarios where operations and maintenance personnel construct long commands using combinations of pipe and redirection symbols, splitting sensitive file reading and network outbound operations into multiple isolated subshell execution spaces, thereby bypassing traditional single-process auditing mechanisms. This invention extracts pipe transmission nodes through lexical analysis, tracks system calls triggered when each subshell is created, identifies file descriptor transmission links crossing isolation barriers, extracts process context nodes mapped to input and output file descriptors, constructs a path reconstruction graph of process call chains and file access sequence, and reconnects previously broken local audit logs at the isolation barrier according to the transmission order, reconstructing the complete cross-subshell data flow trajectory. Finally, by evaluating sensitive file reading and network outbound loads in the trajectory, accurate identification of violations is achieved, effectively solving the problem of blind spots in data leakage monitoring in complex pipe command scenarios. Attached Figure Description

[0014] Figure 1 This is a flowchart of a method for identifying and controlling violations of bastion host operation instructions according to the present invention.

[0015] Figure 2 This is a schematic diagram of a method for identifying and controlling violations of bastion host operation instructions according to the present invention.

[0016] Figure 3 This is another schematic diagram of a method for identifying and controlling violations of bastion host operation instructions according to the present invention. Detailed Implementation

[0017] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0018] like Figures 1-3 This embodiment of a method for identifying and controlling violations of bastion host operation instructions may specifically include: S101. Obtain the surface characters input by the bastion host terminal, and extract the combination structure of redirection symbols and pipe symbols composed of each instruction through lexical analysis.

[0019] The system acquires the original character sequence entered by the operations and maintenance personnel in the bastion host terminal session channel. Echo control characters and cursor movement characters are extracted from the original character sequence, retaining the printable character portion that constitutes a complete instruction. Boundary marking is performed on quotation-enclosed paragraphs mixed in with the printable character portion to distinguish between literal text within quotation marks that do not participate in syntax segmentation and command text outside quotation marks that do participate in syntax segmentation, thus obtaining the command text defined at the surface character level. For the command text, a lexical analysis method is used to scan and match a preset operator vocabulary character by character. Two types of symbolic units are identified from the command text: redirection symbols and pipe symbols. The redirection symbols cover input redirection, output redirection, append redirection, and file descriptor copy variants. The pipe symbols cover unidirectional pipes and error stream merging pipes. The identified symbolic units, along with their corresponding character offset positions, are recorded in the lexical unit sequence. Based on the order of appearance of symbolic lexical units in the lexical unit sequence, the command word and parameter item between two adjacent symbolic lexical units are merged into an instruction fragment. Each instruction fragment is then concatenated with the redirection symbol and the pipeline symbol preceding and following it according to their appearance positions. The connection relationship between the output end of the left instruction fragment and the input end of the right instruction fragment is marked in the concatenation result, thus obtaining a symbol combination structure for subsequent tracking of pipeline transmission nodes.

[0020] In bastion host operation and maintenance auditing scenarios, operations and maintenance personnel interact with the target host through a session channel, which carries a stream of characters typed on the keyboard. Specific implementations perform lexical-level processing on this character stream to extract the symbol combination structure upon which subsequent tracing pipeline nodes rely. In one implementation, the acquisition of the original character sequence occurs at the bastion host proxy forwarding layer. Each time an operations and maintenance personnel press a key, the terminal protocol transmits the corresponding bytes to the proxy forwarding layer. In addition to printable characters constituting commands, the byte stream also includes bytes used for terminal display control, such as cursor left / right movement, character deletion, line jump, and screen refresh control bytes. The stripping process identifies and removes these control bytes, retaining only the characters falling within the printable range.

[0021] It's important to note that the stripping process doesn't change the relative order of printable characters; it only removes interspersed control bytes, ensuring the remaining character sequence precisely corresponds to the command text intended for execution by the operations and maintenance personnel. If the personnel press the backspace key to correct a character during input, the stripping process will remove the preceding character according to the semantics of the backspace control byte. When multiple backspaces are encountered consecutively, the system deletes the corresponding number of characters sequentially from the current cursor position backward until all deletions are completed or the buffer begins. If the backspace operation reaches the beginning of the line but there are still remaining backspace instructions, the excess backspace requests are ignored, preserving the current line as blank. Characters re-entered after a backspace are appended to the current cursor position, forming a new command sequence. After the above processing, the final remaining operations and maintenance text is consistent with the actual command submitted for execution. Further processing is applied to the text content, including paragraphs enclosed in quotation marks. In operations and maintenance commands, single and double quotation marks are often used to enclose parameters containing spaces, special symbols, or operator literals.

[0022] For example, when filtering logs, operations and maintenance personnel might pass a string containing pipe literals as a parameter to a text filtering tool. The tool scans the printable character portion character by character, entering the quote-in-quote state when encountering a single or double quote, and exiting the quote-in-quote state when encountering a closed quote of the same type. This distinguishes between literal text within quotes that does not participate in syntax segmentation and command text outside quotes that does participate in syntax segmentation.

[0023] Preferably, for backslash escaping scenarios, the scanning process needs to identify two escaping cases: when a single backslash is encountered immediately followed by a quotation mark, the backslash escapes the quotation mark, and the quotation mark is not considered a boundary; when two consecutive backslashes are encountered, the first backslash escapes the second backslash, and if the double backslash is immediately followed by a quotation mark, the quotation mark is considered a normal boundary marker. Specifically, the determination method is to count the number of consecutive backslashes *n* before the quotation mark. If *n* is odd, the quotation mark is escaped and not considered a boundary; if *n* is even, the quotation mark is treated as a boundary. The command text obtained after boundary marking is used as the input for subsequent lexical parsing.

[0024] In one possible implementation, the operator lexicon pre-includes all operator literals supported by the command interpreter running on the target host audited by the bastion host. Each entry in the operator lexicon contains three fields: literal string, lexical category, and longest match length. The lexical category distinguishes between redirection operators and pipe symbols. For the command text, the lexical parsing method proceeds character by character, attempting to match the entry in the operator lexicon with the character segment starting at the current position in a longest match length priority manner at each progress point.

[0025] Specifically, the redirection symbols cover input redirection, corresponding to the literal form of a less-than sign; output redirection, corresponding to the literal form of a greater-than sign; append redirection, corresponding to the literal form of two consecutive greater-than signs; and file descriptor copy variant form, corresponding to the literal form of a greater-than sign followed by an AND sign followed by a number. The pipe symbol covers unidirectional pipes, corresponding to the literal form of a vertical bar; and error stream merging pipes, corresponding to the literal form of a vertical bar followed by an AND sign. After a match is found, the lexical parsing method writes the matched literal string, its corresponding term category, and the character offset position of the match in the command text into the lexical unit sequence.

[0026] Understandably, character segments that miss the operator lexicon are classified as command words or parameter items, and are also written into the lexical unit sequence with whitespace as the natural boundary. Furthermore, based on the order of occurrence of symbolic lexical units in the lexical unit sequence, command words and parameter items between two adjacent symbolic lexical units are merged into a single instruction segment.

[0027] For example, if the command text entered by the operations and maintenance personnel contains two pipe symbols, the command text will be divided into three instruction fragments, each corresponding to a continuous sequence of command words and parameter items. For boundary cases, if the command text begins directly with a pipe symbol, the first instruction fragment is treated as an empty fragment with a reserved position marker; if it ends with a pipe symbol, the last instruction fragment also retains an empty fragment marker; if multiple consecutive pipe symbols appear, empty instruction fragments are inserted between the consecutive pipe symbols as placeholders. Finally, each instruction fragment is concatenated with its preceding and following redirection symbols and pipe symbols according to their position. For each pipe symbol position, the concatenation result indicates that the output of the command unit on its left is connected to the input of the execution unit on its right; for each redirection symbol position, the number on the left of the symbol is parsed as the file descriptor number, and the lexical unit on the right of the symbol is parsed as the target file path, establishing a mapping relationship between the two. This concatenation result forms a tree structure containing node identifiers, connection relationships, and data flow direction, which is provided externally for subsequent steps to track the subshell isolated execution space corresponding to the pipe transmission nodes. The structure contains three types of elements: the order of pipe nodes, the redirection mapping from file descriptors to paths, and the position index of each node in the syntax tree.

[0028] S102. Identify the pipe transfer node formed by the pipe symbol connection based on the symbol combination structure, track the process creation system call triggered when the pipe transfer node is executed, identify multiple mutually isolated subshell execution spaces created by it and carrying the split execution of long commands, as well as file read and write operations in each subshell execution space.

[0029] The symbol combination structure obtained in the previous step is obtained. Positions labeled as pipe symbols are selected from this structure. Each position, along with the connection between its left-hand instruction fragment output and right-hand instruction fragment input, is registered as a pipe delivery node. These positions are arranged in the order of their appearance in the command text to obtain a pipe delivery node sequence. When this pipe delivery node sequence is executed on the target host audited by the bastion host, the kernel audit subsystem is used to attach hooks to process creation system calls. The child process identifier and parent process identifier returned by each process creation system call are captured. The captured child process identifiers are organized into a process derivation tree according to their parent process identifiers. Several child processes continuously derived from the same parent process under the trigger of pipe delivery node execution are identified as subshell isolated execution spaces. Each subshell isolated execution space is mapped one by one to the position points in the pipe delivery node sequence according to the order of derivation to determine the instruction fragment carried by each subshell isolated execution space. For file open, read, and write system calls issued by each of the aforementioned subshell isolated execution spaces during operation, the kernel audit subsystem is used to attach corresponding hook points to capture the file descriptor returned when opening a file and the file path bound to the file descriptor. The file descriptor, along with the process identifier, read / write direction, and operation time of the subshell isolated execution space to which it belongs, is registered as a file read / write operation record within that subshell isolated execution space.

[0030] On the target host audited by the bastion host, if a long command submitted by the operations and maintenance personnel contains a pipe symbol, the command interpreter will split the entire command into multiple child processes that do not share memory space and execute them in parallel. The child processes exchange data only through anonymous pipes maintained by the kernel; traditional command character matching methods cannot trace data across the boundaries between these child processes. Specifically, this implementation extends from symbol recognition to kernel behavior capture to address the split execution characteristic. In one implementation, for the symbol combination structure obtained in the previous step, each lexical position is traversed. When a position marked as a pipe symbol is reached, the output identifier of the nearest instruction fragment to the left and the input identifier of the nearest instruction fragment to the right of that position are read. The instruction fragment numbers at both ends, their offset positions in the corresponding command texts, and the connection direction are written into a pipe transmission node record.

[0031] Specifically, the pipeline delivery node record is stored in a ternary structure, carrying the left-end instruction fragment number, the pipeline symbol offset position, and the right-end instruction fragment number. Multiple pipeline delivery node records are arranged in ascending order of offset position, forming the pipeline delivery node sequence. Audit data collection requires capturing system call behavior at the operating system kernel level. The target host operating system kernel provides audit infrastructure, including a loadable kernel probe mechanism and audit framework, collectively referred to as the kernel audit subsystem. This subsystem implants hook points at the entry and exit addresses of specified system calls. The hook points contain callback function pointers and trigger condition identifiers. When a user-mode process executes a system call, the kernel scheduler detects the hook point identifier and synchronously triggers the callback function. The callback function obtains the call context data by reading the CPU general-purpose register group, the parameter area in the kernel stack frame, and the return value register. In specific implementation, the hook point locates the address of the system call processing function through the kernel symbol table, replaces the first 5 bytes of the original instruction sequence with a jump instruction, redirects the execution flow to the callback function entry point, and restores the original instructions and continues execution after the callback function completes its execution.

[0032] In one possible implementation, system calls related to the process are created, such as entry points for cloning, forking, and execution replacement system calls, and corresponding hook point callback functions are registered for each. When a callback function is invoked, the current process identifier that triggered the call is read from the kernel task structure as the parent process identifier, and the newly created process identifier is read from the return value as the child process identifier. The parent process identifier and the child process identifier, along with the call time and call type, are written into the derived event stream.

[0033] Specifically, the derived event stream accumulates all process creation events in chronological order. Based on the dependency relationship between child process identifiers and parent process identifiers, the derived event stream is merged and organized. Several child process identifiers under the same parent process identifier are treated as branches of the node corresponding to that parent process identifier, thus constructing a process derivation tree with the command interpreter main process as the root node and the derived child processes as their subordinate nodes. For several child process nodes directly subordinate to the command interpreter main process node in the process derivation tree, the difference in the invocation time of each derived event in the derived event stream is used to determine whether they are within a continuous derivation sequence triggered by the execution of the same long command: if the difference in the invocation time of the several child process nodes is less than a preset derivation interval threshold, then the several child process nodes are identified as a subshell isolated execution space carrying the split execution of the same long command. One feasible value for the derivation interval threshold is 50 milliseconds. The derivation interval threshold is determined by: collecting samples of the derivation time differences of adjacent subprocesses continuously split and derived by the same command interpreter when executing several commands containing pipe symbols, and taking the upper envelope of the sample distribution (such as the mean plus twice the standard deviation) as the upper bound of the threshold, so that the adjacent intervals of the same long command split and derived fall within the threshold and the derivation intervals between independent commands fall outside the threshold. The derivation interval threshold is determined accordingly, and one feasible value is 50 milliseconds.

[0034] It is understood that each of the subshell isolated execution spaces has its own independent virtual address space, file descriptor table, and signal processing context, and there is no memory sharing between them. Furthermore, according to the order in which each subshell isolated execution space appears in the derived event stream, it is aligned one by one with the instruction fragment numbers at the left and right ends of the position point in the pipelined node sequence. The subshell isolated execution space that appears earlier is mapped to the instruction fragment number at the left end of the position point, and the subshell isolated execution space that appears later is mapped to the instruction fragment number at the right end, thereby determining the specific instruction fragment carried by each subshell isolated execution space. Furthermore, for the three types of system calls (file open, file read, and file write) issued by the subshell isolated execution space during operation, hook points are respectively attached to the corresponding entry and exit points of the kernel audit subsystem. When a file open system call is triggered, the hook point callback function reads the absolute path of the target file from the call parameters, reads the kernel-allocated file descriptor integer number from the return value, and writes the binding relationship between the file descriptor integer number and the absolute path into the path binding table.

[0035] For example, for system calls of file reading and file writing, the hook callback function reads the integer number of the file descriptor referenced in this operation from the call parameters, looks up the absolute path corresponding to the file descriptor according to the path binding table, and assembles the file descriptor, the absolute path, the operation direction identifier, the operation time, and the process identifier of the current subshell isolated execution space into a file read / write operation record. The operation direction identifier is used to distinguish between reading and writing; for example, a read system call corresponds to a read direction, and a write system call corresponds to a write direction.

[0036] It should be noted that the file read / write operation records are collected separately according to their respective subshell isolated execution spaces, covering all file open, read, and write actions within each subshell isolated execution space. These file read / write operation records, together with the aforementioned pipeline transmission node sequence and the mapping relationship between subshell isolated execution spaces and instruction fragments, constitute the data foundation for subsequent auditing.

[0037] S103. Analyze the data entry and exit actions generated by file read and write operations in each subshell isolated execution space, determine the local audit logs formed by the monitoring system recording them one by one, and analyze the transmission links that cross the isolation barrier of the subshell isolated execution space after connecting them in time sequence.

[0038] Retrieve file read / write operation records from each of the subshell isolated execution spaces registered in the previous step. Parse each record to extract the file descriptor integer number, absolute path, operation direction identifier, operation time, and process identifier. For records with an operation direction identifier of "read," extract the data entry action; for records with an operation direction identifier of "write," extract the data exit action. Group these data entry and exit actions according to their respective process identifiers to obtain a local audit log indexed by the process identifier. Sort all actions in the local audit log in ascending order by operation time to obtain a sequentially linked local audit log sequence. Iterate through adjacent action pairs belonging to different subshell isolated execution spaces in the local audit log sequence. For data exit actions generated in the earlier subshell isolated execution space and data entry actions generated in the later subshell isolated execution space, determine whether the difference in their operation times falls within a preset connection window. If they do, mark them as belonging to the same transmission stage that crosses the isolation barrier. For the marked transmission links, the file descriptor integer number and the process identifier corresponding to the data outgoing action are registered as outgoing direction blocks, and the file descriptor integer number and the process identifier corresponding to the data entering action are registered as incoming direction blocks. The outgoing direction blocks and the incoming direction blocks are concatenated in chronological order of operation time to obtain a transmission link sequence that crosses the isolation barriers of each subshell's isolated execution space.

[0039] On the target host audited by the bastion host, a long command executed by splitting it using pipe symbols appears on the kernel side as several independent child processes that do not share memory, each initiating read and write operations to a file or anonymous pipe. The bastion host is a security audit gateway device deployed in the operations and maintenance network, responsible for collecting and forwarding system call logs generated by the kernel audit subsystem on the target host to the monitoring system for analysis. The long command refers to a sequence of shell commands containing at least one pipe symbol or redirection symbol in a single line of input, split by the command interpreter into multiple isolated child processes. These child processes are created via the fork system call during execution, each possessing an independent process identifier and file descriptor table, and exchanging data streams with each other through anonymous pipes. The monitoring side can only register local actions process by process, lacking cross-process data stream concatenation. To address this local registration characteristic, this solution implements timing-based concatenation and cross-process identification. In one implementation, for the set of file read and write operation records assembled in the previous stage, the monitoring system sorts all records in ascending order according to the operation time field, and marks the write records and read records that belong to different subshell isolated execution spaces and whose operation time differences fall within the connection time window as a pair of candidate adjacent action pairs; whether the candidate adjacent action pairs are indeed connected by the same anonymous pipe is left to be verified by the kernel pipe object index number in subsequent stages, and the transmission stage sequence is obtained accordingly.

[0040] Specifically, when the operation direction identifier is set to "read," it indicates that the operation pulls bytes from the read end of a file or anonymous pipe, and the corresponding record is classified as a data entry action. When the operation direction identifier is set to "write," it indicates that the operation pushes bytes out of the file or anonymous pipe, and the corresponding record is classified as a data exit action. Further, the data entry and data exit actions are bucketed according to their respective process identifiers, with each process identifier corresponding to one bucket. Within each bucket, the original order of the operations is maintained. These buckets serve as local audit logs indexed by process identifiers, reflecting the byte entry and exit history within the isolated execution space of a single subshell.

[0041] It should be noted that each individual local audit log only reflects actions within the scope of its own process identifier and does not yet present the data relay relationship between processes. All local audit logs are sorted in ascending order by operation time to obtain a sequentially linked sequence of local audit logs.

[0042] In one possible implementation, adjacent action pairs belonging to different subshell isolated execution spaces are traversed in the local audit log sequence.

[0043] Specifically, the method for determining adjacent action pairs is as follows: Take any data outgoing action in the sequence, scan forward for the next action with a different process identifier. If the operation direction identifier of the next action is read in, then the outgoing action and the read in action constitute an adjacent action pair. For each adjacent action pair, compare the time difference between the data outgoing action and the data entering action to determine whether the time difference falls within a preset connection window. The connection window refers to the maximum allowable latency for transmission across isolation barriers. Its value is determined by statistically analyzing the measured latency distribution of the write and read actions at the same anonymous pipe end and taking the upper quantile of the distribution as the upper limit. The determination uses a closed interval, that is, if the time difference is less than or equal to the connection window, it is considered to fall within the time limit. One possible value is 10 milliseconds. If the time difference falls within the connection window, then the adjacent action pair is marked as belonging to the same transmission link across the isolation barrier. For each marked transmission step, the file descriptor integer number corresponding to the data outgoing action and its associated process identifier are packaged to obtain an outgoing block; the file descriptor integer number corresponding to the data entering action and its associated process identifier are packaged to obtain an incoming block. The outgoing block carries the output identity information of the transmission step, and the incoming block carries the input identity information of the transmission step. Finally, according to the chronological order of the data outgoing actions, the outgoing and incoming blocks in all transmission steps are sequentially concatenated to assemble a transmission step sequence that crosses the isolation barriers of each subshell's execution space. This transmission step sequence provides clear continuation nodes at the breaks in the local audit log, serving as the connecting anchor points upon which the subsequent reconstruction of the complete data flow trajectory depends.

[0044] S104. Extract the input and output file descriptors in the transmission process, compare the entry and exit points of the subshell isolated execution space mapped by the input and output file descriptors, and obtain the process context node associated with the entry and exit points.

[0045] Obtain the sequence of transmission links obtained in the previous step, traverse each transmission link in the sequence, read the integer number of the output file descriptor and the process identifier from the outgoing block of the transmission link, and read the integer number of the input file descriptor and the process identifier from the incoming block of the transmission link to obtain the input and output file descriptor tuples aligned step by step. For the input / output file descriptor tuple, in the kernel-side file descriptor table held by the subshell isolated execution space corresponding to the process identifier of the outgoing block, the write end of the kernel pipe object bound to the output file descriptor is looked up in reverse as the space exit of the subshell isolated execution space; in the kernel-side file descriptor table held by the subshell isolated execution space corresponding to the process identifier of the incoming block, the read end of the kernel pipe object bound to the input file descriptor is looked up in reverse as the space entry of the subshell isolated execution space; the kernel pipe object index number pointed to by the space exit and the space entry is compared. The kernel pipe object index number is a kernel allocation number shared by both ends of the same anonymous pipe. If the two point to the same kernel pipe object index number, it is determined that the two are connected on the same pipe transmission node. For the space exit and space entrance that form a docking, extract the executable image path, command line parameter string, effective user identifier, and parent process identifier corresponding to the process identifier of the space exit, and assemble them into an exit-side process context node; assemble the entrance-side process context node with the same fields, and associate the exit-side process context node and the entrance-side process context node according to the index number of the docked kernel pipe object to obtain the process context node associated with the exit and entrance.

[0046] The timing of data outgoing and incoming actions in the transmission process alone is insufficient to confirm that the two ends are indeed connected through the same anonymous pipe. On the target host audited by the bastion host, when multiple pipe commands are executed concurrently, ingoing and outgoing actions occurring at similar times may belong to unrelated pipes. A specific implementation introduces bidirectional reverse lookup of file descriptors and verification of kernel pipe object index numbers above the transmission sequence to establish the correspondence between entry / exit points and process identities. In one implementation, for the transmission sequence obtained in the previous step, each transmission step is read sequentially. The outgoing block of each transmission step carries the output end identity information, and the incoming block carries the input end identity information. The integer number of the output file descriptor and the process identifier of its parent process are read from the outgoing block, and the integer number of the input file descriptor and the process identifier of its parent process are read from the incoming block. The four values ​​are then assembled into an input / output file descriptor tuple according to the transmission sequence.

[0047] It should be noted that the kernel-side file descriptor table is a resource index table independently maintained by the operating system for each process in kernel mode. Each entry in the table uses the integer number of the file descriptor as the key and a pointer to a specific kernel object as the value. Kernel objects can be open disk files, sockets, anonymous pipes, etc. For input and output file descriptor tuples, the kernel object is searched in the subshell isolated execution space corresponding to the outgoing and incoming blocks, respectively. The subshell isolated execution space refers to the independent file descriptor table space owned by the child process created by the fork system call. In the kernel-side file descriptor table, the table entry is directly located using the integer number as an index to obtain the kernel object pointed to by the entry. If the kernel object corresponding to the output descriptor is the write end of an anonymous pipe, it is recorded as the space exit; if the kernel object corresponding to the input descriptor is the read end of an anonymous pipe, it is recorded as the space entry. When the operating system creates an anonymous pipe, it simultaneously generates two kernel objects, one for the read end and one for the write end, with both ends sharing the same kernel allocation number K, which serves as the pipe index number. The pipe index number K carried by the space exit and space entry are compared. out and K in If K out equals K in If the adjacent action pairs are equal, they are determined to have formed a valid connection at the pipeline transfer node and added to the transfer sequence; otherwise, the action pairs are removed. The transfer sequence consists of verified pipeline connections in chronological order, with each link containing a pipeline index number and inlet / outlet information. For the space exit that forms a connection, four fields are extracted from the process control block: executable image path, command-line argument string, valid user identifier (UID), and parent process identifier (PPID). The executable image path reflects the binary program loaded by the process, the command-line argument string reflects the instruction fragment received when the process starts, the UID reflects the permissions and identity claimed by the process, and the PPID reflects the process's origin. These four fields, along with the process identifier, are packaged into an exit-side process context node. Similarly, for the space entry that forms a connection, the corresponding values ​​are extracted from the process control block corresponding to the process identifier of the space entry according to the same set of fields and packaged into an entry-side process context node. The method for determining the docking relationship is as follows: when the write end file descriptor of the kernel pipe object associated with the space exit and the read end file descriptor of the kernel pipe object associated with the space entry point to the same underlying pipe structure, a docking relationship is established. In this case, the unique identifier assigned by the system to the shared kernel pipe object at creation time serves as the index number. Based on this index number, a pointer association is established between the exit-side process context node and the entry-side process context node to obtain the process context nodes associated with the exit and entry points. The process context node carries the real process identities and runtime contexts at both ends of the transmission link, providing the node objects upon which subsequent path reconstruction depends.

[0048] S105. Extract the process call chain and file access sequence between the context nodes of each process, determine the transmission order of the process call chain, and construct a path reconstruction graph. Connect the transmission links where the local audit logs are broken at the isolation barrier in the path reconstruction graph.

[0049] The process context nodes associated with the entry and exit points obtained in the previous step are retrieved. For each process context node, the parent process identifier and the process identifier it belongs to are traced step by step from the process identifier to the parent process identifier, leading to the command interpreter main process identifier, to obtain the parent-child derivation chain corresponding to that process context node. The parent-child derivation chains of several process context nodes are overlapped and merged according to the common process identifier to obtain the process call chain that connects the process context nodes. For the file read and write operation records in the isolated execution space of the subshell to which each process context node belongs, the operation time field is read, and the access actions of each process context node are arranged in ascending order of time to obtain the access time sequence string. The access time sequence string is traversed, and the process context nodes to which they belong are assigned a transmission sequence number in ascending order according to the order of access actions. The transmission sequence number determines the transmission order between the process context nodes in the process call chain. A directed graph structure is used to lay out the process call chain with additional transmission sequence numbers. Each process context node is used as a vertex, and the kernel pipe object index number connecting two adjacent vertices with transmission sequence numbers is used as an edge to obtain a path reconstruction graph. For positions in the path reconstruction graph where adjacent vertices lack edge connections, a transmission link with a matching kernel pipe object index number is retrieved from the transmission link sequence and added to the path reconstruction graph as a supplementary edge to complete the connection of transmission links where the local audit log is broken at the isolation barrier.

[0050] In bastion host auditing scenarios, simply obtaining the process context nodes associated with the entry and exit points is insufficient to present the complete data flow of a long command on the kernel side. Multiple sets of process context nodes appear as parallel entries during kernel accounting, and their parent-child derivation relationships and temporal sequences are not naturally connected. Specific implementations address this by tracing the parent-child derivation and organizing the temporal sequence of these parallel entries, and then integrating them using a graph structure. In one implementation, for each process context node, starting from its own process identifier, the system searches for the node corresponding to that process identifier in the process derivation tree, reads the parent process identifier of that node, and then repeatedly performs the search using the parent process identifier as the new own process identifier until the command interpreter main process identifier is encountered, recording all process identifiers encountered along the way.

[0051] Specifically, all process identifiers along the path are arranged in order from leaf to root, forming a parent-child derivation chain corresponding to the process context node. The head entry in the parent-child derivation chain is the process identifier to which the process context node itself belongs, the tail entry is the command interpreter main process identifier, and the middle entries are several intermediate process identifiers that the command interpreter main process traverses during the process of deriving the process identifier to which it belongs.

[0052] It should be noted that the parent-child derivation chains of multiple process context nodes share the same command interpreter main process identifier as a common root, and they often share several intermediate process identifiers as well. By merging the parent-child derivation chains of several process context nodes according to their shared process identifiers—that is, treating nodes with the same process identifier as the same node and not registering them repeatedly—a process call chain is obtained, linking the process context nodes. This process call chain presents as a tree with the command interpreter main process identifier as the root and the process identifiers to which each process context node belongs as leaves.

[0053] It is understandable that the process call chain described only depicts the derivation relationship at the process derivation level, and does not yet depict the sequential relationship at the file access level.

[0054] In one possible implementation, for each process context node's isolated execution space within its subshell, file read / write operation records are read one by one, including the operation time field. The access actions for each process context node are stored as binary entries, each containing the operation time and a reference to its respective process context node. All binary entries for access actions are sorted in ascending order by operation time to obtain an access sequence string. This access sequence string reflects the sequential arrangement of access actions along a timeline.

[0055] Specifically, the access sequence string is traversed, and the process context nodes are assigned ascending transmission sequence numbers according to the order of access actions. When a process context node first appears in the access sequence string, the next unused integer starting from 1 is assigned as its transmission sequence number in order of appearance; when the process context node appears again, the previously assigned transmission sequence number remains unchanged. Once assigned, the transmission sequence number is not reassigned in subsequent stages. The value of the transmission sequence number reflects the relative position of the process context node in the data flow direction. The transmission sequence number is appended to the attribute field of the corresponding node in the process call chain. Further, a directed graph structure is used to lay out the process call chain with the appended transmission sequence numbers. The directed graph structure consists of a vertex set and a directed edge set. Each vertex in the vertex set corresponds to a process context node, and the vertex's attribute field contains four items: the process identifier, the transmission sequence number, the executable image path, and the command-line parameter string.

[0056] Specifically, for two vertices with adjacent transmission sequence numbers, a directed edge is placed from the former vertex to the latter. The attribute field of the directed edge is written with the kernel pipeline object index number determined during the entry / exit association phase between the two vertices. All directed edges successfully placed based on the kernel pipeline object index number are included in the directed edge set of the path reconstruction graph. For positions in the path reconstruction graph where no directed edge is placed between adjacent vertices, these positions are identified as missing edge positions. The missing edge positions correspond to the transmission links where the local audit logs are broken at the isolation barrier. For each missing edge position, the transmission link entries corresponding to the preceding vertex of the process identifier of the outgoing block and the following vertex of the process identifier of the incoming block are retrieved from the transmission link sequence. The kernel pipeline object index number shared by the outgoing and incoming blocks is read from the retrieved transmission link entries and added to the directed edge set of the path reconstruction graph in the form of supplementary directed edges. The kernel pipeline object index number is also written in the attribute field of the supplementary directed edge. At this point, there are directed edges connecting adjacent vertices in the reconstructed path graph. The transmission links that were previously broken at the isolation barrier of the subshell execution space have all been connected. The reconstructed path graph carries a complete and traversable data flow skeleton.

[0057] S106. Traverse the path reconstruction graph according to the process call chain, and connect the interconnected branches to obtain the complete data flow trajectory.

[0058] The path reconstruction graph obtained in the previous step is obtained. From the vertex set of the path reconstruction graph, the vertex with the smallest transmission sequence number is selected as the starting vertex. Starting from the starting vertex, directed edges originating from the current vertex are read one by one in ascending order of transmission sequence number. The vertices and directed edges traversed are sequentially concatenated to form a trajectory segment. For positions in the path reconstruction graph with multiple parallel branches, the kernel pipeline object index number registered in the directed edge attribute field of each parallel branch is read. Several parallel branches with the same kernel pipeline object index number are determined to be common branches. These common branches are aggregated to the same vertex according to their transmission sequence number to obtain the interconnection relationship between branches. Regarding the trajectory segment and the interconnection relationship, the last vertex of the trajectory segment with the highest transmission sequence number is connected to the starting vertex of the trajectory segment immediately following it according to the interconnection relationship. This connection is repeated until all trajectory segments in the path reconstruction graph are concatenated into a single sequence, resulting in a complete data flow trajectory.

[0059] The reconstructed path graph obtained after the connection is broken is structurally complete, carrying all process context nodes and their directed connections, but it is still stored in the form of a graph. A specific implementation performs an ordered traversal of the reconstructed path graph, converging each connected branch into a single sequence according to the data flow direction. In one implementation, the transmission sequence number attribute field of each vertex is retrieved from the vertex set of the reconstructed path graph, and all vertices are sorted in ascending order according to the transmission sequence number. The vertex at the top of the sorted list is taken as the starting vertex of the traversal process. The starting vertex corresponds to the process context node where the file access action occurs first in the complete data flow direction.

[0060] Specifically, starting from the starting vertex, the process proceeds sequentially in ascending order of the transmission sequence number. For the current vertex, the set of directed edges in the reconstructed path graph is queried, and the directed edges originating from the current vertex are extracted. The path continues along these directed edges to the target vertex, and the starting vertex, directed edges, and target vertex are written into a trajectory segment in the order of arrival. The current trajectory segment is closed when the target vertex no longer contains any subsequent directed edges originating from it.

[0061] It should be noted that long commands often involve multiple pipelines running in parallel during execution, resulting in multiple parallel branches in the path reconstruction graph. Without convergence processing, these trajectory segments are parallel to each other and cannot be presented as a single flow sequence.

[0062] In one possible implementation, for locations in the path reconstruction graph where multiple parallel branches exist, the kernel pipeline object index number registered in the directed edge attribute field of each branch is read one by one. The kernel pipeline object index number is a kernel allocation number shared by both ends of the same anonymous pipeline established during the transmission phase.

[0063] Specifically, the kernel pipeline object index numbers registered on the directed edges of each parallel branch are compared, and several parallel branches with the same kernel pipeline object index number are determined to be common-origin branches. Common-origin branches represent the set of edges mapped by the same anonymous pipeline in the path reconstruction graph. These common-origin branches are aggregated according to their transmission sequence number to the same target vertex, which serves as the convergence point of the common-origin branches. The convergence point, along with the common-origin branches, is registered in the inter-branch connectivity registration table. Further, for each trajectory segment and the inter-connectivity registration table, each trajectory segment is traversed. The last vertex of the trajectory segment with the earlier transmission sequence number is compared with the starting vertex of the trajectory segment with the immediately following transmission sequence number. If both vertices belong to the same convergence point in the inter-connectivity registration table, then the end of the preceding trajectory segment and the start of the following trajectory segment are aligned according to the convergence point, completing the head-to-tail connection. This head-to-tail connection operation is repeatedly performed, merging two trajectory segments with adjacent transmission sequence numbers into a longer segment each time. During merging, the connecting vertices of adjacent segments are checked. If a vertex exists in both segments, it is retained only once to avoid duplication. The merging is considered complete when only one trajectory segment remains in the path reconstruction graph and that segment contains all vertices. The merged single sequence is laid out in ascending order of transmission sequence number, covering all vertices and directed edges in the path reconstruction graph, forming a complete data flow trajectory. This trajectory records the execution process of compound commands, including pipe operators, on the target host, starting with the file read action at the source end, passing through multiple processes via pipes, and finally reaching the terminal file write action, completely reconstructing the flow of the byte stream between processes. The bastion host, acting as an intermediate layer for operation and maintenance auditing, is responsible for collecting and forwarding audit logs from the target host, and the monitoring system completes the above trajectory reconstruction on top of it.

[0064] S107. Evaluate the sensitive file reading path in the complete data flow trajectory, extract the network outgoing actions at the trajectory endpoint, classify the load scale of the network outgoing actions, and obtain the conclusion of violation identification.

[0065] Obtain the complete data flow trajectory obtained in the previous step, traverse the file read / write operation records registered at each vertex in the complete data flow trajectory, extract the absolute path corresponding to the entry whose operation direction is identified as read, match the absolute path with the character prefix of the pre-established sensitive file path library, if the absolute path falls within the coverage of any entry in the sensitive file path library, then register the vertex along with the absolute path as a sensitive read hit point, thus obtaining the sensitive file read path on the complete data flow trajectory. For the complete data flow trajectory containing the sensitive read hit point, the vertex with the largest transmission sequence number is located as the trajectory endpoint. From the file read / write operation records of the subshell isolated execution space to which the trajectory endpoint belongs, entries with the bound object type "socket" and the operation direction identifier "write" are selected. The destination Internet Protocol address carried by the entry is compared with a pre-established internal address range list. If the destination Internet Protocol address does not fall within the coverage of the internal address range list, the entry is identified as a network outbound action. The number of bytes carried in each write operation is accumulated to obtain the payload scale of the network outbound action. For the payload scale, the low-order threshold and high-order threshold registered in a pre-established payload scale classification threshold table are compared. If the payload scale does not exceed the low-order threshold, a low-order outbound identifier is assigned; if the payload scale is greater than the low-order threshold but does not exceed the high-order threshold, a medium-order outbound identifier is assigned; if the payload scale exceeds the high-order threshold, a high-order outbound identifier is assigned. The sensitive read hit point, along with the assigned outbound identifiers, is summarized as a violation identification conclusion.

[0066] The complete data flow trajectory only depicts the physical path of bytes from the source file to the terminal action, without providing a value judgment on the sensitivity and outbound volume of the content carried by the path. A specific implementation introduces sensitivity matching and outbound volume classification on top of the complete data flow trajectory, specifically for identifying unauthorized exports, downloads, and transfers. In one implementation, a sensitive file path library pre-includes the directory prefixes and file extensions of customer information, financial data, and business logs on the target host. For example, the data directory prefix for customer information, the report archive directory prefix for financial data, and the operation record directory prefix for business logs; each entry contains two fields: a string pattern and its corresponding sensitivity category. The complete data flow trajectory is organized in the form of a directed graph, where nodes represent the operation subject, edges represent data transmission relationships, and node attributes record file read / write operations. By traversing each node in the graph, entries with the operation direction identifier indicating "read in" are filtered, and their corresponding absolute paths are extracted. The absolute path is prefix-matched with each entry in the sensitive file path library. The comparison length is equal to the string pattern length in the library. If a match is successful, the node, along with its absolute path and sensitive category, is registered as a sensitive read hit point. These are then aggregated to form the sensitive file read path on the trajectory. After a sensitive file is read, it is necessary to check whether the endpoint of this data flow trajectory has performed an external push operation. For a complete data flow trajectory containing a sensitive read hit point, the transmission sequence number assigned to each vertex is used. The vertex with the largest transmission sequence number is directly located as the trajectory endpoint. This vertex is the last vertex to perform a byte write operation in the data flow direction. If the trajectory endpoint performs external push operations such as network sending or external storage writing, it is considered a violation.

[0067] In one possible implementation, entries are filtered from the file read / write operation records of the subshell isolated execution space to which the trajectory endpoint belongs. The monitoring system analyzes the bound object category field carried by each entry, retaining only entries with a value of "socket"; it also analyzes the operation direction identifier of the retained entries, retaining only entries with a value of "write". The destination Internet Protocol address carried by the remaining entries is compared with a pre-established list of internal address segments. The internal address segment list includes internal subnet segments protected by the bastion host, such as the operation and maintenance dedicated network segment, the office terminal network segment, and the internal business service network segment. If the destination Internet Protocol address does not fall within the coverage of any segment in the internal address segment list, the entry is identified as a network outbound action. The number of bytes written for each network outbound action is extracted sequentially and accumulated in chronological order to obtain the payload scale of the network outbound action. The payload scale reflects the total number of bytes pushed to the external network by the trajectory endpoint.

[0068] For example, a load size classification threshold table pre-registers two boundaries: a low threshold and a high threshold. The values ​​of the low threshold and the high threshold are based on statistical analysis of the single load volume distribution of historical normal operation and maintenance outbound traffic and known illegal outbound traffic: the upper quantile of the normal single outbound load distribution is taken as the low threshold to distinguish tentative outbound traffic, and the lower quantile of the batch export scenario load distribution is taken as the high threshold to distinguish large-scale transfers; one feasible value is a low threshold of 1 megabyte and a high threshold of 100 megabytes. The relative size of the load size with the low threshold and the high threshold is compared: if the load size does not exceed the low threshold, a low outbound identifier is assigned, corresponding to a small amount of information tentatively sent out; if the load size is greater than the low threshold but not greater than the high threshold, a medium outbound identifier is assigned, corresponding to a medium-scale data export; if the load size exceeds the high threshold, a high outbound identifier is assigned, corresponding to a large-scale data transfer. Finally, the absolute path carried by the sensitive read hit point, its corresponding sensitive category, the process context node identity of the trajectory endpoint, the destination Internet Protocol address of the network outbound action, and the assigned outbound identifier are assembled into a violation identification conclusion according to the order of the complete data flow trajectory. This violation identification conclusion depicts the complete violation chain from sensitive file reading to external network push, and carries the bastion host's output of its judgment on sensitive data leakage.

[0069] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for identifying and controlling violations of bastion host operation instructions, characterized in that, The method includes: Obtain the surface characters input from the bastion host terminal, and extract the combination structure of redirection symbols and pipe symbols composed of each instruction through lexical analysis; Based on the symbol combination structure, identify the pipe passing nodes formed by pipe symbol connections, trace the process creation system calls triggered when the pipe passing nodes are executed, identify multiple mutually isolated subshell isolated execution spaces created by them and carrying the split execution of long commands, as well as file read and write operations within each subshell isolated execution space; Analyze the data entry and exit actions generated by file read and write operations in each subshell isolated execution space, determine the local audit logs formed by the monitoring system recording each item, and analyze the transmission links that cross the isolation barrier of the subshell isolated execution space after connecting them in time sequence. Extract the input and output file descriptors in the transmission process, compare the entry and exit points of the subshell isolated execution space mapped by the input and output file descriptors, and obtain the process context nodes associated with the entry and exit points; Extract the process call chain and file access sequence between the context nodes of each process, determine the transmission order of the process call chain, and construct a path reconstruction graph. Connect the transmission links where the local audit logs are broken at the isolation barrier in the path reconstruction graph. Traverse the path reconstruction graph according to the process call chain, and connect the interconnected branches to obtain the complete data flow trajectory. Evaluate the sensitive file reading path in the complete data flow trajectory, extract the network outgoing actions at the trajectory endpoint, classify the payload scale of the network outgoing actions, and obtain the conclusion of violation identification.

2. The method for identifying and controlling violations of bastion host operation instructions according to claim 1, characterized in that, The process of obtaining the surface characters input from the bastion host terminal and extracting the combination structure of redirection symbols and pipe symbols composed of each instruction through lexical analysis includes: Obtain the original character sequence entered by the operation and maintenance personnel in the bastion host terminal session channel, extract the echo control character and cursor movement character from the original character sequence, mark the boundaries of the paragraphs enclosed in quotation marks, and distinguish the literal text inside the quotation marks from the command text outside the quotation marks. The command text is scanned character by character using a lexical analysis method and matched against a preset operator vocabulary. Two types of symbolic words, namely redirection symbols and pipe symbols, are identified and recorded along with their character offset positions into the lexical unit sequence. Based on the order of appearance of symbolic lexical units in the lexical unit sequence, the command words and parameter items between two adjacent symbolic lexical units are merged into instruction fragments. Each instruction fragment is then concatenated with the redirection symbols and pipe symbols preceding and following it according to their appearance positions. The connection relationship between the output end of the left instruction fragment and the input end of the right instruction fragment is marked in the concatenation result to obtain the symbol combination structure.

3. The method for identifying and controlling violations of bastion host operation instructions according to claim 1, characterized in that, The process involves identifying pipe connection nodes formed by symbolic linking based on symbolic combination structures, tracing process creation system calls triggered when these nodes are executed, identifying multiple isolated subshell execution spaces created by these nodes and carrying out the split execution of long commands, and file read / write operations within each subshell execution space, including: The positions marked as pipeline symbols by word categories are selected from the symbol combination structure and arranged in order of appearance to obtain the pipeline transmission node sequence; The kernel audit subsystem is used to mount the hook point of the process creation system call, capture the child process identifier and parent process identifier returned by each process creation, organize the capture results into a process derivation tree, and identify several child processes continuously derived by the same parent process under the execution trigger of the pipe passing node as the subshell isolated execution space; For each of the subshell isolated execution spaces, corresponding hook points are attached to the file open, read, and write system calls issued during operation. The file descriptor and the bound file path are captured, and the file descriptor, along with its process identifier, read / write direction, and operation time, are registered as file read / write operation records.

4. The method for identifying and controlling violations of bastion host operation instructions according to claim 1, characterized in that, The process involves analyzing the data inflow and outflow actions generated by file read / write operations within each subshell's isolated execution space, identifying the local audit logs recorded by the monitoring system, and analyzing the transmission links that cross the isolation barriers of the subshell's isolated execution space after sequential connection. This includes: The file descriptor integer number, absolute path, operation direction identifier, and operation time are parsed from the read and write operation records of each file, and the local audit log is obtained by grouping them according to the process identifier. The actions in the local audit log are arranged in ascending order of operation time. Adjacent action pairs belonging to different subshell isolated execution spaces are traversed. It is determined whether the difference in the order of operation time between the data outgoing action in the preceding space and the data entering action in the following space falls into a preset connection time window. Those that fall into the window are marked as belonging to the same transmission link that crosses the isolation barrier. The outgoing direction block and the incoming direction block are connected in series to obtain the transmission link sequence.

5. The method for identifying and controlling violations of bastion host operation instructions according to claim 1, characterized in that, The process extracts input and output file descriptors in the transmission process, compares the entry and exit points of the subshell isolated execution space mapped by the input and output file descriptors, and obtains the process context nodes associated with the entry and exit points, including: Traverse the sequence of transmission links, read the integer number of the file descriptor and the process identifier of the process from the outgoing and incoming blocks, and obtain the input and output file descriptor tuples; In the kernel-side file descriptor table held by the subshell isolated execution space corresponding to the process identifier of the outgoing block, the write end of the kernel pipe object bound to the output file descriptor is looked up as the space exit, and the read end of the kernel pipe object bound to the input file descriptor is looked up as the space entry in the corresponding table of the incoming block. Compare the kernel pipe object index numbers pointed to by the space exit and space entrance. If they point to the same index number, it is determined that a docking has been formed. Extract the executable image path, command line argument string, effective user identifier, and parent process identifier corresponding to the process identifiers of both ends of the docking, and assemble them into process context nodes associated with the entry and exit points according to the index number of the docked kernel pipe object.

6. The method for identifying and controlling violations of bastion host operation instructions according to claim 1, characterized in that, The process call chain and file access sequence between the context nodes of each process are extracted to determine the transmission order of the process call chain and to construct a path reconstruction graph. This path reconstruction graph connects the transmission links where the local audit logs are broken at the isolation barrier, including: For each process context node, trace the process identifier to the parent process identifier and then back to the command interpreter main process identifier to obtain the parent-child derivation chain, and then merge the common process identifier nodes into a process call chain. Arrange the access actions in ascending order of access time to obtain the access sequence string, and assign them a transmission sequence number from small to large. A path reconstruction graph is constructed using process context nodes as vertices and kernel pipe object index numbers that connect adjacent vertices by transmission sequence numbers as edges. For positions lacking edge connections, matching index numbers are retrieved from the self-transmission link sequence and added.

7. The method for identifying and controlling violations of bastion host operation instructions according to claim 1, characterized in that, The method of traversing the path reconstruction graph according to the process call chain order, and connecting the interconnected branches to obtain the complete data flow trajectory, includes: From the set of vertices in the path reconstruction graph, select the vertex with the smallest transmission sequence number as the starting vertex, and read the directed edges in ascending order of transmission sequence number to form a trajectory segment. For positions with multiple parallel branches, read the kernel pipeline object index number from the directed edge attribute field, determine branches with the same index number as common source branches and aggregate them to the same vertex, and repeatedly perform the head-to-tail connection until all trajectory segments are concatenated into a single sequence.

8. The method for identifying and controlling violations of bastion host operation instructions according to claim 1, characterized in that, The evaluation process involves identifying sensitive file reading paths within the complete data flow trajectory, extracting network outbound actions at the trajectory endpoints, classifying the payload scale of these outbound actions, and obtaining conclusions regarding violation identification, including: Traverse the file read and write operation records of each vertex in the complete data flow trajectory, extract the operation direction identifier as the absolute path of the read, and perform character prefix matching with the pre-established sensitive file path library. Those that fall within the coverage area are registered as sensitive read hit points. The vertex with the largest sequence number is used as the trajectory endpoint. Entries with the binding object category of socket and the operation direction identifier of write are filtered. The destination Internet Protocol address is compared with the list of internal address ranges. Those not included are identified as network outbound actions and the number of bytes is accumulated to obtain the payload scale. By comparing the low-level and high-level thresholds in the load scale classification threshold table, low-level, medium-level, or high-level outward identification labels are assigned respectively, and the results are summarized as the conclusion of violation identification.

9. The method for identifying and controlling violations of bastion host operation instructions according to claim 2, characterized in that, The redirection symbols cover input redirection, output redirection, append redirection, and file descriptor copy variants, and the pipe symbols cover unidirectional pipes and error stream merging pipes.

10. The method for identifying and controlling violations of bastion host operation instructions according to claim 3, characterized in that, The process involves mapping each of the subshell isolated execution spaces to the position points in the pipeline transmission node sequence in the order of derivation, thereby determining the instruction fragment carried by each subshell isolated execution space.