An API semantic perception-based humanoid robot safety control method and system
Patent Information
- Application Number
- CN202611063121.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-17
AI Technical Summary
[0004]本发明旨在解决基于API调用行为的人形机器人安全控制技术对混淆攻击脆弱的问题
[0060]1) 现有技术严重依赖机器人API调用的名称、顺序或图结构等表层特征,这些特征极易通过API重命名、指令重排或插入“垃圾”API 调用等代码混淆手段进行改变,从而导致安全控制模型失效,引发机器人失控风险。本发明通过对API的功能语义进行建模,直接理解API调用背后的控制意图,而非其表层名称。由于API的底层功能描述是相对稳定的,因此无论恶意控制程序如何变换其代码形式,本方法都能识别其内在控制行为的一致性,从而具备对代码混淆攻击天然的鲁棒性。
Smart Images

Figure CN122571587B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of robot safety and artificial intelligence technology, specifically to a humanoid robot safety control method and system based on API semantic perception. Background Technology
[0002] With the rapid development of humanoid robot technology, its applications in home services, industrial production, and public services are becoming increasingly widespread. The core motion control, sensor data interaction, and external command execution of humanoid robots all rely on Application Programming Interfaces (APIs). The integration of third-party control programs and functional plugins further enriches the application scenarios of robots. Among these methods, analyzing robot API call sequences to identify malicious control behaviors (such as tampering with motion trajectories, stealing sensor data, and hijacking interaction commands) is a mainstream approach to ensuring the safe operation of robots. These methods typically construct the API call sequences as a graph structure and use models such as Graph Neural Networks (GNNs) to learn patterns within them, distinguishing between benign and malicious control programs or plugins.
[0003] However, existing technologies generally rely on superficial features such as API names, call order, or local adjacency relationships. Adversaries can significantly weaken the effectiveness of these models by inserting irrelevant API calls, rearranging call order, or using wrapper or alias interfaces to alter the sequence and graph structure. This allows malicious control programs or plugins to bypass detection, leading to security risks such as robot malfunction and data breaches. The root cause lies in the models' lack of representation of the functional intent behind API calls: they "know which API was called," but they don't understand the actual function that API performs in robot control. Therefore, when attackers change the superficial form of the API without altering the core control intent, the model struggles to recognize its inherent behavioral consistency. There is an urgent need for a humanoid robot safety control method that can penetrate code obfuscation and directly model from the functional semantic level. Summary of the Invention
[0004] This invention aims to address the vulnerability of humanoid robot security control technology based on API call behavior to obfuscation attacks. To this end, a functional semantic awareness scheme is proposed: First, the functional descriptions of robot API calls are obtained from authoritative technical documentation, and functional semantic embeddings are obtained using a pre-trained language model. Then, on a semantic behavior graph containing concurrency and temporal relationships, a Graph Attention Network (GAT) is used to learn the combination patterns of API calls in complex control scenarios. Finally, multi-instance learning is employed to perform voting aggregation at the thread subgraph granularity, achieving accurate identification of malicious control programs or plugins and ensuring the safe operation of the humanoid robot.
[0005] To achieve the above objectives, the safety control method provided by the present invention includes the following steps:
[0006] A safety control method for humanoid robots based on API semantic awareness includes the following steps:
[0007] S1: Construct a training and evaluation dataset containing benign and malicious samples, and deploy a dynamic analysis environment based on container technology; run the samples in the dataset in the dynamic analysis environment, capture dynamic behavior data through system call monitoring tools, and extract API call sequences sorted by time and their process or thread information;
[0008] S2: For the set of target APIs in the API call sequence extracted by S1, obtain the functional description text corresponding to each API from at least one authoritative technical document source in the robotics field, and build an API functional description corpus;
[0009] S3: Using a general pre-trained bidirectional language model based on the Transformer architecture, and using the API function description corpus constructed in step S2 as training data, the pre-trained bidirectional language model is adaptively fine-tuned in the robotics domain. Through fine-tuning, the model can output a high-dimensional dense vector representing the API control function intent; the high-dimensional dense vector serves as the functional semantic embedding of the corresponding robot API call.
[0010] S4: For the robot control program or plugin to be analyzed, capture the API call sequence generated during its runtime, and construct a multi-layer semantic behavior heterogeneous graph that characterizes the concurrent behavior and temporal control behavior of the robot control program or plugin; the multi-layer semantic behavior heterogeneous graph is layered by thread, and the initial features of its API call nodes are initialized by the functional semantic embedding obtained in step S3.
[0011] S5: Input the semantic behavior heterogeneous graph constructed in step S4 into the graph attention network for end-to-end feature learning; calculate the correlation between nodes in the graph through a multi-layer, multi-head attention mechanism, and perform pooling operations on the subgraphs corresponding to each thread in the graph to obtain a set of thread subgraph representation vectors with fixed dimensions.
[0012] S6: Input the set of thread subgraph representation vectors generated in step S5 into the aggregation classifier based on multi-instance learning. The aggregation classifier performs independent malicious probability prediction on each thread subgraph representation vector. Then, the aggregation function fuses all independent prediction results to obtain the sample-level malicious probability of the robot control program or plugin. The cross-entropy loss function is used to optimize the learnable parameters of the pre-trained bidirectional language model, graph attention network and aggregation classifier end-to-end to complete model training and sample determination.
[0013] Furthermore, step S1 specifically includes:
[0014] S101: Construct training and evaluation datasets. Benign samples are obtained from the official control programs of robot manufacturers and ROS / ROS2 open source community certified plugins; malicious samples are obtained from artificially constructed malicious control plugins and publicly available robot malicious sample sharing platforms.
[0015] S102: Deploy the LXC container dynamic analysis platform, which integrates the Gazebo robot simulation environment to simulate the real operation scenario of robot control samples;
[0016] Specifically,
[0017] S1021: Container-isolated execution environment: Multiple isolated container instances are built based on LXC for independently running the robot control samples to be analyzed;
[0018] S1022: Robot Simulation Interaction Module: Integrates the Gazebo robot simulation environment inside each LXC container, communicates with the sample running inside the container through the ROS / ROS2 interface, and simulates the control behavior of the sample on a real robot.
[0019] S1023: System Behavior Monitoring Module: Configure the strace monitoring tool at the host level to trace the process of each LXC container and capture its system call sequence; at the same time, configure the rosbag tool to capture topic messages published by samples in the container through ROS / ROS2;
[0020] S1024: Centralized Log Collection Module: Used to collect log data generated by strace and rosbag from each container after the sample runs out of time or end, and archive it by sample ID.
[0021] S103: Deploy benign and malicious samples one by one into independent LXC containers for execution. Use the strace tool to execute preset monitoring commands and combine it with the rosbag tool to monitor the runtime behavior of the samples, capture API call sequences, process IDs (PIDs), thread IDs (TIDs), and timestamps; set the maximum runtime for each sample, and forcibly terminate the container and collect runtime logs after the timeout.
[0022] S104: Analyze the runtime logs generated by the strace and rosbag tools, extract the API call sequence of each sample arranged in chronological order, and associate it with the corresponding process and thread identifiers to form structured API call sequence data, which serves as the basis for subsequent semantic analysis and graph modeling.
[0023] Furthermore, step S2 specifically includes:
[0024] S201: For the API call sequence compiled in S104, retrieve the functional description text of each API from authoritative documents in the robotics field, including ROS / ROS2 official documents, robot hardware SDK manuals, and manufacturer API technical specifications.
[0025] S202: Use automated scripts to crawl the functional description text of each API in authoritative technical documents. For APIs whose functional descriptions are not found in the documents, manually search and supplement their functional description texts. Store all API functional description texts in a centralized manner to form an API functional description corpus.
[0026] Furthermore, step S3 specifically includes:
[0027] S301: The BERT (Bidirectional Encoder Representations from Transformers) model is used as the semantic extraction model. The BERT model learns contextual semantics through the Masked Language Model (MLM) pre-training task. The BERT model is initialized with general pre-trained weights, and then the API function description corpus constructed in step S2 is used as training data to perform adaptive fine-tuning of the BERT model in the robotics domain so that the model can adapt to the semantic representation requirements of robot API functions.
[0028] S302: Input the functional description text of each API into the fine-tuned BERT model for encoding, and generate a high-dimensional dense functional semantic embedding vector for the corresponding API call. The vector is used to characterize the core function of the API call (such as "robotic arm movement" and "camera data reading") semantics, so as to achieve robust recognition of API renaming obfuscation methods.
[0029] Furthermore, step S4 specifically includes:
[0030] S401: Construct a directed heterogeneous graph as a multi-layer semantic behavior heterogeneous graph. The directed heterogeneous graph includes three types of nodes: the root node representing the file ID of the robot control program or plug-in, the intermediate node representing the thread TID of the concurrent execution flow, and the leaf node representing the API call of the actual control operation.
[0031] S402: Construct a graph topology based on three types of nodes. The construction logic includes: establishing directed edges from the root node of the file ID to all intermediate nodes of the thread TID derived from it to realize concurrency relationship modeling; establishing directed edges from the intermediate node of each thread TID to all API call leaf nodes called within its corresponding thread to realize dependency relationship modeling; and establishing directed edges between adjacent API call leaf nodes in the same thread according to the API call time order to realize temporal relationship modeling.
[0032] S403: Use the functional semantic embedding vector generated in step S3 as the initial feature vector of the corresponding API call leaf node to complete the construction of the semantic behavior heterogeneous graph.
[0033] Furthermore, step S5 specifically includes:
[0034] S501: Input the semantic behavior heterogeneous graph constructed in step S4 into the graph attention network encoder. The graph attention network encoder includes at least two GATConv layers, and each GATConv layer adopts a multi-head attention mechanism. The graph attention network encoder learns the complex relationship between nodes by dynamically allocating attention weights between nodes, and at the same time combines the API call frequency as the edge weight to assist in the calculation, and updates the initial features of each node to higher-order features that fuse neighborhood information.
[0035] The attention coefficient e between node i and its neighbor node j satisfies:
[0036] ;
[0037] in, and Let be the feature vectors of node i and node j, respectively, and W be the learnable weight matrix shared by all nodes. It is an attention function;
[0038] The attention coefficients e are normalized using the Softmax function to obtain the final attention weights. :
[0039] ;
[0040] in, It is the set of all neighboring nodes of node i, || represents vector concatenation, a is the learnable weight vector in the attention mechanism, and LeakyReLU is the activation function;
[0041] S502: Employs a multi-instance learning strategy to decompose the semantic behavior heterogeneous graph into multiple thread-level subgraphs. Each thread-level subgraph consists of thread TID intermediate nodes and their associated API call leaf nodes. Pooling is performed on each thread-level subgraph to aggregate the high-order features of all nodes within the subgraph into a fixed-dimensional subgraph representation vector. ,in The control behavior representation of the i-th thread is used to form a set of thread subgraph representation vectors.
[0042] Furthermore, step S6 specifically includes:
[0043] S601: Convert the subgraph representation vector set generated in step S502 into a single vector set. Input a voting-based aggregation classifier; use a weight-sharing sub-classifier within the aggregation classifier to independently predict the representation vector of each thread subgraph, thus obtaining the malicious prediction probability of each thread subgraph. , ∈[0,1]; This is considered malicious voting in the subgraph;
[0044] S602: Employs a max-pooling aggregation function to calculate the malicious prediction probability for all thread subgraphs. By merging the data, the final probability of malicious activity of the robot control program or plugin can be obtained. , n is the number of thread subgraphs;
[0045] The preferred rule of this invention, "if any subgraph is malicious, then the sample is determined to be malicious," ensures that as long as the model detects a high probability of malice on any thread subgraph (such as the existence of control behaviors such as "reading camera data + sending to overseas IPs"), the overall probability of malice of the sample will increase accordingly, thus avoiding the omission of hidden malicious control behaviors.
[0046] S603: Predict the final probability of the aggregated samples Substituting the true label (benign or malicious) of the sample into the binary cross-entropy loss function:
[0047] ;
[0048] In the real labels, y=1 indicates malicious intent, and y=0 indicates benign intent.
[0049] The gradient of the loss function with respect to all learnable parameters of the model is calculated using the backpropagation algorithm. The parameters are then updated using the optimizer to minimize the loss function. This optimization process is repeated until the model performance converges.
[0050] A humanoid robot safety control system based on API semantic awareness, used to implement the aforementioned humanoid robot safety control method based on API semantic awareness, the system comprising:
[0051] The dynamic behavior capture module is used to build training and evaluation datasets, deploy a dynamic analysis environment based on LXC containers, monitor the runtime behavior of robot control samples through strace, rosbag and Gazebo robot simulation environment, capture API call sequences, process identifiers, thread identifiers and timestamps, and output structured API call sequence data.
[0052] The functional semantic embedding module is used to obtain API functional description text from authoritative technical documents in the robotics field, build an API functional description corpus, and encode the API functional description text into high-dimensional dense functional semantic embedding vectors through a BERT model that is adaptively fine-tuned in the robotics field.
[0053] The semantic behavior graph construction module is used to construct a directed heterogeneous graph containing file ID root nodes, thread TID intermediate nodes, and API call leaf nodes based on structured API call sequence data. The functional semantic embedding vector is used as the initial feature of the API call leaf nodes to complete the construction of a multi-layer semantic behavior heterogeneous graph.
[0054] The graph attention feature learning module receives the semantic behavior heterogeneous graph output by the semantic behavior graph construction module. Through a graph attention network containing multiple GATConv layers and a multi-head attention mechanism, it learns high-order relationships between nodes, performs pooling operations on thread-level subgraphs, and outputs a set of thread subgraph representation vectors.
[0055] The classification prediction and model optimization module is used to input the set of thread subgraph representation vectors output by the graph attention feature learning module into the aggregation classifier to complete the prediction of the malicious probability of the sample. Through the cross-entropy loss function and backpropagation algorithm, it performs end-to-end optimization on all learnable parameters in the system to realize the benign and malicious determination of the robot control program or plug-in.
[0056] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the API-based semantic awareness-based humanoid robot safety control method.
[0057] The design concept of this invention is as follows:
[0058] This method runs robot control programs or plugin samples in a lightweight container environment (integrated with the Gazebo robot simulation environment). It collects dynamic behavioral data such as robot API call sequences, process and thread information, and timestamps using a system call monitoring tool (strace combined with rosbag). API function descriptions are extracted from authoritative robot documentation (ROS / ROS2 official documentation, hardware SDK manuals, and manufacturer API specifications), and functional semantic embedding vectors are generated using a Transformer model (such as BERT, adaptively fine-tuned for the robotics domain). A multi-layered semantic behavior heterogeneous graph is constructed with file IDs (corresponding to control programs or plugins), thread IDs, and robot API calls as nodes, integrating semantic features into the graph structure. A multi-layered multi-head attention mechanism graph attention network is used for feature learning to extract key control behavior patterns and generate global feature representations. Finally, a multi-instance-based aggregation classifier is used to identify malicious robot control programs or plugins. This method can understand robot API calls at the functional semantic level, effectively improving the robustness and recognition accuracy of the security control model against obfuscation and variant attacks, providing reliable protection for the safe operation of humanoid robots.
[0059] The advantages and benefits of this invention compared to existing technologies are as follows:
[0060] 1) Existing technologies heavily rely on surface-level features such as the names, sequences, or graph structures of robot API calls. These features are easily altered through code obfuscation techniques such as API renaming, instruction rearrangement, or the insertion of "junk" API calls, leading to the failure of the safety control model and the risk of robot loss of control. This invention models the functional semantics of APIs to directly understand the control intent behind API calls, rather than their surface names. Since the underlying functional description of APIs is relatively stable, this method can identify the consistency of its inherent control behavior regardless of how malicious control programs change their code form, thus possessing inherent robustness against code obfuscation attacks.
[0061] 2) Traditional methods know which API the robot calls, but not what function that API actually performs in control. This invention innovatively utilizes a deep language model based on the Transformer architecture (such as BERT) to transform the functional descriptions of robot APIs in authoritative technical documents into high-dimensional, dense "functional semantic embedding" vectors. This vector encodes the core control functions of the API, enabling the model to understand the robot's control behavior at a functional level, thus eliminating dependence on API names.
[0062] 3) This invention not only endows individual robot APIs with functional semantics but also constructs a heterogeneous graph capable of simultaneously representing concurrent and temporal control behaviors. Based on this, GAT (Generic Attention Graph) is used to learn the deep dependencies between API calls. Furthermore, this invention innovatively incorporates multi-instance learning, treating each concurrent thread as an independent subgraph for evaluation. GAT's multi-head attention mechanism can focus on key control behavior combinations within each subgraph, ultimately accurately locating key malicious sub-behaviors hidden within a large number of benign control threads through a voting aggregation strategy of "if any subgraph is malicious, then the sample is malicious." Compared to the traditional method of fuzzily aggregating all control behaviors into a single global vector, this significantly improves the sensitivity of safety control and the ability to detect hidden malicious control behaviors, providing a more reliable guarantee for the safe operation of human robots.
[0063] This invention utilizes deep language models and graph neural networks to model the functional semantics of API call sequences for humanoid robots, thereby realizing a method, system, and computer-readable storage medium for humanoid robot safety control with strong robustness against code obfuscation. Attached Figure Description
[0064] Figure 1 This is a schematic diagram of the overall process of the humanoid robot safety control method provided in an embodiment of the present invention;
[0065] Figure 2 This is a schematic diagram illustrating the use of the BERT model to process robot API description text in an embodiment of the present invention;
[0066] Figure 3 This is a schematic diagram illustrating the construction of a multi-layer semantic behavior heterogeneous graph in an embodiment of the present invention;
[0067] Figure 4 This is a schematic diagram of subgraph voting and aggregation classification based on multi-instance learning in an embodiment of the present invention;
[0068] Figure 5 The execution diagram of the core functional modules provided in the embodiments of the present invention. Detailed Implementation
[0069] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0070] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0071] I. Experimental Environment and Basic Configuration
[0072] 1. Hardware Environment
[0073] Server configuration: Intel(R) Core(TM) i9-14900HX @ 2.20GHz (1 core, 24 cores and 32 threads).
[0074] Memory: 16GB DDR5 5600MHz;
[0075] GPU: NVIDIA GeForce RTX 4060 (8GB GDDR6 video memory);
[0076] Storage: 1TB PCIe 4.0 SSD (with one reserved M.2 2280 PCIe 4.0 expansion slot, supporting dual hard drive expansion)
[0077] 2. Software Environment
[0078] Operating system: Ubuntu 20.04 LTS (64-bit);
[0079] Container technology: LXC 4.0.12 (lightweight container engine);
[0080] Robot simulation environment: Gazebo 11.10.2 + ROS Noetic Ninjemys (robot operating system);
[0081] Monitoring tools: strace 5.11 (system call monitoring), rosbag 1.15.11 (ROS topic / service capture);
[0082] Deep learning framework: PyTorch 1.12.1 + TorchGeometric 2.2.0 (graph neural network library);
[0083] Other dependencies: Python 3.8.10, NumPy 1.21.6, Scikit-learn 1.0.2 (model evaluation).
[0084] II. Detailed Implementation of Methods and Steps
[0085] See attached document Figure 1 The overall flow of the humanoid robot safety control method provided in this embodiment is as follows:
[0086] S1: Build a dataset containing benign and malicious samples, and deploy a dynamic analysis environment based on container technology; run samples in this environment, capture dynamic behavior data through system call monitoring tools, and extract API call sequences sorted by time and their process or thread information.
[0087] S2: For the target API set, obtain the functional description text corresponding to each API from one or more authoritative technical documentation sources, and build an API functional description corpus.
[0088] S3: Using a pre-trained bidirectional language model based on the Transformer architecture, the text describing each API function in the corpus is processed into a high-dimensional, dense feature vector, which serves as the functional semantic embedding of this API.
[0089] S4: For the robot control program or plugin to be analyzed, capture the API call sequence generated during its runtime and construct a multi-layered semantic behavior heterogeneous graph that can characterize its concurrent and temporal control behavior. The behavior graph is constructed hierarchically according to thread ID, and its initial node features are represented by the corresponding functional semantic embedding vectors generated in step S3.
[0090] S5: Input the sample semantic behavior graph constructed in step S4 into a graph attention network for end-to-end feature learning. The graph attention network learns the deep dependencies between nodes through a multi-layer, multi-head attention mechanism, and performs pooling operations on each subgraph representing concurrent threads in the graph, aggregating the information of each subgraph into a subgraph representation vector, and finally generating a set of fixed-dimensional subgraph representation vectors.
[0091] S6: Input the set of subgraph representation vectors generated in step S5 into an aggregation classifier based on multi-instance learning to obtain the judgment result. The classifier first makes independent predictions (votes) for each subgraph vector, then obtains the final malicious probability of the control program or plugin through an aggregation function, and uses the cross-entropy loss function to perform end-to-end optimization training on all learnable parameters of the entire model.
[0092] See attached document Figure 1 The specific steps of the humanoid robot safety control method are as follows:
[0093] S1.1: Dataset Construction. A comprehensive dataset for model training and evaluation is constructed. Beneficial samples (2000 in total) are sourced from official control programs of robot manufacturers and ROS / ROS2 open-source community certified plugins, specifically including: UBTECH Walker X official motion control program (500); Boston Dynamics Spot SDK certified plugin (300); MoveIt! robotic arm control plugin (400), RViz visualization plugin (300), and voice interaction plugin (500).
[0094] Labeling criteria: 100% benign as determined by VirusTotal platform (tested in March 2024), with no sensitive API calls (such as APIs for file tampering or network data theft).
[0095] The malicious samples (2000 in total) are derived from artificially constructed and publicly available robot malicious sample sharing platforms. Specifically, they include: artificially constructed malicious plugins (800): plugins that tamper with the movement trajectory of robotic arms (300), plugins that steal camera / LiDAR data (300), and plugins that hijack ROS control commands (200); and publicly available malicious sample sharing libraries (1200), including ransomware plugins and data leakage plugins targeting humanoid robots.
[0096] Labeling criteria: The behavior is clearly malicious (such as sending sensor data to overseas IPs or forcibly modifying joint angles beyond the safe range), and has been verified by manual reproduction.
[0097] Sample preprocessing rules: Samples with a runtime of <10 seconds (no valid API calls) or >120 seconds (infinite loop) were removed, totaling 127 samples, resulting in 3873 valid samples (1920 benign and 1953 malicious). The data was divided into a 6:2:2 ratio for the training set (2324 samples), validation set (775 samples), and test set (774 samples) to ensure a consistent class distribution across all datasets (malicious samples ≈ 50.4%).
[0098] S1.2: Deployment of the LXC container dynamic analysis platform. Deploy an LXC container dynamic analysis platform dedicated to the analysis of malicious samples from bots. The platform consists of the following components:
[0099] S1.2.1: Container-isolated execution environment: Multiple isolated container instances are built based on LXC, with each container allocated 2 CPU cores, 4GB memory, and 20GB storage; the container network adopts bridged mode, allowing access only to the local simulation environment (external network access is prohibited to avoid the spread of malicious samples).
[0100] S1.2.2: Robot Simulation Interaction Module: This module integrates the Gazebo robot simulation environment (version 11) within each LXC container, loading a Walker X humanoid robot model (containing 28 joints, an RGB-D camera, and a LiDAR). It communicates with the sample running within the container via ROS topics, simulating the sample's control behavior on a real robot. Configured topics include: / cmd_vel (motion control), / joint_state (joint state), and / camera / rgb / image_raw (image data).
[0101] S1.2.3: System Behavior Monitoring Module: Configure the strace monitoring tool at the host level to trace the process of each LXC container and execute the command: strace -f -tt -T -s 1024 -o . / logs / sample_${ID}_strace.log ${sample_path}; At the same time, configure the rosbag tool to capture ROS topic messages published by samples within the container.
[0102] S1.2.4: Centralized Log Collection Module: Set the maximum runtime for each sample to 120 seconds. After the timeout, the shell script will automatically execute lxc-stop -n ${container_name}. Logs are archived according to "sample ID + tool type", such as sample_0012_strace.log and sample_0012_rosbag.bag.
[0103] S1.3: Sample Execution and Monitoring. Valid samples were deployed one by one to independent LXC containers for execution. For each sample, the sample program was executed immediately after the container started, and strace and rosbag monitoring were initiated simultaneously. The sample's process ID (PID), thread ID (TID), timestamp, and API call parameters (such as file path, network address, etc.) were recorded. A total of 3873 valid logs were collected, with an average of approximately 5000 API call records per sample.
[0104] S1.4 Log Parsing and Structuring: A custom Python script is used to parse strace and rosbag logs, extracting the API call sequence for each sample in chronological order and associating it with the corresponding PID / TID. Taking sample ID 0012 (malicious plugin "camera data theft") as an example, the parsing results are shown in Table 1:
[0105] Table 1 Analysis Results ;
[0106] Meanwhile, the ROS API calls obtained from rosbag are shown in Table 2:
[0107] Table 2 Parsing Results Call ;
[0108] Structured API call sequence (associated with PID / TID):
[0109] Thread 1 (PID=2500, TID=2501): openat() → read() → close()
[0110] Thread 2 (PID=2500, TID=2502): socket() → connect() → sendto() → close()
[0111] See attached document Figure 2 The specific steps for constructing the API function description corpus in S2 include:
[0112] S2.1: API Search. For all independent APIs appearing in S104 (a total of 1287 different robot-related APIs, including 832 system APIs and 455 ROS-specific APIs), their functional descriptions were retrieved one by one from authoritative technical documentation sources. These sources included: the official ROS Noetic documentation; the Walker X SDK 2.0 manual; and the Intel RealSense camera API technical specifications. For example, for the API `openat()`, the description from the documentation was: "Open a file relative to a directory file descriptor, used to access hardware devices or files."
[0113] S2.2: Corpus Formation. API descriptions were crawled from the aforementioned document sources using an automated web crawler script (Python + Scrapy framework). For APIs with missing documentation (37 in total), developers manually searched and supplemented the descriptions. The final result was a corpus containing 1287 API function descriptions, with each description averaging 25 words in length.
[0114] See attached document Figure 2 The specific steps of generating the API functional semantic embedding in S3 include:
[0115] S3.1: BERT model fine-tuning. A pre-trained BERT model (bert-base-uncased, 12-layer Transformer, 768-dimensional hidden layers, 12 attention heads) was used as the semantic extraction model. The API function description corpus constructed in S2.2 was used as training data to perform robotics domain-adaptive fine-tuning of BERT. The fine-tuning parameters are shown in Table 3.
[0116] Table 3 Fine-tuning parameters ;
[0117] S3.2: Semantic Embedding Generation. The functional description text of each API is input into the fine-tuned BERT model, and the 768-dimensional output vector of the last layer [CLS] token is taken as the functional semantic embedding vector for that API. For example, the embedding vectors (first 10 dimensions) of some APIs are shown in Table 4:
[0118] Table 4 Embedding Vectors ;
[0119] Referring to Figure 3, the specific steps for constructing the multi-layer semantic behavior heterogeneous graph S4 include:
[0120] S4.1: Node definition, constructing a directed heterogeneous graph for each sample to be analyzed, containing three types of nodes:
[0121] Root node: A file ID node fid, representing the entire control program / plugin.
[0122] Intermediate nodes: Each thread corresponds to a thread ID node tid_i (i=1,…,n), where n is the number of threads in the sample.
[0123] Leaf node: Each API call corresponds to an API node api_j, where j is the call sequence number.
[0124] Taking sample ID 0012 (malicious plugin) as an example, the node information is shown in Table 5:
[0125] Table 5 Node Information ;
[0126] S4.2: The logic for constructing directed edges is as follows:
[0127] Concurrent relationship modeling: Add directed edges from root node F0012 to thread nodes T2501 and T2502.
[0128] Subordination modeling: Add directed edges from each thread node to the API node it calls, i.e., T2501→A001, T2501→A002, T2502→A003, T2502→A004.
[0129] Temporal relationship modeling: Within the same thread, directed edges are added between adjacent API nodes according to the API call time order, i.e., A001→A002 (thread 2501), A003→A004 (thread 2502). The edge weight can be set to the API call frequency (set to 1 in this embodiment, representing that the temporal relationship occurs once).
[0130] S4.3: Node feature initialization. Assign the semantic embedding vectors generated in step S3 to the corresponding API leaf nodes. The initial features of the root node and intermediate nodes are set to learnable embedding vectors (768 dimensions), which are updated during training. This completes the construction of the semantic behavior heterogeneous graph.
[0131] See attached document Figure 4 The specific steps of subgraph voting and aggregation classification in multi-instance learning include:
[0132] S5.1: GAT Encoder, which inputs the heterogeneous graph described above into the Graph Attention Network (GAT) encoder. The encoder consists of two GATConv layers, each using a multi-head attention mechanism. Specific parameters are shown in Table 6.
[0133] Table 6 Specific Parameters ;
[0134] Attention coefficient calculation between node i and its neighbor node j (taking the first layer as an example):
[0135] ;
[0136] in , (Dimensions after stitching: 128) The input features for node i are 768 dimensions. After normalization, the attention weights are obtained. Then update the node features: ;
[0137] σ is the ELU activation function.
[0138] After two layers of GAT, each node obtains a final 128-dimensional representation, which incorporates neighborhood information.
[0139] S5.2 Thread Subgraph Pooling employs a multi-instance learning strategy to decompose the heterogeneous graph into multiple thread-level subgraphs: each subgraph consists of the thread node tid_i and all its descendant API nodes. Mean pooling is performed on each subgraph, averaging the 128-dimensional features of all API nodes within that subgraph to obtain its representation vector. For sample 0012, two subgraph vectors are obtained. and ,gather .
[0140] S6: Input the set of subgraph representation vectors generated in step S5 into an aggregation classifier based on multi-instance learning to obtain the judgment result. The classifier first makes independent predictions (votes) for each subgraph vector, then obtains the final malicious probability of the control program or plugin through an aggregation function, and uses the cross-entropy loss function to perform end-to-end optimization training on all learnable parameters of the entire model.
[0141] S6.1 Subgraph Independent Prediction: Set the subgraph vectors... Input is a voting-based aggregation classifier. This classifier contains a weight-shared sub-classifier consisting of a two-layer fully connected network: 128-dimensional input → 64-dimensional hidden layer (ReLU activation) → output dimension (Sigmoid activation). Output Let be the probability of malicious activity for the i-th thread. For sample 0012, the calculated probability is... , .
[0142] S6.2 Aggregation function: Using the max pooling aggregation function, the sample-level malicious probability is obtained.
[0143] ;
[0144] because The sample is determined to be malicious. This strategy conforms to the optimal rule that "if any subgraph is malicious, then the sample is malicious".
[0145] S6.3 Model optimization: During the training phase, the aggregated probability of each sample in the training set is... With real labels (1 represents malicious, 0 represents benign) Substitute into the binary cross-entropy loss function:
[0146] ;
[0147] The gradient of the loss with respect to all parameters of BERT, GAT, and the classifier is calculated through backpropagation. The parameters are then updated using the Adam optimizer (learning rate 1e-4, training epochs 50). The training process is monitored as follows:
[0148] Training set loss: decreased from an initial 0.693 to 0.087 in round 50.
[0149] Validation set accuracy: 85.2% in round 10 → 94.1% in round 30 → 95.3% in round 50 (no overfitting).
[0150] Early stopping strategy: If the F1 score on the validation set does not improve for 5 consecutive rounds (threshold 0.001), the process stops. The optimal model is the 46th round.
[0151] The model was evaluated on the test set (774 samples, 384 benign and 390 malicious) and compared with existing models. The results are shown in Table 7.
[0152] Table 7 Comparison Results ;
[0153] The table above shows a comparison of the classification accuracy of a method implemented according to an embodiment of the present invention with several other models on the dataset. Bold text indicates the highest accuracy on that dataset. The data in the table shows that the method implemented according to an embodiment of the present invention exhibits high classification accuracy on the tested dataset. This result demonstrates that the technical solution proposed in this invention has significant technical effects in robot control sample data processing and malicious intent detection, providing a reliable technical solution for the safety control of humanoid robots.
[0154] Referring to Figure 5, the core functional module execution diagram provided in this embodiment of the invention includes the following module contents:
[0155] The dynamic behavior capture module is used to safely run robot control samples in an LXC-based isolated environment. This module uses tools such as strace and rosbag to comprehensively capture key dynamic behavior data during sample runtime, including API call sequences, process and thread IDs, and precise timestamps.
[0156] The functional semantic embedding module is responsible for converting the natural language descriptions of robot APIs into machine-understandable numerical features. This module first obtains the functional description texts of the APIs from authoritative sources such as the ROS / ROS2 official documentation and hardware SDK manuals, constructing a corpus. Then, using a pre-trained language model based on the Transformer architecture, each functional description is processed into a high-dimensional, dense functional semantic embedding vector, which encodes the core control intent of the API.
[0157] The Semantic Behavior Graph Construction Module is used to build a directed heterogeneous graph that accurately depicts the concurrency and sequence of each robot control program / plugin. This module constructs a three-layer topology with file ID as the root node, thread ID as the intermediate node, and API calls as leaf nodes. Its core task is to use the semantic vectors generated by the Functional Semantic Embedding Module as the initial features of the corresponding API nodes in the graph, thereby integrating the functional intent into the graph structure.
[0158] The graph attention feature learning module, as the core analysis engine of the system, is responsible for executing the process shown in Figure 4. It receives the semantic behavior graph as input and processes it through a deep encoder containing multiple layers of GATConv and a multi-head attention mechanism. This module can dynamically assign different importance levels to API neighbor nodes, learn deep compositional patterns and dependencies between API calls, and finally aggregate the information of the entire graph into a representation vector that accurately represents the global control behavior of the samples through graph pooling operations.
[0159] The classification prediction and model optimization module is responsible for making the final judgment based on the learned features. This module inputs the global representation vector output by the "Graph Attention Feature Learning Module" into a fully connected classifier to obtain the prediction probability that controls whether the program / plugin is benign or malicious. During training, it uses the cross-entropy loss function to evaluate the gap between the prediction and the true label, and performs end-to-end optimization of all parameters of the entire model through the backpropagation algorithm and optimizer until the model performance converges.
[0160] Those skilled in the art will understand that all or part of the steps described in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.
[0161] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0162] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A safety control method for humanoid robots based on API semantic awareness, characterized in that, Includes the following steps: S1: Construct a training and evaluation dataset containing benign and malicious samples, and deploy a dynamic analysis environment based on container technology; run the samples in the dataset in the dynamic analysis environment, capture dynamic behavior data through system call monitoring tools, and extract API call sequences sorted by time and their process or thread information; S2: For the set of target APIs in the API call sequence extracted by S1, obtain the functional description text corresponding to each API from at least one authoritative technical document source in the robotics field, and build an API functional description corpus; S3: Using a general pre-trained bidirectional language model based on the Transformer architecture, and using the API function description corpus constructed in step S2 as training data, the pre-trained bidirectional language model is adaptively fine-tuned in the robotics domain. Through fine-tuning, the model can output a high-dimensional dense vector representing the API control function intent; the high-dimensional dense vector serves as the functional semantic embedding of the corresponding robot API call. S4: For the robot control program or plugin to be analyzed, capture the API call sequence generated during its runtime, and construct a multi-layer semantic behavior heterogeneous graph that characterizes the concurrent behavior and temporal control behavior of the robot control program or plugin; the multi-layer semantic behavior heterogeneous graph is layered by thread, and the initial features of its API call nodes are initialized by the functional semantic embedding obtained in step S3. S5: Input the semantic behavior heterogeneous graph constructed in step S4 into the graph attention network for end-to-end feature learning; The correlation between nodes in the graph is calculated by multi-layer multi-head attention mechanism, and pooling operation is performed on the subgraphs corresponding to each thread in the graph to obtain a set of thread subgraph representation vectors with fixed dimensions. Step S5 specifically includes: S501: Input the semantic behavior heterogeneous graph constructed in step S4 into the graph attention network encoder. The graph attention network encoder includes at least two GATConv layers, and each GATConv layer adopts a multi-head attention mechanism. The graph attention network encoder learns the complex relationship between nodes by dynamically allocating attention weights between nodes, and at the same time combines the API call frequency as the edge weight to assist in the calculation, and updates the initial features of each node to higher-order features that fuse neighborhood information. The attention coefficient e between node i and its neighbor node j satisfies: ; in, and Let be the feature vectors of node i and node j, respectively, and W be the learnable weight matrix shared by all nodes. It is an attention function; The attention coefficients e are normalized using the Softmax function to obtain the final attention weights. : ; in, It is the set of all neighboring nodes of node i, || represents vector concatenation, a is the learnable weight vector in the attention mechanism, and LeakyReLU is the activation function; S502: Employs a multi-instance learning strategy to decompose the semantic behavior heterogeneous graph into multiple thread-level subgraphs. Each thread-level subgraph consists of thread TID intermediate nodes and their associated API call leaf nodes. Pooling is performed on each thread-level subgraph to aggregate the high-order features of all nodes within the subgraph into a fixed-dimensional subgraph representation vector. ,in The control behavior representation of the i-th thread is formed into a set of thread subgraph representation vectors; S6: Input the set of thread subgraph representation vectors generated in step S5 into the aggregation classifier based on multi-instance learning. The aggregation classifier performs independent malicious probability prediction on each thread subgraph representation vector. Then, the aggregation function fuses all independent prediction results to obtain the sample-level malicious probability of the robot control program or plugin. The cross-entropy loss function is used to optimize the learnable parameters of the pre-trained bidirectional language model, graph attention network and aggregation classifier end-to-end to complete model training and sample determination. Step S6 specifically includes: S601: Convert the subgraph representation vector set generated in step S502 into a single vector set. Input a voting-based aggregation classifier; use a weight-sharing sub-classifier within the aggregation classifier to independently predict the representation vector of each thread subgraph, thus obtaining the malicious prediction probability of each thread subgraph. , ∈[0,1]; This is considered malicious voting in the subgraph; S602: Employs a max-pooling aggregation function to calculate the malicious prediction probability for all thread subgraphs. By merging the data, the final probability of malicious activity of the robot control program or plugin can be obtained. , n is the number of thread subgraphs; S603: Predict the final probability of the aggregated samples Substituting the true labels of the samples into the binary cross-entropy loss function: ; In the real labels, y=1 indicates malicious intent, and y=0 indicates benign intent. The gradient of the loss function with respect to all learnable parameters of the model is calculated using the backpropagation algorithm. The parameters are then updated using the optimizer to minimize the loss function. This optimization process is repeated until the model performance converges.
2. The method for safety control of a humanoid robot based on API semantic awareness according to claim 1, characterized in that, Step S1 specifically includes: S101: First, build training and evaluation datasets. Benign samples come from the official control programs of robot manufacturers and certified plugins from the ROS / ROS2 open source community; malicious samples come from artificially constructed malicious control plugins and publicly available robot malicious sample sharing platforms. S102: Deploy an LXC container dynamic analysis platform dedicated to bot malicious sample analysis, the platform comprising: S1021: Container-isolated execution environment: Multiple isolated container instances are built based on LXC for independently running the robot control samples to be analyzed; S1022: Robot Simulation Interaction Module: Integrates the Gazebo robot simulation environment inside each LXC container, communicates with the sample running inside the container through the ROS / ROS2 interface, and simulates the control behavior of the sample on a real robot. S1023: System Behavior Monitoring Module: Configure the strace monitoring tool at the host level to trace the process of each LXC container and capture its system call sequence; at the same time, configure the rosbag tool to capture topic messages published by samples in the container through ROS / ROS2; S1024: Centralized Log Collection Module: Used to collect log data generated by strace and rosbag from each container after the sample runs out of time or ends, and archive it by sample ID; S103: Deploy benign and malicious samples one by one into independent LXC containers for execution. Use the strace tool to execute preset monitoring commands and combine it with the rosbag tool to monitor the runtime behavior of the samples, capture API call sequences, process IDs (PIDs), thread IDs (TIDs), and timestamps; set the maximum runtime for each sample, and forcibly terminate the container and collect runtime logs after the timeout. S104: Analyze the runtime logs generated by the strace and rosbag tools, extract the API call sequence of each sample arranged in chronological order, and associate it with the corresponding process and thread identifiers to form structured API call sequence data, which serves as the basis for subsequent semantic analysis and graph modeling.
3. The method for safety control of a humanoid robot based on API semantic awareness according to claim 1, characterized in that, Step S2 specifically includes: S201: For the API call sequence compiled in S104, retrieve the functional description text of each API from authoritative documents in the robotics field, including ROS / ROS2 official documents, robot hardware SDK manuals, and manufacturer API technical specifications. S202: Use automated scripts to crawl the functional description text of each API in authoritative technical documents. For APIs whose functional descriptions are not found in the documents, manually search and supplement their functional description texts. Store all API functional description texts in a centralized manner to form an API functional description corpus.
4. The method for safety control of a humanoid robot based on API semantic awareness according to claim 1, characterized in that, Step S3 specifically includes: S301: The BERT model is used as the semantic extraction model. The BERT model learns contextual semantics through the masked language model pre-training task. The BERT model is initialized with general pre-trained weights. Then, the API function description corpus constructed in step S2 is used as training data to perform robot domain adaptive fine-tuning on the BERT model so that the model can adapt to the semantic representation requirements of robot API functions. S302: Input the functional description text of each API into the fine-tuned BERT model for encoding, and generate a high-dimensional dense functional semantic embedding vector for the corresponding API call. The vector is used to characterize the core functional semantics of the API call, so as to achieve robust identification of API renaming obfuscation methods.
5. A humanoid robot safety control method based on API semantic awareness according to claim 1, characterized in that, Step S4 specifically includes: S401: Construct a directed heterogeneous graph as a multi-layer semantic behavior heterogeneous graph. The directed heterogeneous graph includes three types of nodes: the root node representing the file ID of the robot control program or plug-in, the intermediate node representing the thread TID of the concurrent execution flow, and the leaf node representing the API call of the actual control operation. S402: Construct a graph topology based on three types of nodes. The construction logic includes: establishing directed edges from the root node of the file ID to all intermediate nodes of the thread TID derived from it to realize concurrency relationship modeling; establishing directed edges from the intermediate node of each thread TID to all API call leaf nodes called within its corresponding thread to realize dependency relationship modeling; and establishing directed edges between adjacent API call leaf nodes in the same thread according to the API call time order to realize temporal relationship modeling. S403: Use the functional semantic embedding vector generated in step S3 as the initial feature vector of the corresponding API call leaf node to complete the construction of the semantic behavior heterogeneous graph.
6. A humanoid robot safety control system based on API semantic awareness, characterized in that, The system is used to implement the API semantic awareness-based humanoid robot safety control method according to any one of claims 1-5, the system comprising: The dynamic behavior capture module is used to build training and evaluation datasets, deploy a dynamic analysis environment based on LXC containers, monitor the runtime behavior of robot control samples through strace, rosbag and Gazebo robot simulation environment, capture API call sequences, process identifiers, thread identifiers and timestamps, and output structured API call sequence data. The functional semantic embedding module is used to obtain API functional description text from authoritative technical documents in the robotics field, build an API functional description corpus, and encode the API functional description text into high-dimensional dense functional semantic embedding vectors through a BERT model that is adaptively fine-tuned in the robotics field. The semantic behavior graph construction module is used to construct a directed heterogeneous graph containing file ID root nodes, thread TID intermediate nodes, and API call leaf nodes based on structured API call sequence data. The functional semantic embedding vector is used as the initial feature of the API call leaf nodes to complete the construction of a multi-layer semantic behavior heterogeneous graph. The graph attention feature learning module receives the semantic behavior heterogeneous graph output by the semantic behavior graph construction module. Through a graph attention network containing multiple GATConv layers and a multi-head attention mechanism, it learns high-order relationships between nodes, performs pooling operations on thread-level subgraphs, and outputs a set of thread subgraph representation vectors. The classification prediction and model optimization module is used to input the set of thread subgraph representation vectors output by the graph attention feature learning module into the aggregation classifier to complete the prediction of the malicious probability of the sample. Through the cross-entropy loss function and backpropagation algorithm, it performs end-to-end optimization on all learnable parameters in the system to realize the benign and malicious determination of the robot control program or plug-in.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the API semantic awareness-based humanoid robot safety control method according to any one of claims 1-5.
Citation Information
Patent Citations
Malicious software packer identification method based on heterogeneous graph neural network
CN121211448A
Platform-enabled orchestration and optimization of digital workflows
WO2025064639A1