Data security protection system and method based on anti-AI penetration technology

By introducing identity pre-verification, AI behavior whitelist rules and real-time monitoring mechanisms when registering AI entities in the data security protection system, combined with deep learning algorithms to judge abnormal API call behavior, the problem of difficulty in distinguishing between legitimate AI and malicious AI in the existing technology is solved, and more efficient AI access control and data security protection are achieved.

CN120151098AActive Publication Date: 2025-06-13ZHEJIANG HANNAO DIGITAL TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510512368.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-06-13
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing technology is difficult to distinguish between legal AI and malicious AI, resulting in insufficient targeted protection measures.

Method used

By introducing the identity pre-verification mechanism, AI behavior whitelist rules and real-time monitoring mechanism when registering AI entities, combined with deep learning algorithms to judge abnormal API call behavior, generate AI identity identification and limit access scope.

Benefits of technology

Effectively identifying and blocking the fake identity and abnormal API call behavior of AI entities, improving the system's adaptability to diversified access control needs, and reducing the manual maintenance cost of permission management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151098A_ABST
    Figure CN120151098A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data security and network security, and discloses a data security protection system and method based on an anti-AI penetration technology, and the method comprises the following steps: S1, receiving user identity data sent by a client; s2, generating a unique ePASS-ID through a national secret algorithm or other encryption algorithms; s3, according to a preset organizational structure information database, analyzing the digital service file, and generating authority data corresponding to the ePASS-ID; s4, receiving a registration request of an AI entity, obtaining model fingerprint data of the AI entity, and generating an AI identity label through a national cryptographic algorithm or other encryption algorithms; s5, judging whether abnormity exists or not based on a deep learning algorithm; and S6, processing the service data. According to the method and the device, an identity pre-verification mechanism, an AI behavior white list rule and a real-time monitoring mechanism during AI entity registration are introduced, so that the limitation that legal AIs and malicious AIs are difficult to distinguish in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of data security and network security, and particularly to a data security protection system and method based on anti-AI penetration technology. Background Art

[0002] In the prior art, data security protection systems usually protect data resources through access control lists (ACLs), firewall rules, and static authentication mechanisms. For example, role-based access control (RBAC) technology restricts data access by predefined user permissions, and combines symmetric encryption algorithms (such as AES) or asymmetric encryption algorithms (such as RSA) to protect data during transmission and storage. In addition, some systems use logging and rule matching methods to monitor abnormal behaviors. For example, potential malicious access is detected by setting API call frequency thresholds. These technologies have been widely applied in traditional network security scenarios and have achieved protection against known threats to a certain extent.

[0003] In the prior art, attack methods for AI entities to forge identities include, but are not limited to, simulating user behaviors through generative adversarial networks (GANs), forging biometric data, or probing system vulnerabilities through high-frequency API calls. These attacks usually occur in scenarios of unregistered or anonymous access, and existing systems lack the ability to pre-verify the identities of AI entities and extract behavioral characteristics. For example, traditional systems are difficult to distinguish between legitimate AIs (such as intelligent customer services deployed by enterprises) and malicious AIs (such as automated scripts of external attackers), resulting in insufficient pertinence of protection measures. Summary of the Invention

[0004] To make up for the above deficiencies, the present invention provides a data security protection system and method based on anti-AI penetration technology, aiming to improve the problem that traditional systems are difficult to distinguish between legitimate AIs (such as intelligent customer services deployed by enterprises) and malicious AIs (such as automated scripts of external attackers).

[0005] In a first aspect, the present invention provides the following technical solution. A data security protection method based on anti-AI penetration technology, which is applied to the server side, includes the following steps: S1. Receive user identity data sent by a client, where the user identity data includes biometric data and behavioral data; S2. Based on the user identity data, generate a unique ePASS-ID through a national cryptographic algorithm or other encryption algorithms, and encrypt and store the biometric data by using a national cryptographic algorithm or other encryption algorithms; S3. According to a preset organizational structure information database, parse a digital appointment document, generate permission data corresponding to the ePASS-ID, and the permission data is dynamically allocated based on a four-dimensional permission control model; S4. Receive the registration request of the AI entity, pre-verify the identity of the AI entity, obtain its model fingerprint data after passing the verification, generate an AI identity identifier through a national cryptography algorithm or other encryption algorithms, and restrict the access scope of the AI entity according to the AI behavior whitelist rules; S5. Real-time monitor the API call behavior of the AI entity, judge whether there is an abnormality based on the deep learning algorithm, and block the access of the AI entity and trigger an alarm if there is an abnormality; S6. Process business data, encrypt and transmit it through a national cryptography algorithm or other encryption algorithms, encrypt and store it through a national cryptography algorithm or other encryption algorithms, and verify the data integrity through a national cryptography algorithm or other encryption algorithms.

[0006] Preferably, the biometric data in step S1 includes face, iris, voiceprint and finger vein data, the behavior data includes user operation habit data, and the user identity data is verified through a multi-modal credential system.

[0007] Preferably, the four-dimensional permission control model in step S3 includes: Organizational structure dimension, generating a permission inheritance relationship according to the enterprise hierarchical structure; Data sensitivity dimension, classifying and encrypting data based on a national cryptography algorithm or other encryption algorithms; Business scenario dimension, encrypting and sandboxing different business environments using a national cryptography algorithm or other encryption algorithms; Time-space dimension, restricting access permissions according to GPS and IP addresses.

[0008] Preferably, the AI behavior whitelist rules in step S4 include: Register the model fingerprint data of the AI entity and limit its data interaction scope; Prohibit the AI entity from accessing sensitive data beyond the permissions corresponding to its ePASS-ID.

[0009] Preferably, the real-time monitoring in step S5 includes: Isolate the operations of the AI entity in the AI behavior sandbox; If an abnormal API call is detected, trigger secondary biometric verification.

[0010] Preferably, step S6 further includes: Add a national cryptography algorithm or other encryption algorithm hash digital watermark to the exported data; Set up a multi-level approval process for sensitive data operations and chain the evidence.

[0011] In a second aspect, the present invention provides the following technical solution, a data security protection system based on anti-AI penetration technology, configured on the server side, including: A receiving module, configured to receive user identity data and a registration request of an AI entity sent by a client; An identity generation module, configured to generate an ePASS-ID based on the user identity data and generate an identity identifier for the AI entity; A permission allocation module, configured to generate and allocate permission data according to an organizational structure information database; A monitoring module, configured to monitor the API call behavior of the AI entity in real time and determine anomalies; A data processing module, configured to perform encrypted transmission and storage on service data.

[0012] Preferably, the identity generation module generates an ePASS-ID through a national cryptographic algorithm or other encryption algorithms, and encrypts biometric data by using a national cryptographic algorithm or other encryption algorithms.

[0013] In a third aspect, the present invention provides the following technical solution: a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the data security protection method based on anti-AI penetration technology described above is implemented.

[0014] In a fourth aspect, the present invention provides the following technical solution: a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the data security protection method based on anti-AI penetration technology described above is implemented.

[0015] The present invention has the following beneficial effects: 1. In the present invention, by introducing an identity pre-verification mechanism during AI entity registration, an AI behavior whitelist rule, and a real-time monitoring mechanism, it is possible to effectively identify and block the forged identities and abnormal API call behaviors of AI entities (such as malicious generation models), assign a traceable identity identifier to legitimate AIs through digital signature and behavior log verification, and generate model fingerprint data by combining behavior feature extraction, solving the limitation in the prior art that it is difficult to distinguish between legitimate and malicious AIs.

[0016] 2. In the present invention, the proposed four-dimensional permission control model (organizational structure, business scenario, data sensitivity, time-space) realizes dynamic allocation and real-time update of permissions through data logic and database parsing technology, overcomes the limitation of single-dimensional static configuration in traditional permission management, can quickly adapt to permission requirements in complex organizational structures and multi-business scenarios, improves the adaptability of the system to diverse access control requirements, and reduces the manual maintenance cost of permission management at the same time.

[0017] 3. In the present invention, a multi-modal credential system is adopted, which combines biometric data (face, iris, voiceprint, finger vein) with behavioral data. An ePASS-ID is generated through national cryptographic algorithms or other encryption algorithms, and national cryptographic algorithms or other encryption algorithms are used to respectively implement asymmetric encryption transmission, symmetric encryption storage, and integrity verification. Compared with traditional single encryption methods, this technical solution significantly improves the anti-theft and anti-tampering capabilities of business data through multi-level encryption and verification mechanisms, and can ensure the security of the entire data life cycle especially in cross-network transmission and distributed storage environments.

[0018] 4. In the present invention, by generating a unique identity identifier for the AI entity and combining behavioral sandbox isolation technology, real-time monitoring of AI operations and rapid positioning of abnormal behaviors are achieved. Compared with traditional access control methods with undifferentiated management, this technical solution can record the model fingerprints and API call logs of AI entities, facilitating post-event auditing and responsibility tracing, thereby improving the management efficiency and security of the system in scenarios where AI is widely used. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flowchart of the data security protection method based on anti-AI penetration technology proposed by the present invention; Figure 2 is a system architecture diagram of the data security protection system based on anti-AI penetration technology proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] Embodiment 1 Refer to Figure 1 , in the first embodiment of the present invention, the present invention provides a data security protection method based on anti-AI penetration technology, which is applied to the server side and includes the following steps: S1. Receive user identity data sent by the client, where the user identity data includes biometric data and behavioral data; S2. Based on the user identity data, generate a unique ePASS-ID through national cryptographic algorithms or other encryption algorithms, and encrypt and store the biometric data using national cryptographic algorithms or other encryption algorithms; S3. According to the preset organizational structure information database, parse the digital appointment document to generate permission data corresponding to the ePASS-ID, and the permission data is dynamically allocated based on a four-dimensional permission control model; S4. Receive the registration request of the AI entity, pre-verify the identity of the AI entity, obtain its model fingerprint data after passing the verification, generate an AI identity identifier through the national cryptography algorithm or other encryption algorithms, and restrict the access scope of the AI entity according to the AI behavior whitelist rules; S5. Real-time monitor the API call behavior of the AI entity, judge whether there is an abnormality based on the deep learning algorithm, and block the access of the AI entity and trigger an alarm if there is an abnormality; S6. Process the business data, encrypt and transmit it using the national cryptography algorithm or other encryption algorithms, encrypt and store it using the national cryptography algorithm or other encryption algorithms, and verify the data integrity through the national cryptography algorithm or other encryption algorithms.

[0022] Specifically, in practical applications, the server side can be a high-performance computing device deployed on an enterprise private cloud or public cloud, such as a server cluster that supports multi-core CPUs and GPU acceleration. Its operating system can be Linux or Windows Server to meet the high-concurrency data processing requirements. In step S1, the user identity data sent by the client is transmitted in JSON format via the HTTPS protocol. The biometric data (such as iris images) is uploaded after preprocessing (such as grayscale conversion and feature point extraction), and the behavioral data records dynamic information such as the user's keyboard input frequency and mouse movement trajectory through the client log collection module. In step S2, the national cryptographic algorithm is used to generate the ePASS-ID. A hash algorithm that complies with national cryptographic standards is adopted to generate a unique identifier of a fixed length to ensure collision resistance. The national cryptographic algorithm also includes identity-based encryption technology. The key generation center (KGC) assigns a key to each user, and the encrypted biometric data is stored in a distributed database (such as MongoDB) to improve query efficiency. The organizational structure information database in step S3 can be an enterprise employee information table built based on a relational database (such as MySQL). The digital appointment document is stored in XML format and contains fields such as position, department, and permission level. The four-dimensional permission control model implements the dynamic allocation logic through a scripting language (such as Python) and supports real-time update and revocation of permissions. In steps S4, the AI entity registration request is submitted through a dedicated API interface. The request data includes the meta-information of the AI entity (such as model type, deployment environment) and a digital signature. The RSA algorithm (2048-bit key) is used to generate the signature. The server side verifies the legality of the digital signature through a preset trust root certificate and only allows registration requests from trusted sources (such as in-house enterprise AI services) to pass. If the verification fails, the system records the request IP and adds it to a temporary blacklist to prevent subsequent attempts. After successful verification, the server side requires the AI entity to execute a standardized test task in a controlled environment. The task is a dynamic challenge sequence, including 100 predefined API call requests (such as GET / api / data, POST / api / update, with the input in the fixed JSON format {"query":"test"}) and 15 adversarial challenge requests. The adversarial challenge requests are generated based on the historical behavior templates of legitimate in-house enterprise AIs (extracted from the Elasticsearch database, covering 100,000 calls of 1,000 AI entities). The templates include statistical features of the API call sequences (such as mean response time, error rate distribution).Challenge requests include non-standard parameters (such as malformed JSON {"query": null}), high-load computing tasks (such as matrix operation APIs), and forgery behavior simulation (such as extremely high-frequency calls, 10 times per second). Legitimate AIs can handle them efficiently due to optimizations for enterprise APIs (such as fine-tuning or caching mechanisms), with an error rate lower than 5% and stable response times (mean < 200 milliseconds); forged AIs exhibit high error rates (> 20%) or inconsistent resource usage (memory fluctuation > 30%) due to lack of optimization. The test tasks run in isolated Docker containers (4-core CPU, 8GB memory, Ubuntu 20.04), and the behavior logs are recorded for each call. The logs include response time (in milliseconds), memory usage (in MB), CPU utilization rate (%), and API call error rate (proportion of error responses). The behavior analysis algorithm uses the k-means clustering algorithm (k-means clustering) to extract model fingerprint data. k-means is selected for its efficiency and robustness to high-dimensional data, and it can extract a compact representation of AI computing features from the behavior logs. The specific steps are as follows: 1) Normalize the response time, memory usage, CPU utilization rate, and error rate and map them to [0, 1]; 2) Each log record is represented as a four-dimensional feature vector, forming a feature matrix of 115 records (100 API calls + 15 challenge requests); 3) Apply the k-means algorithm (k = 5, based on the elbow method, k-means++ initialization, maximum 100 iterations, convergence error threshold 0.001) to generate 5 cluster centers; 4) Concatenate the cluster centers into a 20-dimensional vector (5 × 4 dimensions) as fingerprint data. The fingerprint data reflects the computing features of the AI (such as response time stability, error rate tendency). The fingerprint vectors of legitimate AIs form a compact cluster (intra-class distance < 0.1) due to low error rates and high stability, while forged AIs are outliers (inter-class distance > 0.3) due to high error rates and resource anomalies. To verify the fingerprint discrimination ability, the system calculates the Euclidean distance between the fingerprint vector and the legitimate AI template. The mean distance of the forged AIs exceeds 2 standard deviations. The test tasks are repeated 3 times, and the fingerprint mean is taken to ensure consistency; if the number of logs is less than 90, the system rejects registration and records the anomaly. The fingerprint data generates a 32-byte AI identity identifier through the national cryptographic algorithm (SHA-256) and is stored in MongoDB. The whitelist rules are based on the function scope (such as "read operations only") and fingerprint data, in JSON format, and the Redis cache supports 100,000 queries per second. If the registration fails, the system sends an alert via SMTP. Step S5: Continuously monitor the API call behavior of the AI entity, and based on the deep learning algorithm, determine whether there is an anomaly. If there is an anomaly, block the access of the AI entity and trigger an alert.Specifically, the real-time monitored AI behavior sandbox is built based on Linux's cgroups and namespace technologies. Each AI entity is allocated independent CPU and memory resources (2 cores and 4GB) to prevent resource competition. A lightweight proxy program (written in Go language) runs inside the sandbox, collecting API call logs in real time, including call time, interface name, parameters, and error rate (from the challenge task definition in step S4). The logs are transmitted to the server side through Kafka and stored in Elasticsearch. For anomaly detection, a Long Short-Term Memory (LSTM) model is used, trained with the PyTorch framework, and the training objective is binary classification (normal vs. abnormal). The training samples include: 1) Normal samples: 100,000 API calls of legitimate AIs within the enterprise (30 days, 1000 AI entities), with features including call time, interface name, parameter hash value, error rate (introduced from S4, <5%), labeled as "normal"; 2) Abnormal samples: 5500 records generated from red team tests, covering high-frequency probes (1000), privilege escalation (1000), GAN forgery (1000), low-frequency forgery (500), and mixed attacks (1000, such as forgery + privilege escalation), with error rate >20%, labeled as "abnormal". The samples are stored in Elasticsearch. To adapt to LSTM, the data preprocessing is as follows: The call time is normalized to a relative time interval of [0,1]; the interface name is converted into a 64-dimensional vector through Word2Vec (based on 1000 API interface names, window size 5, 64-dimensional vector, 50 epochs, and the corpus is the enterprise API documentation); the parameter hash value (256-bit SHA-256) is reduced to 32 dimensions through PCA (retaining 90% of the variance); the error rate is directly used as a 1-dimensional feature; they are concatenated to form a 97-dimensional feature vector (64 + 32 + 1), and combined with a 50-call window, mapped to a 128-dimensional input vector through a fully connected layer. The training uses cross-entropy loss, L=-[y·log(ŷ)+(1-y)·log(1-ŷ)], where y is the label (0 for normal, 1 for abnormal). The model is a three-layer LSTM (256 hidden units in each layer, determined by grid search, balancing accuracy and efficiency), training set:validation set = 8:2, Adam optimizer (learning rate 0.001), 100 epochs, accuracy 95%, F1 score 0.93. The 50-call window is optimized through experiments, taking into account both temporal information and computational efficiency. During real-time detection, the system extracts a 50-call sequence, converts it into a 128-dimensional vector according to the preprocessing, and the LSTM outputs the anomaly probability (threshold 0.7). If it is abnormal, it is blocked through iptables, and a warning is pushed through WebSocket (including the anomaly type and timestamp). If the number of logs is less than 50, zero vectors are filled or the detection is paused. When an anomaly occurs, secondary biometric verification (iris or voiceprint, 30-second timeout, permanent ban for failure) is triggered.In step S6, the national cryptographic algorithm and other encryption algorithms are used for the asymmetric encryption of data transmission. The key length complies with national standards to ensure transmission security. The national cryptographic algorithm is also used for symmetric encryption. After the data is encrypted in blocks, it is stored in the local encrypted file system (such as LUKS). The national cryptographic algorithm verifies the data integrity by calculating the check value during data writing and reading to verify the consistency and prevent tampering. In step S5, the implementation of anomaly judgment based on the deep learning algorithm relies on a pre-trained long short-term memory (LSTM) model. The training samples of the model come from the following two types of data:. Normal samples: Generated from the historical API call logs of legitimate AI entities within the enterprise (such as intelligent customer service, data analysis models). The logs include features such as call time, interface name, parameter hash value, etc. Each record is labeled as "normal". The sample collection period is 30 days, covering at least 100,000 calls of 1,000 AI entities. The data is stored in the Elasticsearch database.

[0023] Abnormal samples: Generated by simulating attack scenarios, including API calls with forged identities (such as high-frequency probing, unauthorized access), forged behavior data (such as simulating user mouse trajectories), and known malicious AI behavior patterns (such as call sequences based on open-source malicious scripts). Abnormal samples are generated through red team testing, and the simulated attacks cover scenarios such as GAN forgery, SQL injection, etc. Each type of abnormal sample contains at least 5,000 records and is labeled as "abnormal".

[0024] The training process uses the PyTorch framework. The model input is the vector representation of the API call sequence (with a dimension of 128). The time series features are extracted through a three-layer LSTM network, and the output is a binary classification result (normal / abnormal). The training set and validation set of the model are divided at a ratio of 8:2. After training for 100 epochs, the accuracy reaches over 95%. During anomaly judgment, the system extracts the API call sequence of the AI entity in real time and compares it with the training model. If the classification result is abnormal, the blocking and warning mechanisms are triggered.

[0025] The biometric data in step S1 includes face, iris, voiceprint, and finger vein data, and the behavior data includes user operation habit data. The user identity data is verified through a multi-modal credential system.

[0026] Specifically, the biometric data collection devices can include iris scanners (such as IrisGuard IG-AD100), voiceprint input microphones (such as Shure SM58), and finger vein recognition devices (such as Hitachi H1). These devices are connected to the client via USB or Bluetooth, and the data formats collected are JPEG (iris), WAV (voiceprint), and binary stream (finger vein) respectively. The iris boundary is extracted from the iris data through the Hough transform, the voiceprint features are extracted from the voiceprint data through Mel Frequency Cepstral Coefficients (MFCC), and the finger vein data generates a vascular texture map through infrared imaging technology. The behavior data collection module runs on the client operating system (such as Windows 10 or Android), captures the keyboard keystroke intervals and mouse click coordinates through hook functions (such as SetWindowsHookEx of Windows API), generates time series data, and the sampling frequency can be set to 10 times per second. The multi-modal credential system combines the dual verification of biometric and behavior data. First, the biometric features are preliminarily screened through the Support Vector Machine (SVM) algorithm. After successful matching, the Hidden Markov Model (HMM) is used to analyze the consistency of the behavior data. If both pass, the identity is confirmed as legal. The advantage of multi-modal verification is to improve the robustness of identity authentication. For example, when iris recognition fails due to insufficient light, the voiceprint and behavior data can still be used as supplementary evidence.

[0027] The four-dimensional permission control model in step S3 includes: The organizational structure dimension generates a permission inheritance relationship according to the enterprise hierarchical structure; The data sensitivity dimension classifies and encrypts data based on national secret algorithms or other encryption algorithms; The business scenario dimension encrypts and isolates different business environments using national secret algorithms or other encryption algorithms in a sandbox; The time-space dimension restricts access permissions according to GPS and IP addresses.

[0028] Specifically, the implementation of the four-dimensional permission control model relies on a multi-level permission management system deployed on the server side, developed using the Java Spring framework, and supporting RESTful interface calls. The organizational structure dimension stores enterprise-level information through a tree structure. For example, group B is subordinate to department A. The permission inheritance relationship is automatically calculated through SQL query statements (such as WITH RECURSIVE) to ensure that child nodes inherit the permissions of parent nodes. The business scenario dimension creates an independent national cryptography algorithm encryption sandbox through Docker container technology. Each sandbox is assigned a unique key and IP address, and the sandboxes are isolated through a virtual network. For example, the NetworkPolicy of Kubernetes is used to restrict cross-scenario access. The data sensitivity dimension classifies data into three levels: public, internal, and confidential. The hierarchical keys generated by the national cryptography algorithm are stored in a hardware security module (HSM, such as YubiHSM), and the encryption process supports batch processing, capable of processing 1000 records per second. The time-space dimension combines the client's GPS coordinates (accuracy ±10 meters) and IP address (resolved through the GeoIP database). The permission rules are stored in JSON format, and access requests outside the scope will be rejected.

[0029] The AI behavior whitelist rules in step S4 include: Register the model fingerprint data of the AI entity and limit its data interaction scope; Prohibit the AI entity from accessing sensitive data that exceeds the permissions corresponding to its ePASS-ID.

[0030] Specifically, the generation process of the AI behavior whitelist rules includes two stages: AI entity registration and rule matching. During registration, the AI entity needs to submit model fingerprint data. For example, a deep learning model trained based on PyTorch can export parameters through torch.save, and the server side uses the national cryptography algorithm to generate a fingerprint hash value as the AI identity identifier. The data interaction scope is limited through a configuration file and stored in the AI permission collection in MongoDB. The matching of the whitelist rules is implemented using the Trie tree algorithm, which supports fast prefix queries. For example, the rule for prohibiting access to sensitive data tables (such as the user information table) can be written as "deny:select:user_*", capable of processing 100,000 matching requests per second. The implementation of prohibiting access to sensitive data that exceeds the permissions depends on the access control list (ACL) of the database. Double verification is performed in combination with the AI identity identifier. If an unauthorized behavior is detected, the system will record a log and send an alert email to the administrator through the SMTP protocol.

[0031] The real-time monitoring in step S5 includes: Isolate the operations of the AI entity in the AI behavior sandbox; If an abnormal API call is detected, trigger secondary biometric verification.

[0032] Specifically, the real-time monitoring AI behavior sandbox is built based on Linux's cgroups and namespace technologies. Each AI entity is allocated independent CPU and memory resources (such as 2 cores and 4GB) to prevent resource competition. A lightweight proxy program (written in Go language) runs inside the sandbox to collect API call logs in real time, including call time, interface name, and parameters. The logs are transmitted to the server-side analysis through the Kafka message queue. The deep learning model for anomaly detection is based on the PyTorch framework. The training dataset includes normal API call samples (such as GET / api / data) and abnormal samples (such as high-frequency POST / api / admin), and the model accuracy can reach over 95%. If an anomaly is detected, for example, the API call frequency exceeds 100 times per second, the system blocks the network connection of the AI entity through iptables rules and simultaneously triggers secondary biometric verification. The secondary verification pushes a notification to the client (based on the WebSocket protocol), asking the user to resubmit iris or voiceprint data. The verification timeout is 30 seconds, and if it fails, the AI entity's access will be permanently blocked.

[0033] Step S6 further includes: Adding hash digital watermarks of national cryptography algorithms or other encryption algorithms to the exported data; Setting up a multi-level approval process for sensitive data operations and storing the evidence on the blockchain.

[0034] Specifically, the digital watermark generated by the national cryptography algorithm is achieved by calculating the hash value of the exported data (such as a PDF document), and then embedding the hash value into the file metadata or invisible area, such as the XMP field of PDF. The embedding process is completed through the open-source library PyMuPDF. When verifying the watermark, the system recalculates the file hash and compares it with the embedded value to ensure that the data has not been tampered with. The multi-level approval process is implemented based on a workflow engine (such as Activiti). For example, the export of sensitive data requires two-level approvals from the department manager and the security officer. The approval records are stored on the blockchain after generating a digest through the hash algorithm. The blockchain process uses the Hyperledger Fabric blockchain platform, and the data is stored in JSON format. Each record occupies about 1KB of storage space. The advantage of storing on the blockchain is to provide an immutable audit log, which is suitable for scenarios with high compliance requirements, such as data management in government agencies. In addition, the multi-level approval supports dynamic adjustment, such as automatically increasing the approval levels according to the data sensitivity to ensure the flexibility and security of the process.

[0035] To achieve the protection against the forged identity of AI entities, the following closed-loop mechanism works collaboratively in this embodiment: Identity pre-verification: In step S4, the registration of AI entities needs to be verified through digital signatures and behavior logs to ensure that only legitimate AIs can enter the system, excluding the initial risk of forged identities.

[0036] Behavior feature monitoring: In step S5, a deep learning model is used to analyze the API call sequence, and combined with the model fingerprint data, forgeries (such as simulating the call patterns of legitimate AIs) are dynamically detected.

[0037] Quick response mechanism: After anomaly detection, the system blocks access through iptables and triggers secondary biometric verification to prevent forged AIs from bypassing the protection through temporary identities.

[0038] For example, in the scenario deployed within an enterprise, a legitimate AI (such as an automatic report generation model) generates a unique identifier through registration, and its API call pattern conforms to the whitelist rules; when an external malicious AI attempts to forge its identity, its call sequence is recognized as abnormal by the deep learning model due to the lack of a legitimate behavior pattern, the system completes the block within 500 milliseconds, and the alarm information is sent to the administrator via SMTP.

[0039] Embodiment 2: Refer to Figure 2 , in the second embodiment of the present invention, the present invention provides a data security protection system based on anti-AI penetration technology, configured on the server side, including: A receiving module, used to receive the user identity data and the registration request of the AI entity sent by the client; An identity generation module, used to generate an ePASS-ID based on the user identity data and generate an identity identifier for the AI entity; A permission allocation module, used to generate and allocate permission data according to the organizational structure information database; A monitoring module, used to monitor the API call behavior of the AI entity in real time and judge anomalies; A data processing module, used to encrypt the transmission and storage of business data.

[0040] Specifically, the physical hardware on the server side where the data security protection system is deployed can be a rack-mounted server supporting RAID5 disk arrays, equipped with at least 64GB of memory and 1TB of SSD storage to meet the high-concurrency access requirements. The receiving module realizes load balancing through Nginx reverse proxy, supports processing 5,000 client requests per second, and the user identity data and AI registration requests are transmitted in Protobuf format to reduce bandwidth occupancy. The identity generation module runs on an independent microservice instance (based on Spring Boot), accelerates the generation of ePASS-ID through Redis cache, and can generate 10,000 identities per second; the generation of AI identity identifiers is linked with the AI model fingerprint database (PostgreSQL) to ensure consistency. The permission allocation module synchronizes the organizational structure information through a scheduled task (CronJob), updates the permission data daily, and supports a user scale of 100,000. The monitoring module integrates Prometheus and Grafana tools to display the API call frequency and exception rate in real time, and the exception judgment logic is deployed as an independent Python script, supporting dynamic loading of new rules. The data processing module implements data encryption based on national secret algorithms and other encryption algorithms, with an encryption speed of up to 50MB per second, and the stored procedure realizes high availability through a distributed file system (such as HDFS).

[0041] The identity generation module generates ePASS-ID through national secret algorithms or other encryption algorithms, and encrypts biometric data using national secret algorithms or other encryption algorithms.

[0042] Specifically, the specific implementation of the identity generation module includes an encryption service sub-module that runs inside a Docker container on the server side. The container image is built based on Ubuntu20.04 and pre-installed with a national secret algorithm library. The input of the national secret algorithm is the binary stream of user identity data, such as an iris feature vector (about 2KB), and the output is a hash value of a fixed length. The generation process takes about 5 milliseconds. The national secret algorithm also includes identity-based encryption technology. The key pair is generated by the KGC, the key length complies with national standards, and the private key is stored in the HSM to prevent leakage. The encrypted biometric data is saved to the database in Base64 encoded format, supporting fast retrieval. For example, the corresponding record can be queried within 1 millisecond through the index field "user_id". The module supports multi-threaded processing and can concurrently process 1,000 encryption requests per second, suitable for high-load scenarios. In addition, the identity generation module provides an error handling mechanism. For example, if the input data format is incorrect, a JSON-formatted error code is returned to ensure the robustness of the system.

[0043] Embodiment Three The third embodiment of the present invention, based on the same inventive concept, provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the data security protection method based on the anti-AI penetration technology in the above embodiment.

[0044] Embodiment 4 The fourth embodiment of the present invention, based on the same inventive concept, provides a computer device, which includes: a processor and a memory; the processor and the memory communicate with each other; the memory is used to store instructions; the processor is used to execute the instructions in the memory to implement the data security protection method based on the anti-AI penetration technology in the above embodiment.

[0045] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0046] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data security protection method based on anti-AI penetration technology, characterized in that: Applied to the server side, it includes the following steps: S1. Receive user identity data sent by a client, where the user identity data includes biometric data and behavior data; S2. Based on the user identity data, generate a unique ePASS-ID using a national secret algorithm or other encryption algorithm, and encrypt and store the biometric data using the national secret algorithm or other encryption algorithm; S3. According to a preset organizational structure information database, the digital appointment document is parsed to generate permission data corresponding to the ePASS-ID, and the permission data is dynamically allocated based on a four-dimensional permission control model; S4. Receive the registration request of the AI ​​entity, pre-verify the identity of the AI ​​entity, obtain its model fingerprint data after the verification, generate the AI ​​identity through the national secret algorithm or other encryption algorithm, and limit the access scope of the AI ​​entity according to the AI ​​behavior whitelist rules; S5. Monitor the API call behavior of the AI ​​entity in real time, determine whether there is an abnormality based on the deep learning algorithm, and if there is an abnormality, block the access of the AI ​​entity and trigger an alarm; S6. Process the business data, encrypt and transmit it using the national secret algorithm or other encryption algorithms, encrypt and store it using the national secret algorithm or other encryption algorithms, and verify the data integrity using the national secret algorithm or other encryption algorithms.

2. The data security protection method based on anti-AI penetration technology according to claim 1 is characterized in that: The biometric data in step S1 includes face, iris, voiceprint and finger vein data, the behavioral data includes user operation habit data, and the user identity data is verified through a multimodal credential system.

3. The data security protection method based on anti-AI penetration technology according to claim 1 is characterized in that: The four-dimensional authority control model in step S3 includes: Organizational structure dimension, generating authority inheritance relationship based on the enterprise hierarchy; Data sensitivity dimension: encrypt data in different levels based on national encryption algorithms or other encryption algorithms; In the business scenario dimension, use the national encryption algorithm or other encryption algorithms to encrypt the sandbox to isolate different business environments; Time-space dimension, restricting access rights based on GPS and IP addresses.

4. The data security protection method based on anti-AI penetration technology according to claim 1 is characterized in that: The AI ​​behavior whitelist rules in step S4 include: Register the model fingerprint data of the AI ​​entity and limit its data interaction scope; The AI ​​entity is prohibited from accessing sensitive data beyond the permissions corresponding to its ePASS-ID.

5. The data security protection method based on anti-AI penetration technology according to claim 1 is characterized in that: The real-time monitoring in step S5 includes: Isolating the operations of said AI entity in an AI behavioral sandbox; If an abnormal API call is detected, a secondary biometric verification is triggered.

6. The data security protection method based on anti-AI penetration technology according to claim 1 is characterized in that: The step S6 further comprises: Add a hash digital watermark of the national secret algorithm or other encryption algorithm to the exported data; Set up a multi-level approval process for sensitive data operations and store them on the chain.

7. Data security protection system based on anti-AI penetration technology, characterized by: The data security protection method based on anti-AI penetration technology as described in any one of claims 1 to 6 is configured on the server side, comprising: A receiving module, used to receive user identity data and a registration request for an AI entity sent by a client; An identity generation module, configured to generate an ePASS-ID based on the user identity data and generate an identity identifier for the AI ​​entity; The authority allocation module is used to generate and allocate authority data according to the organizational structure information database; The monitoring module is used to monitor the API call behavior of AI entities in real time and identify anomalies; The data processing module is used to encrypt, transmit and store business data.

8. The data security protection system based on anti-AI penetration technology according to claim 7 is characterized in that: The identity generation module generates an ePASS-ID using a national secret algorithm or other encryption algorithms, and encrypts biometric data using the national secret algorithm or other encryption algorithms.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the data security protection method based on anti-AI infiltration technology as described in any one of claims 1 to 6 is implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by the processor, the data security protection method based on anti-AI infiltration technology as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Intelligent power grid information security protection method

    CN117640207A

  • Data security protection method for photovoltaic power station in smart grid environment

    CN119696899A

  • System for providing zero trust model based seruity management service

    KR102655993B1

  • Systems and methods of security for trusted artificial intelligence hardware processing

    US20200250312A1

  • Ai-enhanced simulation and modeling experimentation and control

    US20240348663A1