Method and system for protecting instructions of a large model based on four-layer cooperative closed loop

The four-layer collaborative closed-loop protection mechanism solves the problems of fragmentation and compliance adaptation in large-scale model security protection, and achieves full-link system protection and cross-model adaptation, which is suitable for high-security scenarios.

CN122153881APending Publication Date: 2026-06-05高鹏
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
高鹏
Filing Date
2026-03-10
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing large-scale security protection solutions suffer from fragmented protection systems, intrusive dependencies, single-modality identity verification, and insufficient compliance adaptability, making it difficult to cope with combined attacks and compliance requirements for cross-border deployments.

Method used

Deploy a non-intrusive independent protection layer and implement a four-layer collaborative closed-loop protection mechanism. Through a shared semantic tag library, dynamic threshold synchronization mechanism and cross-layer anomaly feedback channel, achieve full-link security verification. Combined with dynamic control of instruction weights, multi-modal public authority trusted verification and rigid locking of security rule priorities, a systematic protection is formed.

Benefits of technology

It achieves end-to-end systemic protection, improves protection against weight hijacking, context bypassing and identity impersonation, has cross-model adaptability, meets global compliance requirements, and is suitable for high-security scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122153881A_ABST
    Figure CN122153881A_ABST
Patent Text Reader

Abstract

The application discloses a kind of big model instruction security protection method, system and storage medium based on four-layer cooperative closed-loop protection mechanism, belong to artificial intelligence security and trusted computing field.The application is deployed non-intrusive independent protection layer between user input layer and big model inference execution layer, through instruction weight dynamic control, context integrity check, multi-modal public power subject credible verification, four big links cooperative linkage of security rule priority rigid locking, combined with shared semantic label library, dynamic threshold value synchronization mechanism and cross-layer abnormal feedback channel constructs whole-link closed-loop protection system.The application realizes high semantic weight content attenuation by instruction weight system, resists split bypass attack by context governance, guarantees public power identity credible by multi-modal double verification, realizes permission control and algorithm transparent by rule priority rigid locking, built-in global multi-jurisdiction compliance rule library, adapt to global regulatory requirements and have anti-black-box auditable characteristics.The application does not need to modify model bottom code, supports hardware-level isolation and national encryption algorithm acceleration, can be widely applied to government, finance, military industry and other high-end security scene, provides non-intrusive, global, high-security compliance protection solution for big model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence security and trusted computing technology, specifically to a method, system, storage medium, and application for large model instruction security protection based on a four-layer collaborative closed-loop protection mechanism.

[0002] Terminology Definition To ensure consistent terminology and clarify the scope of protection of this invention, the core technical terms involved in this invention are defined as follows: 1. Four-layer collaborative closed-loop protection mechanism: refers to a full-link, closed-loop, and systematic protection system formed by the collaborative linkage of four core links: dynamic control of instruction weight, context integrity verification, multi-modal public authority trust verification, and rigid locking of security rule priority. 2. Non-intrusive independent protection layer: This refers to a secure middleware deployed independently between the user input layer and the large model inference execution layer. Its core features are that it does not modify the underlying model code, does not intrude on the model inference process, and does not fine-tune the model weights, thus achieving decoupled deployment from the model itself. 3. High semantic weight content: refers to text content that has a significant influence weight in the attention mechanism of large models, mainly including legal provisions, regulatory statements, compliance instructions, and commitments of responsibility, which are easily exploited to carry out weight hijacking attacks. 4. Context dilution attack: This refers to the attacker's behavior of splitting malicious core instructions into multiple fragments, nesting inputs, or diluting and hiding them through long texts in order to reduce the security detection's ability to identify malicious intent. 5. Dual verification rule: This refers to the rule that when verifying the identity of a public authority (including certificates, official seals, and official letters), it is necessary to simultaneously satisfy visual feature matching, text structured feature matching, and cross-verification by an authoritative and credible feature database. 6. Rigid Security Rule Priority Locking Mechanism: This mechanism solidifies platform security control rules and mandatory industry regulatory requirements of corresponding jurisdictions as the highest priority for enforcement. It ensures that system policy weights are always higher than the semantic weights of user input. 7. Dynamic weight baseline: refers to the upper limit of instruction semantic weight security that is adaptively adjusted in real time according to the specific application scenario, industry attributes, and regulatory requirements of the target judicial jurisdiction. 8. Incremental Context Reorganization: A processing mechanism that integrates semantic associations and completes logic for multi-turn dialogues and fragmented batch input. 9. Tamper-proof and audit-traceable: This refers to data such as system logs, rule configurations, and permission changes, which not only possess the physical immutability of conventional software modifications, but also require approval processes and retain complete operation traces for compliance management. 10. Jurisdiction: refers to a country, region, or specific legal domain that has an independent legal and data security regulatory system for artificial intelligence. 11. Typical deployment scenarios: refers to the deployment environment of conventional commercial servers using general-purpose x86 servers and single or multiple GPUs, representing the mainstream engineering implementation capabilities. 12. Shared Semantic Tag Library: This refers to a unified database used to store semantic feature identifiers, risk tags, and compliance rule identifiers shared across layers. It provides a unified semantic recognition standard and security strategy basis for the four-layer verification process and supports real-time updates and full-layer synchronization of tags. 13. Dynamic threshold synchronization mechanism: This refers to a mechanism that automatically adjusts the security thresholds of each verification layer based on real-time risk assessment results through a preset adaptive algorithm, and achieves coordinated linkage of thresholds at each layer by sharing feature vectors, thereby ensuring the consistency of the security strategy across the entire chain. 14. Cross-layer anomaly feedback channel: refers to a dedicated data link used to transmit verification results, anomaly events, and risk level data between different verification layers / functional modules. It can synchronize the verification results of any layer to all other layers, triggering full-link policy adjustments or secondary iterative verification to achieve closed-loop protection. Background Technology

[0016] The large-scale deployment of big data models in government offices, financial risk control, judicial rulings, medical assistance, and cross-border trade has made model security, access control, and full-process auditability core requirements for regulatory agencies and industry clients.

[0017] Currently, mainstream large-scale model security protection solutions fall into two categories: one is single-point keyword filtering and text similarity matching, which can only intercept known attacks and cannot cope with combined bypass attacks; the other is multi-layer independent verification solutions, where there is no data linkage or closed-loop feedback between protection links, and the overall protection fails once one layer is bypassed, making it unable to cope with advanced attacks such as weight hijacking and context splitting. Faced with increasingly complex large-scale model attacks and compliance challenges, existing solutions generally suffer from the following technical deficiencies: 1. Fragmented protection system: Most systems adopt single-point defense and multi-layer independent verification, without forming a closed-loop linkage mechanism across the entire chain, making it difficult to cope with combined attacks and bypass strategies. This invention addresses this pain point by using a four-layer collaborative closed-loop protection system. 2. Strong intrusive dependency: Some solutions require modification of the model's internal structure or embedding of the inference process, resulting in poor versatility, high migration costs, and inability to adapt to large models with multiple architectures. This invention addresses this pain point by using a non-intrusive independent protection layer. 3. Lack of independent security link: The external protection capability is weak and there is a lack of independent management and control interface that can be audited throughout the process, making it difficult to meet the compliance requirements of high security level scenarios. In response, this invention solves this pain point by using an independent audit module and hardware-level isolated deployment. 4. Single-modal identity verification: The authentication of public authorities often relies on text or single visual verification, which is vulnerable to forgery and impersonation. This invention addresses this pain point by using multimodal dual verification rules. 5. Lack of weight control: The lack of effective constraints on content with high semantic weight makes it extremely easy for weight hijacking attacks to occur, leading to permission overreach. This invention addresses this pain point by dynamically controlling instruction weight and performing compliance calibration. 6. Insufficient compliance adaptability: Existing solutions are difficult to meet the dynamic adaptation and compliance requirements of multiple jurisdictions in cross-border deployment scenarios. This invention addresses this pain point by using a global compliance rule base and a dynamic adaptation engine.

[0023] Therefore, the industry urgently needs a non-intrusive, architecture-level, systematic, and globally adaptable large-scale instruction security protection solution to address the aforementioned technical pain points. Summary of the Invention

[0024] Large Model Instruction Security Protection Methods This invention deploys a non-intrusive, independent protection layer between the user input layer and the large model inference execution layer. It performs full-link security verification of the user-submitted instructions using a four-layer collaborative closed-loop protection mechanism, achieving closed-loop protection from the source to the execution end.

[0025] The four-layer collaborative closed-loop protection mechanism of this invention provides a unified semantic recognition standard for each layer through a shared semantic tag library. A dynamic threshold synchronization mechanism adjusts the security thresholds of each layer synchronously based on real-time risk assessment. A cross-layer anomaly feedback channel enables cross-layer transmission of verification results and policy linkage between any two layers. For example, when S3 detects a high-risk event suspected of being an impersonator of a public authority, the risk level is synchronized to S1, S2, and S4 through the cross-layer anomaly feedback channel. The dynamic threshold synchronization mechanism simultaneously tightens the security thresholds of each layer: S1 strengthens the downweighting of high-semantic-weight content, S2 increases the sensitivity of contextual relevance verification, and S4 strengthens rule priority locking, triggering a second iteration verification throughout the entire process, forming a closed-loop protection from risk identification to end-to-end response.

[0026] When any layer of verification detects a medium-to-high-risk event, or when there are discrepancies in the verification results from multiple engines, the system can trigger a full-process iterative verification. During the iteration process, security thresholds and verification strategies can be adjusted based on the verification results of the previous round, until the risk is eliminated or the instruction is intercepted. Each layer of verification can choose to execute sequentially, in parallel, or iteratively, depending on actual security requirements. The verification result of any layer can be used as an input parameter or decision basis for any other layer of verification.

[0027] The specific steps are as follows: S1. Dynamic Control and Compliance Calibration of Command Weights: Multi-dimensional semantic weight analysis is performed on user-input commands to accurately identify and extract high-semantic-weight content such as legal provisions, regulatory statements, and compliance texts. Attention-based weighted attenuation or vector normalization is applied to this high-weight content, strictly limiting the upper limit of the command's semantic weight to a system-preset dynamic weight baseline range that is adaptable to the regulatory requirements of different global jurisdictions. This step blocks privilege escalation attacks using weight aggregation at the source, and does not collect sensitive user privacy data throughout the process, thus meeting data compliance requirements.

[0028] S2. Collaborative Verification of Context Integrity and Malicious Segmentation: For user-inputted segmented, fragmented, nested, or long-text diluted contexts, incremental semantic reorganization and cross-segment association analysis are performed. Semantic matching engines and cross-modal association verification engines are used to calculate the semantic relevance and logical integrity between segments. When the semantic relevance of any segment consistently falls below a preset security threshold, the system determines it as context dilution evasion behavior, immediately intercepts the abnormal command, and records an immutable audit log, effectively preventing malicious commands from bypassing security detection through segmentation.

[0029] S3. Multimodal Dual Verification of Public Authority Credibility: For text and image data submitted by users, including the name, credentials, official seal, and official letters of public authorities, visual local features, textual structure features, and spatiotemporal features are extracted simultaneously. These extracted multi-dimensional features are cross-validated with a dynamically updated authoritative and credible feature library within the corresponding judicial jurisdiction. The system strictly enforces the dual verification rules of visual and textual verification, prohibiting the recognition of public authority identities based solely on single textual semantics. Requests that fail dual verification are directly rejected for privilege escalation operations based on the public authority's identity, and sensitive data is never transferred across borders throughout the process.

[0030] S4. Rigid Priority Locking of Security Control Rules: This mechanism parses the legal semantics and persuasive instructions in user input, solidifying the platform's security control rules and the corresponding industry regulatory requirements of their respective jurisdictions into an immutable, highest-priority execution basis. This ensures that the system's security policy weight always exceeds the semantic weight of any user input text. The final execution of all instructions is based on the built-in highest-priority rule, thereby preventing attempts to override system control rules through legal or misleading language. This rigid priority locking mechanism can be implemented via hardware, making it particularly suitable for high-security scenarios involving physically isolated confidential internal networks.

[0031] Further optimization of the plan: In S1, semantic importance is calculated based on the attention mechanism of the pre-trained language model, and a dynamic baseline database is maintained to achieve adaptive weight reduction and traceable auditing. The dynamic weight baseline database supports the installation of industry-specific compliance rule packages, and weight thresholds can be configured independently for different scenarios. - In S2, incremental context reorganization for multi-turn dialogues is supported, effectively identifying cross-turn context dilution attacks and ensuring the security of long sessions; multi-engine collaborative mode is supported, which uses at least two of the following engines: pre-trained semantic matching engine, incremental context reorganization engine, and cross-modal context association verification engine. When any engine detects an anomaly, regardless of whether the comprehensive association index meets the standard, the interception mechanism is immediately triggered to avoid missed detection by a single engine. - In S3, cross-domain matching of anti-counterfeiting features, format specifications and filing information of certificates and official seals is supported, and the authoritative database supports real-time online updates; feature extraction and comparison are completed locally, and only the encrypted feature hash value is transmitted to the central database for matching. - In S4, an independent rule engine is built, with the highest priority rules built-in and fixed, supporting the adaptation and approval record of regulatory rules in multiple jurisdictions; rule version management complies with the requirements of Section 5.2 of ISO / IEC 23053 standard, supports the export of rule change lists in SBOM format, and meets the traceability review requirements of regulatory agencies. - In typical deployment scenarios, the latency of the entire detection process is controlled at the millisecond level, which has a negligible impact on the native inference performance of the model.

[0036] Large Model Instruction Security Protection System This invention provides a large model instruction security protection system that implements the above method. Specifically, it is a non-intrusive architecture-level security protection gateway that is deployed in series between the user input layer and the large model inference execution layer.

[0037] The system comprises four core functional modules. These modules achieve closed-loop linkage through a shared semantic tag library, a dynamic threshold synchronization mechanism, and a cross-layer anomaly feedback channel. The verification result of any module can be used as an input parameter or decision basis for any other module, collaboratively forming a complete closed-loop protection system. 1. Instruction Weight Detection and Calibration Module: Responsible for executing step S1, with a built-in global compliance rule base, enabling adaptive weighting across multiple scenarios. 2. Context integrity verification module: responsible for executing S2 steps, supporting multi-engine collaborative verification and multi-round session reassembly. 3. Public Authority Entity Credible Identification Module: Responsible for executing the S3 steps, connecting to the authoritative filing database, and realizing multimodal dual verification. 4. Legal semantic priority locking module: responsible for executing the S4 steps to achieve rigid rule locking and multi-jurisdictional adaptation.

[0041] System deployment types and characteristics: It supports various forms such as software plug-ins, cloud-native security components, hardware isolation devices, chip-embedded logic, FPGA acceleration modules, or integrated hardware and software systems.

[0042] The hardware isolation device includes an external security gateway, a PCIe physical isolation card, an independent security chip, and a hardware acceleration unit. It can run protection logic in an independent environment, independent of the host CPU / GPU resources, supporting transparent plug-and-play and offline rule updates, and is suitable for physically isolated classified intranet environments. The hardware isolation device has a built-in national cryptographic algorithm acceleration engine, supporting SM2, SM3, and SM4 national cryptographic algorithms. This engine features dedicated pipeline optimization for the entire data processing of four-layer collaborative closed-loop protection. It uses the SM4 algorithm for parallel fragmentation encryption of multimodal feature data, the SM3 algorithm for chain hashing and solidification of audit logs, and the SM2 algorithm for asymmetric encryption and signature verification of rule configurations and keys. The entire encryption process is completed in the hardware isolation environment without consuming host computing resources.

[0043] System core capabilities: It features a built-in tamper-proof and audit-tracking log module. This module uses a consortium blockchain architecture for storage; each log record has its hash value calculated synchronously after generation and then solidified into a block using a Merkle tree structure. Multiple trusted nodes synchronize data through a practical Byzantine fault-tolerant consensus mechanism, ensuring the logs cannot be tampered with. It also supports regulatory agencies verifying log authenticity through nodes and can be integrated with regulatory platforms. Multi-tenant isolated deployment is supported. The instruction weight module has a built-in global compliance rule base; the identity verification module is cross-regional adaptable; and the rule locking module supports canary updates. In typical scenarios, it supports high instruction concurrency and low-latency detection, making it suitable for high-concurrency businesses.

[0044] This system can serve as a mandatory security middleware for the compliant deployment of large-scale overseas models in China, which helps to enhance the country's independent and controllable capabilities in AI security.

[0045] Non-transitory computer-readable storage medium This invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When executed by a processor, the program implements the large model instruction security protection method described in this invention. The program can be embedded in a large model service platform, browser plugin, endpoint security software, or cloud security service to achieve adaptable deployment across multiple terminals and scenarios, without requiring modification of the original large model program.

[0046] Industry Applications This invention has a wide range of applications, specifically for strengthening the protection against privilege escalation in large models, protecting against impersonation of public authorities, defending against injection attacks, managing compliance in cross-border AI businesses, and high-security scenarios such as government affairs, finance, healthcare, judiciary, military secrets, and cross-border trade.

[0047] This invention can be widely applied to core scenarios such as defense against large model hint injection attacks, security enhancement against privilege escalation, verification of the authenticity of public authority entities, and legal semantic security compliance management. It is particularly suitable for high-security industry scenarios such as government affairs, finance, healthcare, judiciary, military-related classified information, and cross-border trade. This invention is compatible with mainstream large model architectures such as Transformer, RLHF, and DPO, and can independently adapt to the regulatory compliance requirements of major jurisdictions worldwide, including the EU's AI Act, China's Interim Measures for the Administration of Generative Artificial Intelligence Services, and the US Artificial Intelligence Risk Management Framework (AIRF), possessing global technical universality and compliance.

[0048] This invention is compatible with mainstream global AI regulatory requirements, and its data processing complies with local regulations. It can provide security and compliance reinforcement for large model service providers, provide private security deployment solutions for enterprises, and provide auditable AI security management tools for regulatory agencies. Beneficial effects

[0049] This invention has substantial features and significant progress compared to the prior art: 1. Full-link systematic protection: By constructing a four-layer collaborative closed-loop protection mechanism through a shared semantic tag library, dynamic threshold synchronization mechanism, and cross-layer anomaly feedback channel, it solves the defects of fragmented and uncoordinated single-point protection in existing technologies, effectively covers core attack vectors such as weight hijacking, context bypass, identity impersonation, and rule overriding, and significantly improves protection effectiveness. 2. Non-intrusive and highly versatile: It adopts an external protective layer deployment method independent of the large model body, does not depend on the internal structure of the model, does not require modification of the underlying code, and does not require retraining or fine-tuning of the model. It has universal adaptability across models, platforms, and architectures, is easy to deploy, and has minimal impact on the performance of existing systems. 3. Global Compliance Adaptation: Built-in multi-jurisdictional compliance rule library, through dynamic weighted baseline and rule adaptation engine, supports dynamic adaptation to global regulatory rules such as the EU AI Act and China's Interim Measures for the Administration of Generative AI Services, meeting the compliance needs of multinational enterprises for cross-border deployment. 4. High security level and strong feasibility: It supports various high-security deployment forms such as hardware isolation devices, hardware acceleration of national cryptographic algorithms, and FPGA chip-level solidification, fully covering high-end industry scenarios such as government affairs, finance, and military industry, and meeting high-level confidentiality and compliance requirements. 5. Data security and privacy protection: The entire process supports local feature extraction and encrypted hash transmission, does not collect sensitive user privacy data, conforms to global data compliance trends, and effectively reduces compliance risks for enterprises and service providers. 6. High standardization and replicability: It has built a non-intrusive, cross-model compatible standardized large-scale model security protection paradigm, which can be quickly adapted to the deployment needs of different industries and different models, and has broad industry promotion value. Attached Figure Description

[0055] Figure 1 This is a schematic diagram of the overall deployment architecture of the system of the present invention; Figure 2 This is a flowchart of the four-layer collaborative closed-loop protection method of the present invention; Figure 3 This is a schematic diagram of the hardware isolation device of the present invention; Figure 4 This is a schematic diagram of the process for dual verification of multimodal public authority subjects in this invention.

[0056] (Note: When submitting the actual document, corresponding drawings must be provided. The drawings are only schematic diagrams and do not limit the scope of protection of this invention.) Detailed Implementation

[0057] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention.

[0058] Example 1: High-security scenario for government office work Application scenario: A provincial government service center deploys a large-scale model for internal office work and approval assistance.

[0059] Deployment method: A hardware and software integrated security gateway is deployed at the boundary between the government intranet and the Internet.

[0060] Implementation process: 1. The system receives user input commands, automatically identifies government regulations with high semantic weight, performs attention decay, and limits the weight to a government-specific baseline. 2. Perform incremental context reorganization on multi-round approval workflow texts to detect whether there is malicious instruction splicing across rounds. 3. Connect to the authoritative database of the National Government Service Platform to conduct multimodal dual verification of the identity documents and official seals submitted by staff. 4. Set the mandatory rules for government supervision as the highest priority for execution to ensure that the instructions do not exceed the staff's authority and the red lines of the law. 5. When the system verifies an abnormal official seal document, it synchronizes high-risk information through a cross-layer abnormality feedback channel, and the dynamic threshold synchronization mechanism simultaneously tightens the security threshold throughout the entire process, triggering a second full-link iterative verification to form a closed-loop risk response and completely block privilege escalation attacks impersonating government entities.

[0065] Results: Detection latency ≤5ms, performance impact ≤0.8%, meeting the government's "zero violations" and low latency requirements.

[0066] Example 2: Cross-border Financial Business Scenario Application scenario: Multinational banks use large-scale overseas models to provide cross-border financial services within China.

[0067] Deployment model: Cloud-native security components are deployed at the edge of the cloud platform.

[0068] Implementation process: 1. Based on dynamic baselines across multiple jurisdictions (EU, China, and the United States), weighted compliance calibration is performed on cross-border directives. 2. Utilize a dual-engine approach (semantic matching + cross-modal verification) to detect the integrity of long text transaction instructions. 3. Conduct cross-domain authoritative verification of official letters and license documents involving financial regulation. 4. Prioritize global financial regulatory rules to mitigate compliance risks. 5. When compliance risks of cross-border instructions are detected, the weight calibration module is synchronized through the cross-layer anomaly feedback channel to dynamically adjust the weight baseline of the corresponding jurisdiction, triggering a second verification throughout the entire process to ensure compliance with local regulatory requirements.

[0073] Results: Supports tens of thousands of concurrent requests per second with a latency of ≤10ms, meeting the compliance and high concurrency requirements of cross-border business.

[0074] Example 3: Military-related classified FPGA hardware scenario Application scenario: Deployment of classified large-scale models within military research units.

[0075] Deployment method: FPGA hardware acceleration isolation device is used to achieve physical isolation protection.

[0076] Implementation process: 1. The protection logic is completely embedded in the FPGA chip and deployed in the PCIe slot of the classified server. 2. Perform hardware-level weight decay and context verification on all inputs. 3. Built-in national cryptographic SM4 / SM3 algorithm acceleration engine, hardware encryption of feature data and audit logs, with the key stored in HSM. 4. Supports offline rule package updates without connecting to the public network, ensuring the ultimate security of confidential data. 5. The four-layer verification achieves data linkage through a hardware-level shared tag library. When a confidential or illegal instruction is detected, the hardware triggers full-link interception and auditing, without the need for host CPU participation, ensuring that the protection logic cannot be tampered with.

[0081] Results: Detection delay ≤2ms, performance impact ≤0.5%, meeting the high-level confidentiality and auditing requirements of the military industry.

[0082] Example 4: Enterprise-level SaaS platform scenario Application scenario: Vertical industry SaaS platforms provide large model API services.

[0083] Deployment method: Embedded in the platform gateway as a software plug-in.

[0084] Implementation process: 1. Implement weighted baseline isolation for multi-tenant user commands to prevent unauthorized access between tenants. 2. Real-time detection and interception of hint injection attacks targeting the underlying model. 3. Provides visual audit logs and compliance reports. 4. When a tenant triggers a high-risk interception, the security verification threshold for that tenant is tightened in a targeted manner through a dynamic threshold synchronization mechanism, without affecting the normal use of other tenants, thus achieving multi-tenant isolation and protection.

[0088] Results: Zero-code intrusion into the underlying model, rapid deployment, and significant reduction in platform security and maintenance costs.

[0089] Example 5: Privacy Protection Scenarios in the Medical Industry Application scenario: Large-scale models within hospitals to assist in medical record analysis and scientific research.

[0090] Deployment type: Deployed on the intranet server within the hospital.

[0091] Implementation process: 1. Strictly adhere to the medical industry's specific compliance rules and calibrations to protect patient privacy data. 2. Perform context integrity checks on the input text to prevent malicious data theft. 3. Verify the identity of medical staff using multimodal methods to ensure that operational permissions are compliant. 4. No sensitive patient privacy data is collected throughout the entire process. All feature extraction and verification are completed locally within the hospital, and only the encrypted audit hash value is synchronized to the hospital's monitoring platform.

[0095] Results: Meets medical data security and privacy protection regulations, with no risk of patient privacy data leakage.

[0096] Example 6: Evidence Preservation Scenarios in the Judicial Industry Application scenario: Courts and law firms use large-scale models for case retrieval and document generation.

[0097] Deployment method: Physically isolated security gateway.

[0098] Implementation process: 1. Establish specific regulatory rules for the judicial industry to ensure compliance in case handling. 2. Full-process operation auditing to meet the requirements of complete evidence chain. 3. When a query command for information on a violation case is detected, risk information is synchronized through a cross-layer anomaly feedback channel, triggering full-link interception. At the same time, the operation record is solidified into the blockchain audit log to ensure that the operation is traceable and tamper-proof.

[0101] Results: Strict access control and audit trails ensure the security of case information and judicial fairness.

Claims

1. A large-scale instruction security protection method based on a four-layer collaborative closed loop, characterized in that, The method deploys a non-intrusive protection layer independent of the large model itself between the user input layer and the large model inference execution layer, performing a four-layer linked end-to-end security check on user input commands. Each layer's check achieves closed-loop data linkage through a shared semantic tag library, a dynamic threshold synchronization mechanism, and a cross-layer anomaly feedback channel, providing a unified security strategy basis for each layer's check. The check result of any layer can be used as an input parameter or decision basis for any other layer's check. Each layer's check can choose to execute sequentially, in parallel, or iteratively, forming a closed-loop feedback protection mechanism from the source to the execution end, based on actual security needs. The semantic tag library stores shared semantic feature identifiers across layers, the dynamic threshold synchronization mechanism adjusts the security thresholds of each layer based on real-time risk assessment, and the cross-layer anomaly feedback channel ensures the consistency and coherence of the security strategy, achieving auditable and traceable operation throughout the entire process. The method includes the following steps: S1. Dynamic Control and Compliance Calibration of Command Weights: Perform multi-dimensional semantic weight analysis on user-input commands to identify and extract high semantic weight content such as legal provisions, regulatory statements, liability commitments, and compliance texts; perform attention mechanism weighted attenuation or vector normalization processing on the high semantic weight content to limit the upper limit of the semantic weight of user commands to within the system's preset security baseline weight range. The security baseline is dynamically adjusted according to the regulatory requirements of the jurisdiction, blocking privilege escalation attacks achieved through weight superposition from the source. S2. Collaborative Validation of Context Integrity and Malicious Splitting: For user-input segmented, fragmented, nested, and long-text diluted contexts, perform incremental semantic reorganization and cross-fragment association analysis to calculate the semantic matching degree and logical integrity between each context fragment; when the semantic association degree between fragments is lower than the preset security threshold range, it is determined to be a context dilution or malicious splitting evasion behavior, and abnormal instructions are immediately intercepted and recorded in the audit log; S3. Dual Verification of the Credibility of Public Authority Entities: For text and image content submitted by users that includes the name, certificate, official seal, and official letter of a public authority, visual local features, text structure features, and spatiotemporal features are extracted simultaneously. The extracted multi-dimensional features are cross-validated with a dynamically updated compliant, authoritative, and credible feature library within the corresponding judicial jurisdiction. The dual verification rules of visual and text are strictly enforced, and it is prohibited to recognize the user's public authority entity identity solely based on text semantics. If the dual verification fails, the request for privilege escalation based on the public authority entity identity is directly rejected. S4. Rigid Locking and Compliance Adaptation of Security Control Rules Priority: The legal provisions, regulatory definitions, and rule interpretations entered by users are analyzed for weight and priority; the platform's security control rules and the mandatory industry regulatory requirements of the corresponding jurisdictions are solidified as the highest priority execution basis that cannot be tampered with, ensuring that the weight of the system security policy is always higher than the semantic weight of any user input text; all instruction execution is based on the built-in highest priority rule as the final basis, blocking the behavior of overriding the system control rules through legal semantics or wording.

2. The method according to claim 1, characterized in that, The semantic weight analysis described in step S1 includes: calculating the semantic importance and global influence of text segments based on the attention mechanism of a pre-trained language model; establishing and maintaining a dynamic user instruction weight baseline database, wherein the baseline can be adaptively adjusted in real time according to user interaction scenarios, instruction risk levels, industry attributes, and regulatory requirements of the corresponding jurisdiction; comparing user instruction weights with dynamic baseline thresholds in real time, and automatically triggering weight reduction processing for instructions exceeding the threshold range, with the weight reduction process being traceable and auditable; and not collecting any sensitive user privacy data throughout the entire process of step S1.

3. The method according to claim 1, characterized in that, The semantic relevance calculation in step S2 includes: using a semantic similarity calculation engine, combined with a pre-trained semantic matching model, to calculate the semantic similarity between context segments, using at least one of cosine similarity, Euclidean distance, and Jaccard coefficient as a quantification index; the semantic similarity calculation engine integrates at least one of the functions of a pre-trained semantic matching engine, an incremental context reconstruction engine, and a cross-modal context association verification engine, with each engine calculating the relevance index in parallel, and the system using a weighted fusion algorithm to synthesize the results of each engine. When the comprehensive index is lower than a preset threshold, it is determined to be a context dilution or malicious splitting behavior and is intercepted, and the interception record is synchronously written to an immutable audit log.

4. The method according to claim 1, characterized in that, The context integrity verification described in step S2 supports incremental context reorganization in multi-turn dialogues. It can perform full-cycle semantic association analysis on fragmented instructions input by users in batches and multiple turns, and identify cross-turn context dilution attacks.

5. The method according to claim 1, characterized in that, The multimodal credibility verification in step S3 includes: multi-dimensional cross-domain matching of the anti-counterfeiting features, format specifications, and filing information of certificates and official seals; the authoritative and credible feature library supports real-time online updates to adapt to the version iteration and compliance requirements of the identity identifiers of public authorities in the corresponding judicial jurisdictions; the identity of the public authority can only be determined to be credible when both visual features and text features pass verification and there is a valid match with the compliant authoritative and credible feature library; the feature extraction and comparison in step S3 are completed locally, and only the encrypted feature hash value is transmitted to the central database for matching.

6. The method according to claim 1, characterized in that, The rigid priority locking mentioned in step S4 includes: constructing a priority rule engine independent of the large model ontology, with the highest priority rule built into the engine, and no user input or external command can overwrite, modify, or reduce the execution weight of the highest priority rule; the rule engine supports compliance adaptation and adjustment according to the regulatory policies of different jurisdictions, and the adjustment process must be approved by the relevant authorities and fully audited.

7. The method according to claim 2, characterized in that, The dynamic weighted baseline database supports the inclusion of industry-specific compliance rule packages, enabling independent configuration of weight thresholds, protection strategy depth, and compliance verification standards for classified scenarios in government, finance, healthcare, judiciary, and military industries, achieving scenario-based customization and adaptation. The method's end-to-end detection and processing latency adapts to the low-latency requirements of real-time interactive business scenarios, without affecting the user's normal interactive experience and business continuity.

8. The method according to claim 3, characterized in that, The context integrity verification supports a multi-engine collaborative mode, which uses at least two of the following engines: a pre-trained semantic matching engine, an incremental context reorganization engine, and a cross-modal context association verification engine. When any engine detects an anomaly, the interception mechanism is triggered immediately, regardless of whether the comprehensive association index meets the standard.

9. A large-scale instruction security protection system for implementing the method of any one of claims 1-8, characterized in that, The system is a non-intrusive security gateway deployed in series between the user input layer and the large model inference execution layer. It includes functional modules that implement four-layer collaborative closed-loop protection logic. These modules are divided into an instruction weight detection and calibration module, a context integrity verification module, a public authority trust identification module, and a legal semantic priority locking module. Each module achieves closed-loop linkage through a shared semantic tag library, a dynamic threshold synchronization mechanism, and a cross-layer anomaly feedback channel. The verification result of any module can be used as an input parameter or decision basis for any other module. The modules operate collaboratively, achieving full-process instruction security detection and interception without modifying the underlying architecture code of the large model or retraining or fine-tuning the model weights. The system has a built-in, tamper-proof audit log module, making the entire operation process traceable and auditable.

10. The system according to claim 9, characterized in that, The system can be implemented in the following forms: pure software function plug-in, cloud-native security component, hardware isolation device, chip-level solidified logic, field-programmable gate array acceleration module, or integrated protection system combining software and hardware.

11. The system according to claim 10, characterized in that, The hardware isolation device includes an external security gateway, a physical isolation card, an independent security chip, and a hardware acceleration module. The hardware isolation device is connected to the host through a general high-speed interface, supports hot-swappable deployment, and is independently deployed at the network boundary or device boundary between the user terminal and the large model inference server. It has an independent computing unit, storage unit, and communication interface, and can complete all the protection logic described in any one of claims 1-8 in an independent environment free from host resource occupation.

12. The system according to claim 11, characterized in that, The hardware isolation device supports transparent plug-and-play deployment, and the deployment process does not affect the normal operation of existing large model services; the device is based on a hardware parallel computing architecture, which can achieve low-latency processing of instruction detection.

13. The system according to claim 11, characterized in that, The hardware isolation device supports offline updates of compliance rule packages and authoritative and trusted feature libraries. The update process does not require connection to the public network or modification of the device's core protection logic. The entire process is audited and left a trace, making it suitable for deployment scenarios involving physically isolated classified internal networks.

14. The system according to claim 9, characterized in that, The audit log module uses blockchain technology for storage and synchronizes log data through a distributed node network. The hash value of each record is fixed by a Merkle root hash. Each node reaches consensus by verifying the consistency of the Merkle root hash. The consensus mechanism is selected from at least one of Practical Byzantine Fault Tolerance, Proof-of-Stake, and Proof-of-Work to ensure the immutability and verifiability of the logs. The audit log can record the detection process, verification results, interception actions, and rule adaptation details of all instructions. The logs can be exported and connected to the regulatory compliance platforms of the corresponding jurisdictions to meet the AI ​​regulatory audit requirements of different regions around the world.

15. The system according to claim 9, characterized in that, The system supports multi-tenant isolation deployment and can independently configure protection policies, compliance rule packages and security thresholds for different tenants and different business scenarios, meeting the compliance deployment needs of large enterprises with multiple business lines and branches in multiple regions.

16. The system according to claim 9, characterized in that, The instruction weight detection and calibration module has a built-in jurisdictional rule adaptation engine. By parsing the geographic identification information in the user request, it dynamically loads the corresponding regulatory rule set. The rule set includes differentiated compliance weight thresholds, cross-border data restrictions, and localization processing strategies of the EU AI Act, the US NIST AI Standard, and China's "Interim Measures for the Administration of Generative AI Services," ensuring that instruction verification complies with the mandatory requirements of the target jurisdiction. The system is compatible with all generative artificial intelligence model architectures, including large-scale artificial intelligence models based on reinforcement learning alignment mechanisms and direct preference optimization alignment mechanisms, which are based on human feedback.

17. The system according to claim 9, characterized in that, The aforementioned public authority credible identification module supports connection to the authoritative filing database of public authorities within the corresponding judicial jurisdiction, enabling cross-regional verification of compliant identity credibility, and is applicable to the compliance requirements of cross-border government affairs and cross-border financial business.

18. The system according to claim 9, characterized in that, The legal semantic priority locking module supports versioned management and gray-scale updates of regulatory rules. The rule version management complies with the requirements of Section 5.2 of ISO / IEC 23053 standard and supports rule change traceability in SBOM format. Each rule change requires digital signature verification. The update record includes the change time, approver information, and compliance basis, supporting traceability review by regulatory agencies. It enables smooth switching of regulatory rules in different jurisdictions, and the update process is fully traceable, approvable, and rollbackable.

19. The system according to claim 9, characterized in that, The system is applicable to large third-party models developed overseas, and can achieve compliant deployment and security protection of overseas models in China without modifying the underlying structure of the overseas models or retraining them.

20. The hardware isolation device according to claim 11, characterized in that, The hardware isolation device incorporates a national cryptographic algorithm acceleration engine, supporting SM2, SM3, and SM4 national cryptographic algorithms. This engine is specifically optimized for end-to-end data processing in the four-layer collaborative closed-loop protection system, including: sharded parallel SM4 encryption for multimodal feature data, SM3 chained hashing for audit logs, and SM2 asymmetric encryption for rule configuration and key management. It also supports full lifecycle management of key generation, distribution, use, and destruction, and automatic key rotation based on periodic or risk-based strategies. Keys are managed by a hardware security module (HSM) compliant with national cryptographic standards. The four-layer collaborative closed-loop protection logic, security control rules, and national cryptographic encryption engine are implemented through hardware embedding. Electromagnetic leakage protection meets the requirements of the BMB20-2007 standard, and the physical isolation strength meets the graded protection requirements for classified information systems, making it suitable for compliance requirements in military, government, and classified scenarios.

21. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements the large model instruction security protection method according to any one of claims 1-8.

22. The computer-readable storage medium according to claim 21, characterized in that, The computer program can be embedded in large model service platforms, browser plugins, terminal security software, and cloud security services to achieve multi-terminal and multi-scenario adaptation and deployment without modifying the original large model program.

23. The application of the protection method according to any one of claims 1-8, or the protection system according to any one of claims 9-20, in the field of artificial intelligence security, characterized in that, It is applied to the reinforcement of privilege escalation prevention in large models, the enhancement of security alignment strategies, the protection against impersonation of public authorities, the defense against injection attacks, the compliance management of cross-border AI business, and the security management of instructions in government, finance, medical, judicial, military, and cross-border trade scenarios involving classified information.

24. The application according to claim 23, characterized in that, The applications further include: adapting and deploying large overseas models to domestic compliance, hardware-level security hardening in physically isolated environments of classified intranets, and end-to-end data encryption protection based on national cryptographic algorithms.