A defense method and system for large language models with context stripping and dual-track isolation

CN122615872BActive Publication Date: 2026-09-18CHANGCHUN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611095756.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-23
Publication Date
2026-09-18
Estimated Expiration
2046-07-23

AI Technical Summary

Technical Problem

[0006]本发明为克服现有技术中单轨推理架构下的大模型安全防护方案,仅依靠文本语义规则与模型微调实现风险识别,未解决显存KV-cache交叉污染、内存资源占用高、算力传输开销大等计算机底层硬件缺陷,难以在复杂诱导上下文场景下同时实现高风险拦截能力、低误拦率与稳定的推理硬件运行性能等问题

Benefits of technology

本发明所述的上下文剥离与双轨隔离的大语言模型防御方法采用双轨独立GPU显存分区物理隔离架构,硬件层面切断目标生成轨与独立审查轨之间KV-cache张量、会话内存的数据互通通道,从根源避免诱导性上下文缓存张量跨分支污染模型注意力表征,解决单轨架构下模型受诱导语境干扰导致战略级敏感内容漏判的缺陷,大幅提升复杂对抗、多轮诱导场景下安全判定稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122615872B_ABST
    Figure CN122615872B_ABST
Patent Text Reader

Abstract

A method and system for defending large language models using context stripping and dual-track isolation is disclosed, relating to the field of large language model security protection. This addresses the shortcomings of existing large model security protection schemes under single-track inference architectures, which rely solely on textual semantic rules and model fine-tuning for risk identification, resulting in high memory resource consumption and significant computational overhead. The method includes: constructing a dual-GPU physically isolated inference architecture; receiving user input token streams and storing them in a process memory buffer; distributing the token stream to the target generation track process in the first GPU memory partition via hardware routing; loading KV-cache tensors for inference to generate initial token text and temporarily storing it in a streaming sharded memory cache; and using a context stripping hardware module to strip the inducing context memory data block, extracting core semantic data and sending it to a physically isolated second GPU memory partition for independent review. Security determination is performed based on the core semantics, and the cached text is then processed according to the results, performing operations such as allowing, clearing, or rewriting, thus optimizing the inference hardware resource overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model security protection technology, specifically to a large language model defense method and system that combines context stripping and dual-track isolation. Background Technology

[0002] Large language models are widely used in scenarios such as intelligent question answering, code assistance, and enterprise knowledge base retrieval. To suppress the output of illegal and sensitive content by the model, existing security protection measures mainly fall into two categories: one is to perform security alignment fine-tuning during the model training phase, and the other is to add software review logic such as input filters, output text classifiers, and system prompt word constraints to the model inference chain. The above methods rely solely on semantic rules and the model's own recognition capabilities to make security judgments. They have a certain interception effect on straightforward and brief illegal requests, but their reliability drops significantly in scenarios involving leading and complex contexts such as role-playing, academic peer review, error correction discussions, and multi-turn historical dialogues.

[0003] Current mainstream inference architectures all adopt a single-track design where generation and review share a common runtime environment. This design suffers from a coupling defect in underlying computer hardware resources: the task generation branch and the security review branch share the same GPU memory, session memory, and KV-cache tensor cache. Adversarial prefixes, role packaging, and multi-turn inducement history carried by the user's original input are fully loaded into the GPU memory cache tensor, continuously polluting the model's attention representation. When the same model needs to both complete text generation and perform self-security judgments within the polluted context cache, it is highly susceptible to being swayed by the inducement context, resulting in security judgment biases that fail to recognize tactical content compliance, strategic framework, and sensitive policies.

[0004] Existing optimization approaches primarily focus on further strengthening model security alignment. However, simply raising the rejection threshold significantly increases the probability of falsely blocking legitimate requests such as academic discussions, vulnerability maintenance, and enterprise audits, failing to balance security and business availability. Furthermore, existing solutions do not decouple the generation and review processes at the inference hardware resource level, resulting in several hardware performance shortcomings: After the complete text is generated, it is uniformly sent to the review branch, requiring the entire long text to be cached in memory in streaming dialogue scenarios, leading to excessively high peak server memory usage; generation and review share GPU computing power, causing inference tasks to compete for hardware resources time-sharing, resulting in decreased overall throughput; user's original input and the entire dialogue history are all synchronized to the review branch, requiring the review side to process a large number of redundant tokens, leading to high cross-GPU transmission overhead and increased inference latency.

[0005] In summary, in existing single-track architecture large language model inference systems, the model generation branch and the security review branch share the same GPU memory, session memory, and KV-cache tensor cache. User input carrying suggestive context, multi-turn dialogue history, and adversarial packaging content pollutes the GPU memory cache tensor, causing a shift in the model's attention representation. This results in the same model performing self-security checks under induced contexts, making it highly susceptible to missed detections and failures in strategic risk identification. Furthermore, existing solutions rely solely on software logic optimizations such as model alignment and front-end text filtering, without addressing the isolation and modification at the inference hardware resource level. After generating complete text, it is uniformly sent to the review branch. In streaming inference scenarios, all long texts need to be cached in memory, leading to excessively high peak memory usage, GPU computing power contention, increased inference latency due to the transmission of a large number of redundant tokens across branches, and decreased cluster throughput—all deficiencies in underlying computer hardware resources. Simply increasing the model security alignment threshold significantly increases the probability of falsely intercepting legitimate business requests, failing to simultaneously balance inference hardware performance, content security identification accuracy, and business availability. Summary of the Invention

[0006] This invention addresses the shortcomings of existing large-model security protection schemes under single-track inference architectures, which rely solely on textual semantic rules and model fine-tuning for risk identification. These schemes fail to address underlying computer hardware limitations such as KV-cache cross-contamination, high memory resource consumption, and significant computational overhead. Consequently, they struggle to simultaneously achieve high-risk interception capabilities, low false positive rates, and stable inference hardware performance in complex, induced context scenarios. Therefore, this invention proposes a large-language model defense method and system based on context stripping and dual-track isolation. The invention achieves these technical problems through the following technical solutions: Solution 1: This invention proposes a large language model defense method based on context stripping and dual-track isolation, executed by a large language model inference server equipped with at least one GPU, independent video memory partition, and streaming sharded memory cache. The method includes the following steps: Step 1: Receive the user's original input token stream, assign a unique request identifier, and write it to the process memory buffer; Step 2: The user's original input token stream is distributed to the target generation track process running in the first independent GPU memory partition through the hardware routing and distribution module. The target generation track process loads the session KV-cache tensor into the memory partition and calls the target large language model to perform inference to generate the initial generated token text. Step 3: Write the initially generated token text into an independent and isolated streaming sharded memory cache module, and prohibit outputting it to the user terminal hardware before receiving a release instruction; Step 4: The context stripping hardware processing module reads the initially generated token text from the sharded memory cache; identifies the memory data blocks storing misleading context and performs stripping, retaining only the memory data modules storing core semantic content; Step 5: Route the memory data block containing only core semantic content to the independent review track process running in the second independent GPU memory partition; the second independent GPU memory partition is physically isolated from the first GPU memory partition and has no memory tensor communication channel. The independent review track process does not copy or read the KV-cache tensor, session memory and user raw input token stream of the target generation track. Step 6: The independent review track process performs security reasoning based on the stripped core semantic memory data block and outputs a security determination result; Step 7: Based on the security determination result, perform output control, memory clearing, or digest rewriting operations on the initially generated token text in the sharded memory cache.

[0007] Furthermore, a preferred embodiment is provided, in step 4, the context stripping hardware processing module reads the initially generated token text in the sharded memory cache; identifies the memory data blocks storing inductive context and performs stripping, retaining only the memory data modules storing core semantic content. The method is as follows: Hardware-accelerated text normalization is performed on the initially generated token text, and invalid bytes occupied by control characters and abnormal encodings in memory are cleared and redundant memory is released. The normalized token text is segmented based on the hardware semantic segmentation operator, and the memory offset address corresponding to each text segment is marked. The marked memory data modules are divided into packaging fragment memory submodules, restatement fragment memory submodules, core fragment memory submodules, and conclusion fragment memory submodules; Perform memory deletion or digest compression on the packaging fragment memory submodule and the restatement fragment memory submodule; merge the core fragment memory submodule and the conclusion fragment memory submodule to generate an inspection input tensor containing only the core semantic content.

[0008] Furthermore, a preferred embodiment is provided in which the independent review track process and the target generation track process described in step 5 are hardware isolated from at least one of the following: GPU memory partition, model instance, running process, container, virtual machine, inference server, system prompt word, session state, memory cache or KV-cache tensor.

[0009] Furthermore, a preferred embodiment is provided, wherein the independent review track includes at least one of a rule reviewer, a classification model, and a review big model; and the security determination result includes at least one of a security label, risk category, risk level, risk score, determination confidence level, or rule hit identifier.

[0010] Furthermore, a preferred embodiment is provided, wherein the method for performing output control, memory clearing, or digest rewriting operations on the initially generated token text in the sharded memory cache based on the security determination result in step 7 is as follows: When the security determination result meets the release conditions, the complete initially generated token text in the fragmented memory cache will be sent to the user terminal. When the security determination result meets the rewriting condition, the initially generated token text is rewritten securely in memory and then output. When the security determination result meets the interception conditions, the fragment memory cache is cleared directly and the corresponding GPU video memory is released, and a security prompt is returned to the user. When the security determination result meets the review conditions, the core semantic review input tensor is written to the disk audit queue, triggering the enhanced review process or the manual review process.

[0011] Furthermore, a preferred embodiment is provided, wherein when the target generation track process performs streaming token generation, the context stripping hardware processing module strips incremental tokens in the segmented memory cache into segments according to sentences, paragraphs, code blocks, semantic fragments, or fixed token windows; the independent review track process asynchronously performs security checks on the core semantic tensors after segment stripping; if any segment check meets the interception conditions, the subsequent token streaming transmission is immediately stopped and the streaming segmented memory cache is cleared.

[0012] Furthermore, a preferred embodiment is provided, wherein the method further includes an audit logging step: only the compressed core semantic digest is written to disk storage, and the user's original input memory data is hash-desensitized and then overwritten and erased.

[0013] Option 2: A large language model defense system with context stripping and dual-track isolation, the system comprising: The request access module is used to receive the user's original input token stream, assign an independent request identifier, and write it to the process memory buffer; The target generation track module is used to distribute the user's original input token stream to the target generation track process running in the first independent GPU memory partition through the hardware routing and distribution module. The target generation track process loads the session KV-cache tensor into the partition memory and calls the target large language model to perform inference to generate the initial generated token text. The output temporary storage module is used to write the initially generated token text into an independent and isolated streaming sharded memory cache module, and is prohibited from outputting to the user terminal hardware before receiving a release instruction; The context stripper module is used by the context stripping hardware processing module to read the initially generated token text in the sharded memory cache; identify the memory data blocks storing the inducement context and perform stripping, retaining only the memory data modules storing the core semantic content; The independent review track module is used to route memory data blocks containing only core semantic content to an independent review track process running on a second independent GPU memory partition. The second independent GPU memory partition is physically isolated from the first GPU memory partition and has no memory tensor communication channel. The independent review track process does not copy or read the KV-cache tensor, session memory, or user's original input token stream from the target generation track. The isolation determination module is used by the independent review track process to perform security inference based on the stripped core semantic memory data block and output the security determination result; The output control module is used to perform output control, memory clearing, or digest rewriting operations on the initially generated token text in the sharded memory cache according to the security determination result.

[0014] Option 3: A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in Option 1.

[0015] Option 4: A computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the method described in Option 1.

[0016] The advantages of this invention are: The large language model defense method with context stripping and dual-track isolation described in this invention adopts a dual-track independent GPU memory partition physical isolation architecture. At the hardware level, it cuts off the data communication channels between the KV-cache tensor and session memory between the target generation track and the independent review track. This avoids the pollution of the model's attention representation by the inducing context cache tensor across branches from the root, and solves the defect of the model being interfered with by the inducing context under the single-track architecture, which leads to the omission of strategic sensitive content. It also greatly improves the stability of security judgment in complex adversarial and multi-round inducing scenarios.

[0017] The large language model defense method with context stripping and dual-track isolation described in this invention classifies and trims the token data blocks in memory through a hardware-accelerated context stripping unit, removes redundant induced context tensors such as role packaging and paraphrasing, and only transmits the core semantic tensor to an independent review track. This significantly reduces the amount of token data transmitted across GPUs, reduces GPU memory transmission overhead, shortens the inference latency of the review branch, and improves GPU computing power utilization and the concurrent throughput of the AI ​​inference cluster.

[0018] The large language model defense method of context stripping and dual-track isolation described in this invention adopts streaming segmented independent physical memory caching to temporarily store generated content in segments and asynchronously conduct segmented review. It does not require caching complete long texts in memory, effectively reducing the peak memory usage of the inference server, avoiding hardware operation failures such as memory overflow, inference lag, and service interruption in ultra-long multi-turn dialogue scenarios, and improving the operational stability of streaming dialogue services.

[0019] The large language model defense method with context stripping and dual-track isolation described in this invention relies on inference hardware resource scheduling to achieve a security fallback. It does not require large-scale security alignment fine-tuning of the target large model, and will not cause false blocking of normal business requests such as legitimate academic discussions, vulnerability maintenance, and enterprise compliance audits due to raising the rejection threshold. While ensuring high-risk interception capabilities, it fully preserves the original generation capabilities of the model, balancing the accuracy of content security identification with business availability.

[0020] The large language model defense method of context stripping and dual-track isolation described in this invention sets up a context stripper before review to separate the packaging context and substantive output content in the initially generated text, so that the review model can judge based on the core semantic content, thereby improving the stability and objectivity of the review.

[0021] The large language model defense method with context stripping and dual-track isolation described in this invention is based on the underlying inference scheduling logic of GPU memory partitioning isolation, memory sharding caching, and hardware tensor stripping. It does not require modification of the basic network structure of the large model and can be seamlessly adapted to various streaming large model applications such as cloud inference servers, code assistants, and intelligent customer service. It has strong engineering deployment adaptability and low modification cost.

[0022] This invention supports asynchronous review and streaming output review, and can be well adapted to existing large language model APIs, enterprise private model services and intelligent application gateways. This invention retains anonymized audit records, which facilitates enterprises to conduct compliance tracking, risk statistics and subsequent strategy optimization.

[0023] This implementation method is also applicable to application areas that require low-latency responses, such as chatbots, code assistants, and online customer service. Attached Figure Description

[0024] Figure 1This is a flowchart illustrating the large language model defense method of context stripping and dual-track isolation described in Implementation Method 1.

[0025] Figure 2 This is a schematic diagram of the dual-track isolation security defense system described in Implementation Method 1.

[0026] Figure 3 This is a schematic diagram of the context stripper described in Embodiment 1.

[0027] Figure 4 This is a data flow diagram of the independent review track and isolation determination module as described in Implementation Method 1.

[0028] Figure 5 This is a schematic diagram of the asynchronous review process in the streaming output scenario described in Implementation Method 1. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0030] Implementation Method 1, see Figures 1 to 5 This embodiment describes a large language model defense method based on context stripping and dual-track isolation. The method specifically includes the following steps: S101. After receiving user input, the system passes the input to the target generation track for processing.

[0031] S102. The target generation track calls the target large language model to generate the initial generated text and writes the initial generated text into the output temporary storage module.

[0032] S103. The output temporary storage module does not send the initially generated text directly to the user before obtaining permission from the isolation judgment module.

[0033] S104. After reading the initial generated text, the context stripper first performs text normalization processing, and then segments it according to semantic boundaries and structural boundaries.

[0034] S105-S107: For identified role packaging, error correction packaging, prompt word restatements, and narrative guidance, the system marks them as non-core fragments; for substantive explanations, process descriptions, code modules, sensitive entities, and conclusive content, the system marks them as core fragments. Next, the system constructs review input objects based on the core fragments and sends these review input objects to an independent review track.

[0035] S108. The independent review track invokes the independent review model for security determination. Since the independent review track does not receive the user's original prompts and does not share the target generation track's KV-cache, its determination process is not directly affected by the target generation track's original inducement context. The isolation determination module performs allow, rewrite, or block based on the review results.

[0036] S109. All handling results are ultimately generated into audit records, forming a complete closed loop for text security control.

[0037] See Figure 3 This implementation describes a context stripper that combines rules and models. Rules are used to identify content with obvious formatting, such as code blocks, quotes, lists, control characters, invisible characters, and fixed templates. Models are used to determine whether a text fragment is wrapper, restatement, core explanation, or conclusion. For fragments that are difficult to determine, the system includes them in the core semantic content to avoid omissions.

[0038] The context stripper also outputs fragment metadata. Fragment metadata includes the fragment's start and end positions in the initial generated text, fragment type, extraction confidence level, whether it contains code blocks, whether it contains sensitive entities, and candidate risk categories. Independent review tracks utilize this metadata to improve review efficiency.

[0039] See Figure 4 This implementation method is described below. This implementation method further explains the independent review track, which employs a multi-level review structure. The first-level rule reviewer is used to quickly identify obvious risks; the second-level classification model is used to output risk categories and risk scores; the third-level review big model is used to handle semantically complex or unclearly defined content. If a high-confidence interception decision has already been made in the previous level, subsequent levels will not be executed to reduce inference costs.

[0040] The model instances, cache space, and system prompts of the independent review track are all separated from the target generation track. The system achieves this separation through container isolation, service isolation, video memory session isolation, or access permission isolation.

[0041] See Figure 4 This implementation method is described below. This implementation method further explains the isolation decision module, which maintains a policy table. The policy table records security labels, risk categories, risk score ranges, confidence thresholds, and corresponding output actions. Output actions include direct approval, deleting risky segments and then outputting, generating a security summary, returning a security prompt, transferring to manual review, and writing to a high-risk audit queue.

[0042] When multiple reviewers produce inconsistent results, the isolation judgment module employs either the highest risk priority rule or a weighted voting rule. For scenarios with high enterprise compliance requirements, the system is set to prioritize high risk; for general knowledge-based question-and-answer scenarios, the system is set to threshold fusion to reduce false blocking.

[0043] See Figure 5 This implementation method is described below. It is applicable to streaming generation. The target generation track continuously writes fragments to the output buffer module during the generation process. The context stripper performs incremental processing on the fragments. An independent review track asynchronously reviews the current fragment and its preceding and following adjacent fragments. If the review result is safe, the output control module allows the fragment to be output; if the review result is high-risk, the output control module stops streaming and outputs a safety warning.

[0044] In summary, the context stripper described in this embodiment is used to extract the target model output from the original inducing context. This component does not merely perform sensitive word removal, but separates the packaging context from the substantive output content, allowing the review track to receive the output content itself, rather than the packaged task narrative. In one implementation, the context stripper includes a text normalization unit, a fragment segmentation unit, a packaging identification unit, a core content extraction unit, and a review object generation unit. The text normalization unit is used to eliminate encoding and formatting noise; the fragment segmentation unit is used to segment the initially generated text according to paragraphs, sentences, Markdown structures, code block boundaries, or semantic boundaries; the packaging identification unit is used to identify role settings, error correction packaging, academic review packaging, adversarial prefixes, prompt word restates, and disclaimer-style packaging; the core content extraction unit is used to extract the substantive output content; and the review object generation unit is used to form the review input object. The review input object includes fields such as core semantic content, fragment type, source location, extraction confidence, risk candidate labels, whether it contains code blocks, and whether it contains sensitive entities.

[0045] The independent review track must meet isolation requirements. The independent review track does not read the user's original prompts, does not reuse the message history of the target generation track, does not inherit the target generation track's KV-cache, and does not use the attention state formed by the target generation track in this request. The independent review track can adopt a review model heterogeneous with the target generation model. This heterogeneity manifests in different model architectures, parameter sources, training objectives, system prompts, security policies, or deployment environments. The independent review track includes at least one of a rule reviewer, a classification model, and a large review model. The results of multiple reviewers are fused through voting, weighted averaging, priority rules, or policy tables.

[0046] The isolation decision module determines the final action based on the results of the independent review track output. The isolation decision module can employ a threshold strategy, such as setting a first risk threshold and a second risk threshold. When the risk score is below the first risk threshold, the request is allowed; when the risk score is between the two thresholds, rewriting, summarizing, or enhanced review is performed; when the risk score is above the second risk threshold and the confidence level meets the requirements, the request is blocked. Risk levels can include Level 0, Level 1, and Level 2, where Level 0 is a secure output or compliance rejection, allowing the output; Level 1 is a strategic-level risk output, requiring blocking, summarizing, or review; and Level 2 is a tactical-level violation output, requiring mandatory blocking. The above levels are only an example; in actual deployments, more categories are expanded according to industry compliance requirements.

[0047] In streaming generation scenarios, the target generation track writes the generated content into the output temporary storage module in segments. The context stripper strips the current segment or combinations of adjacent segments, while independent review tracks review them in parallel. If the current segment is determined to be safe, the output control module allows continued output; if any segment is determined to be high-risk, the output control module immediately stops streaming output and returns a safety warning. This approach avoids discovering risks only after the entire long text has been generated, while maintaining a low user wait time.

[0048] In enterprise deployments, the target generation track can reside in the business model service cluster, while the independent review track can reside in the security review service cluster. Both transmit review input objects via an internal gateway, message queue, or remote procedure call interface. The security review service cluster only receives the core semantic content after context stripping and does not have permission to read the user's original session, business retrieval content, or tool call results. Even if the target generation track is affected by complex context, the independent review track still makes judgments with relatively clean review input, reducing the propagation of context pollution from a system architecture perspective.

[0049] Implementation Method Two: This implementation method proposes a large language model defense system with context stripping and dual-track isolation. The system includes: The request access module is used to receive the user's original input token stream, assign an independent request identifier, and write it to the process memory buffer; The target generation track module is used to distribute the user's original input token stream to the target generation track process running in the first independent GPU memory partition through the hardware routing and distribution module. The target generation track process loads the session KV-cache tensor into the memory partition and calls the target large language model to perform inference to generate the initial generated token text. The output temporary storage module is used to write the initially generated token text into an independent and isolated streaming sharded memory cache module, and is prohibited from outputting to the user terminal hardware before receiving a release instruction; The context stripper module is used by the context stripping hardware processing module to read the initially generated token text in the sharded memory cache; identify the memory data blocks storing the inducement context and perform stripping, retaining only the memory data modules storing the core semantic content; The independent review track module is used to route memory data blocks containing only core semantic content to an independent review track process running on a second independent GPU memory partition. The second independent GPU memory partition is physically isolated from the first GPU memory partition and has no memory tensor communication channel. The independent review track process does not copy or read the KV-cache tensor, session memory, or user's original input token stream from the target generation track. The isolation determination module is used by the independent review track process to perform security inference based on the stripped core semantic memory data block and output the security determination result; The output control module is used to perform output control, memory clearing, or digest rewriting operations on the initially generated token text in the sharded memory cache according to the security determination result.

[0050] The request access module is used to receive the user's original input and generate a request identifier.

[0051] The target generation track is used to invoke the target large language model to generate initial generated text. The target generation track can utilize existing system prompts, search enhancements, tool call results, and session history.

[0052] The output buffer module is positioned after the target generation track and is used to buffer the initially generated text before it is sent to the user. For streaming output, the output buffer module can buffer segments by sentence, paragraph, code block, or fixed token window.

[0053] The context stripper is used to clean, segment, classify, and extract the initial generated text to generate review input objects. Review input objects include core semantic content and necessary metadata such as fragment type, source location, and risk candidate tags, but do not include the user's original prompt words.

[0054] Independent review tracks are used to perform security checks on the review input objects. Independent review tracks do not share the target generation track's session history, system prompts, and key-value cache. Independent review tracks can be deployed as standalone processes, standalone containers, standalone model services, standalone virtual machines, standalone inference servers, or trusted execution environments.

[0055] The isolation decision module is used to determine the output action based on the security label, risk category, risk score, and confidence level output by the independent review track.

[0056] The output control module is used to perform actions such as release, rewriting, summarization, interception, security prompts, or manual review.

[0057] The audit log module is used to record information about the de-identified review process for subsequent compliance audits, model risk statistics, and strategy optimization.

[0058] Those skilled in the art will understand that the above description is merely a preferred embodiment of the present invention, and the features described in the various embodiments and / or technical solutions of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. This is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0059] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended technical solutions are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the present invention. Clearly, those skilled in the art can make various modifications and variations to the present invention without departing from its spirit and scope. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalents, the present invention also intends to include these modifications and variations.

Claims

1. A method for defending large language models with context stripping and dual-track isolation, executed by a large language model inference server equipped with at least one GPU, independent video memory partition, and streaming sharded memory cache, characterized in that, The method includes the following steps: Step 1: Receive the user's original input token stream, assign a unique request identifier, and write it to the process memory buffer; Step 2: The user's original input token stream is distributed to the target generation track process running in the first independent GPU memory partition through the hardware routing and distribution module. The target generation track process loads the session KV-cache tensor into the partition memory and calls the target large language model to perform inference to generate the initial generated token text. Step 3: Write the initially generated token text into an independent and isolated streaming sharded memory cache module, and prohibit outputting it to the user terminal hardware before receiving a release instruction; Step 4: The context stripping hardware processing module reads the initially generated token text from the sharded memory cache; identifies the memory data blocks storing misleading context and performs stripping, retaining only the memory data modules storing core semantic content; Step 5: Route the memory data block containing only core semantic content to the independent review track process running in the second independent GPU memory partition; the second independent GPU memory partition is physically isolated from the first GPU memory partition and has no memory tensor communication channel. The independent review track process does not copy or read the KV-cache tensor, session memory and user raw input token stream of the target generation track. Step 6: The independent review track process performs security reasoning based on the stripped core semantic memory data block and outputs a security determination result; Step 7: Based on the security determination result, perform output control, memory clearing, or digest rewriting operations on the initially generated token text in the sharded memory cache.

2. The large language model defense method based on context stripping and dual-track isolation according to claim 1, characterized in that, In step 4, the context stripping hardware processing module reads the initially generated token text from the sharded memory cache; identifies the memory data blocks storing suggestive context and performs stripping, retaining only the memory data modules storing core semantic content. Hardware-accelerated text normalization is performed on the initially generated token text, and invalid bytes occupied by control characters and abnormal encodings in memory are cleared and redundant memory is released. The normalized token text is segmented based on the hardware semantic segmentation operator, and the memory offset address corresponding to each text segment is marked. The marked memory data modules are divided into packaging fragment memory submodules, restatement fragment memory submodules, core fragment memory submodules, and conclusion fragment memory submodules; Perform memory deletion or digest compression on the packaging fragment memory submodule and the restatement fragment memory submodule; merge the core fragment memory submodule and the conclusion fragment memory submodule to generate an inspection input tensor containing only the core semantic content.

3. The large language model defense method based on context stripping and dual-track isolation according to claim 1, characterized in that, The independent review track process and the target generation track process mentioned in step 5 are hardware isolated from each other in at least one of the following: GPU memory partition, model instance, running process, container, virtual machine, inference server, system prompt word, session state, memory cache or KV-cache tensor.

4. The large language model defense method based on context stripping and dual-track isolation according to claim 3, characterized in that, The independent review track includes at least one of a rule reviewer, a classification model, and a large review model; the security determination result includes at least one of a security label, risk category, risk level, risk score, determination confidence level, or rule hit identifier.

5. The large language model defense method with context stripping and dual-track isolation according to claim 1, characterized in that, The method for performing output control, memory clearing, or digest rewriting operations on the initially generated token text in the sharded memory cache based on the security determination result in step 7 is as follows: When the security determination result meets the release conditions, the complete initially generated token text in the fragmented memory cache will be sent to the user terminal. When the security determination result meets the rewriting condition, the initially generated token text is rewritten securely in memory and then output. When the security determination result meets the interception conditions, the fragment memory cache is cleared directly and the corresponding GPU video memory is released, and a security prompt is returned to the user. When the security determination result meets the review conditions, the core semantic review input tensor is written to the disk audit queue, triggering the enhanced review process or the manual review process.

6. The large language model defense method based on context stripping and dual-track isolation according to claim 1, characterized in that, When the target generation track process performs streaming token generation, the context stripping hardware processing module segments the incremental tokens in the fragmented memory cache according to sentences, paragraphs, code blocks, semantic fragments, or fixed token windows; the independent review track process asynchronously performs security checks on the core semantic tensors after segmentation; if any segment check meets the interception conditions, the subsequent token streaming transmission is immediately stopped and the streaming fragmented memory cache is cleared.

7. The large language model defense method based on context stripping and dual-track isolation according to claim 1, characterized in that, The method also includes an audit logging step: only the compressed core semantic digest is written to disk storage, and the user's original input memory data is hash-desensitized and then overwritten and erased.

8. A large language model defense system with context stripping and dual-track isolation, characterized in that, The system includes: The request access module is used to receive the user's original input token stream, assign an independent request identifier, and write it to the process memory buffer; The target generation track module is used to distribute the user's original input token stream to the target generation track process running in the first independent GPU memory partition through the hardware routing and distribution module. The target generation track process loads the session KV-cache tensor into the memory partition and calls the target large language model to perform inference to generate the initial generated token text. The output temporary storage module is used to write the initially generated token text into an independent and isolated streaming sharded memory cache module, and is prohibited from outputting to the user terminal hardware before receiving a release instruction; The context stripper module is used by the context stripping hardware processing module to read the initially generated token text in the sharded memory cache; identify the memory data blocks storing the inducement context and perform stripping, retaining only the memory data modules storing the core semantic content; The independent review track module is used to route memory data blocks containing only core semantic content to an independent review track process running on a second independent GPU memory partition. The second independent GPU memory partition is physically isolated from the first GPU memory partition and has no memory tensor communication channel. The independent review track process does not copy or read the KV-cache tensor, session memory, or user's original input token stream from the target generation track. The isolation determination module is used by the independent review track process to perform security inference based on the stripped core semantic memory data block and output the security determination result; The output control module is used to perform output control, memory clearing, or digest rewriting operations on the initially generated token text in the segmented memory cache according to the security determination result.

9. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.

10. A computer device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Voice emotion recognition method and device based on context information, equipment and medium

    CN120636474A

  • Text large model defense method and system for multi-agent collaborative security filtering

    CN122241762A