Large model safety shield system and safety protection method

By building a large-model security shield system with multi-level protection architecture and multi-dimensional alignment expert model, the security protection problem of large-language models in the financial field is solved, real-time identification and repair of deceptive attacks is achieved, security and compliance are improved, and it is suitable for high-security industries such as finance, medical care, and law.

CN120579210APending Publication Date: 2025-09-02APPROVAL TECH CO LTD

Patent Information

Application Number
CN202510751822.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing large language model has problems such as incomplete security coverage, limited language support, delayed audit, lack of targetedness, insufficient static defense mechanisms and confrontational deception methods in the security protection in the financial field, especially in the face of deceptive prompt words and breakthrough attacks.

Method used

Build a large-model security shield system, adopting a multi-level protection architecture of input layer security protection, business scenario processing layer security alignment and output layer security audit, combining multi-dimensional alignment expert model and cross-wheel dialogue security detection to achieve real-time identification and repair of illegal content, and support multilingual and streaming output.

Benefits of technology

The full-process security monitoring of the financial field is realized, the performance and efficiency of the model are improved, and the deceptive attacks in multiple rounds of dialogue are effectively prevented, ensuring the security and compliance of the output content, and the user experience is not affected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579210A_ABST
    Figure CN120579210A_ABST
Patent Text Reader

Abstract

The invention relates to a large model safety shield system and a safety protection method. The large model safety shield system comprises three layers of input layer safety protection, business scene processing layer safety alignment and output layer safety auditing, and the safety of the whole processing flow is ensured through a multi-layer protection framework. According to the large model safety shield system, the cue word shield module, the multi-dimensional alignment expert model module and the output supervision module are utilized, and real-time safety detection and multi-dimensional alignment processing of user input and real-time monitoring and auditing of output content are achieved. The method has the advantages of high-efficiency security detection, high-accuracy interception, optimized user experience and good expandability and adaptivity, and the security and reliability of application of a large model in the fields of finance and the like are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence security protection technology, and particularly relates to a security shield system and security protection method for large (language) models. This technology can be applied to industries such as finance, medical care, and law that require strict content security control, and is especially suitable for large model application scenarios in the financial technology industry. Background Art

[0002] With the rapid development of artificial intelligence (AI) in recent years, the application of large language models in the financial sector has become a key driver of industry transformation. The financial industry, with its rich data resources and diverse application scenarios, offers ample room for the implementation of large language model technology. Its applications span a wide range of sectors, including mobile banking, investment analysis, financial analysis, intelligent customer service, credit assessment, fraud detection, and more. It can process massive amounts of text, interpret financial reports, and improve service efficiency.

[0003] However, the widespread use of large language models in the financial sector has also brought about increasingly prominent security issues. Existing security protection mechanisms mainly include the following solutions: (1) Keyword-based filtering mechanism: User input is filtered through a keyword list or rule engine to block prompt words containing sensitive words. This method is simple and direct, but it is easy to bypass and cannot understand the contextual semantics. Its defense capabilities are seriously insufficient when facing deceptive prompt words and carefully designed breakthrough attacks.

[0004] (2) Filtering mechanism based on prompt word optimization: This method accurately locates and filters inappropriate content by optimizing prompt words. This method is highly flexible and can adjust prompt word content according to different scenarios, making it easier to understand and maintain. However, it generally suffers from difficulties in command alignment and balanced rule formulation, making it prone to omissions or misjudgments.

[0005] (3) Filtering mechanism based on the general security audit model: The general security audit model mainly intercepts crimes, bias discrimination, command attacks, and moral and ethical issues. However, this mechanism lacks support for the security mechanism requirements of the financial sector and is difficult to meet the special security needs of the financial sector.

[0006] While the above solution is effective in general scenarios, in specific areas such as finance, the existing technology has the following obvious shortcomings: (1) Incomplete security coverage: General security filtering mechanisms cannot cover professional fraud patterns and attack methods in specific fields such as finance, resulting in security blind spots.

[0007] (2) Limited language support: Most security mechanisms only support Chinese and English, and lack support for minority languages ​​such as Arabic, which cannot meet the needs of global business.

[0008] (3) Audit delay problem: The traditional audit method is carried out after the model output is completed, which not only increases the response time, but also cannot adapt to the streaming output mode, affecting the user experience.

[0009] (4) Lack of specificity: Existing model alignments mainly focus on general values ​​and lack a deep understanding of and alignment with the ethical and regulatory requirements of specific industries (such as the financial sector).

[0010] (5) Static defense mechanism: Most security mechanisms are statically designed and cannot cope with ever-changing attack methods, especially deception techniques and breakthrough attacks against large models. The security protection effect will decrease over time.

[0011] (6) Insufficient means to combat deception: Existing technologies have limited ability to identify carefully designed deceptive prompt words, making it difficult to prevent users from breaking through the security limitations of large models through roundabout, split or multi-round dialogue methods.

[0012] In addition, although some security protection technologies involving large language models have been developed, for example, the Chinese patent "A System and Method for Identifying and Filtering Inappropriate Content Based on a Large Language Model" (CN118484539A) discloses an inappropriate content identification and filtering system based on a large language model, including a text collection and preprocessing module, a feature extraction and analysis module, an inappropriate content identification module, a filtering decision module, and a user interaction module, these methods still have certain drawbacks in their application. For example, most of them use a single centralized processing flow, which makes it difficult to fully cover all key security nodes and results in large protection blind spots. Their functions mainly focus on relatively straightforward processing measures such as warnings, isolation, and deletion, lacking optimization methods for repairable content. Their application scenarios are relatively limited, mostly targeting scenarios such as social media and online forums, and lack deep adaptation to specific industries. This is especially true for industries such as finance, healthcare, and law that have extremely high security and reliability requirements, and lack customized security protection solutions. Summary of the Invention

[0013] To address these issues, this application proposes a novel large-scale model security shield system and security protection method. This system utilizes a prompt word shield, multi-dimensional alignment, and output supervision mechanisms to build a comprehensive security protection system. This system can identify illegal content in real time and invoke a fallback strategy to ensure input security. It integrates expert models to align semantics and values, ensuring the professionalism and compliance of generated content. Furthermore, it implements monitoring during the output phase to intercept inappropriate information, establish a solid security baseline, adapt to multiple scenarios, and ensure the robust operation of AI.

[0014] In this invention, in particular, in response to the above-mentioned defects in existing large-scale model security protection technologies, 1) a multi-level protection architecture with three levels of input layer security protection, business scenario processing layer security alignment and output layer security audit is adopted to conduct comprehensive security monitoring of the entire processing flow of the large model, effectively reducing protection blind spots; 2) a multi-dimensional alignment expert model is designed, which is deeply integrated with business scenarios in specific fields such as finance, and can integrate business scenario expertise with adversarial protection technology. Through supervised fine-tuning and reinforcement learning, the model instruction alignment capability is improved to ensure that the generated content is both in line with business needs and compliant with value requirements. 3) A cross-round dialogue security detection mechanism is specially introduced, which can realize the recognition of distributed command injection and progressive deception in multi-round dialogues, effectively preventing users from breaking through the security protection of large models through roundabout means, and has the advantage of cross-round dialogue security detection; 4) A sensitive word repair mechanism is adopted, which can intelligently repair content marked as unsafe but repairable, such as using synonym replacement technology, reconstructing sentences, or inserting security qualifiers, etc., to ensure the security of the content while maintaining the availability and coherence of the output content as much as possible. Through the above design, the present invention successfully builds a security protection system for specific fields, which is particularly suitable for application in finance, medical care, law and other industries with extremely high security and compliance requirements.

[0015] Explanation of terms: Large model: refers to a pre-trained artificial intelligence model obtained through training with large-scale data and a huge number of parameters, which has strong generalization and multi-task processing capabilities.

[0016] Expert model: refers to a model that, based on a pre-trained large language model, introduces a trainable low-rank matrix at a specific layer and only fine-tunes a small number of parameters to achieve domain adaptation. Its purpose is to improve the performance, security, controllability, etc. of the model in specific tasks or scenarios.

[0017] Prompt Word Shield: A security mechanism used to detect and filter malicious prompt words.

[0018] Streaming output: refers to the real-time continuous output of model-generated content, rather than waiting for the complete content to be generated and then displayed all at once.

[0019] The present invention aims to solve the following core technical problems: 1. Build a security protection system for specific fields: How to build a security protection system for specific fields such as finance that can effectively identify and defend against special attack patterns in that field, especially deception and breakthrough attacks against large models.

[0020] 2. Implement multi-language security protection support: How to implement security protection support for multiple languages ​​(including minority languages).

[0021] 3. Real-time security audit of streaming output: How to achieve real-time security audit of streaming output while ensuring a smooth user experience.

[0022] 4. Design a multi-dimensionally aligned expert model: How to design a multi-dimensionally aligned expert model that meets the regulatory requirements and ethical standards of specific industries.

[0023] 5. Preventing behaviors that bypass security restrictions: How to effectively prevent users from bypassing security restrictions through multiple rounds of dialogue, hidden instructions, etc.

[0024] In order to achieve the above objectives, this application provides the following technical solutions: A first aspect of the present application provides a large-scale model safety shield system, the safety shield system comprising: The input layer security protection module is used to receive user input and perform preliminary security checks. The input layer security protection module includes a prompt word shield module to detect user input in real time and identify and filter potentially harmful or illegal content; A business scenario processing layer security alignment module is used to perform multi-dimensional alignment processing on inputs that pass preliminary security checks during the business scenario processing process. The business scenario processing layer security alignment module includes a multi-dimensional alignment expert model module, which is used to integrate business scenario expertise with adversarial protection technology, and improve the model instruction alignment capability through supervised fine-tuning and reinforcement learning to ensure that the generated content is compliant and meets value requirements; The output layer security audit module is used to monitor and audit the content output by the business scenario processing layer in real time. The output layer security audit module includes an output supervision module (streaming audit module). The output supervision module includes a lightweight real-time layer and a deep analysis layer. It is used to implement real-time monitoring and review during the large model generation stage, detect potential violations or harmful content in the output information, intercept and correct non-compliant outputs, and ensure that the final output content is safe and compliant.

[0025] Furthermore, in the large-scale model security shield system of the present application, the prompt word shield module includes a financial-specific security framework for building a dedicated private domain knowledge base for the financial field, including a general security audit set, a common financial fraud pattern library, a financial sensitive information identification rule set, and a cross-round dialogue attack pattern library; When the prompt word shield module discovers a risk, it immediately calls the preset safety strategy and private domain knowledge base to extract or generate safety words to ensure the safety of the downstream processing environment.

[0026] Furthermore, in the large model security shield system of the present application, the input layer security protection module also includes: A multilingual matching system module, designed to support multilingual matching mechanisms for global businesses, includes cross-lingual pattern recognition based on vector matching, specialized word segmentation for minority languages, data synthesis, and an expert review model that supports multiple languages. The matching algorithm module is used to effectively intercept harmful input at the user input layer, including a cross-turn dialogue attack pattern library and a prompt word attack pattern library. The matching algorithm module also includes financial field pattern detection functions, including fraud pattern detection, deception pattern detection, and compliance detection.

[0027] Furthermore, in the large model security shield system of the present application, the business scenario processing layer security alignment module also includes: The post-training expert model module is used to achieve domain adaptation by introducing a trainable low-rank matrix (Lora) at specific layers based on the pre-trained model. This module includes scene cue word design and optimization, manual annotation and synthesis of scene data, supervised fine-tuning, reinforcement learning based on human feedback, and expert model performance evaluation. The multi-dimensional alignment assessment module is used to evaluate the alignment capabilities of the model from multiple dimensions, including business instruction alignment, instruction attack, compliance alignment, risk alignment, and value alignment.

[0028] Furthermore, in the large model security shield system of this application, the working process of the post-training expert model module is as follows: Scenario data and safety-related data are used as input data sources, and prompt words are generated by combining the scenario data and safety-related data; Use the generated prompt words to train the base model and perform supervised fine-tuning on the base model to improve the performance of the model on specific tasks; Use human feedback to guide the model’s reinforcement learning process and further improve the model’s performance; The model that has undergone reinforcement learning is evaluated to determine whether it meets specific requirements. Models that meet the requirements are quantized to reduce the model's resource consumption and improve operating efficiency. Models that do not meet the requirements return to the supervised fine-tuning step for further model optimization. After the model quantization is completed, the training process ends.

[0029] Furthermore, in the large model security shield system of the present application, the output layer security audit module also includes: Dynamic adjustment algorithm module, used to adjust output in real time while maintaining a smooth experience; The sensitive word repair module is used to intelligently repair content marked as unsafe but repairable. The intelligent repair includes: replacing sensitive words using synonym replacement technology, reconstructing sentences that may have implicit unsafe meanings, or inserting safe qualifiers while maintaining semantic coherence.

[0030] A second aspect of the present application provides a large model security protection method, the method comprising the following steps: S1. The prompt word shield module receives user input and performs preliminary security checks. This includes real-time detection of user input, identifying and filtering potentially harmful or illegal content. If risks are detected, it immediately invokes pre-set safety strategies and private domain knowledge bases to extract or generate safety-critical phrases, ensuring a secure downstream processing environment. S2. The Multi-Dimensional Alignment Expert Model module performs multi-dimensional alignment on inputs that pass preliminary security checks during business scenario processing. This involves integrating business scenario expertise with adversarial protection technologies. It uses supervised fine-tuning and reinforcement learning to enhance the model's instruction alignment capabilities, ensuring that the generated content is compliant and meets value requirements. S3. The output supervision module monitors and reviews the output content of the business scenario processing layer in real time. This includes real-time monitoring and review during the large model generation phase, detecting potential violations or harmful content in the output information, intercepting and correcting non-compliant outputs, and ensuring the security and compliance of the final output content.

[0031] Furthermore, step S1 of the present method also includes: Build a dedicated private domain knowledge base for the financial sector, including a general security audit set, a common financial fraud pattern library, a set of financial sensitive information identification rules, and a cross-round dialogue attack pattern library; Implement a multilingual matching mechanism to support global business, including cross-language pattern recognition based on vector matching, specialized word segmentation for minority languages, data synthesis, and an expert review model that supports multiple languages. A multi-level matching algorithm is used to effectively intercept harmful input at the user input layer, including a cross-turn dialogue attack pattern library and a prompt word attack pattern library. The multi-level matching algorithm also includes financial field pattern detection functions, including fraud pattern detection, deception pattern detection and compliance detection.

[0032] Furthermore, step S2 of the present method also includes: Based on the pre-trained model, domain adaptation is achieved by introducing a trainable low-rank matrix (Lora) at specific layers, including scene cue word design and optimization, manual annotation and synthesis of scene data, supervised fine-tuning, reinforcement learning based on human feedback, and expert model performance evaluation. Evaluate the model's alignment capabilities from multiple dimensions, including business directive alignment, directive attack, compliance alignment, risk alignment, and value alignment.

[0033] Furthermore, step S3 of the present method also includes: Implement real-time monitoring and review during the large model generation phase, including a lightweight real-time layer and a deep analysis layer; Dynamically adjust the algorithm to adjust the output in real time while maintaining a smooth experience; Intelligently repair content marked as unsafe but repairable, including: replacing sensitive words with synonym replacement technology, reconstructing sentences that may have implicit unsafe meanings, or inserting safe qualifiers while maintaining semantic coherence.

[0034] A third aspect of the present application provides an electronic device, comprising: a memory and a processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the aforementioned large model security protection method.

[0035] The fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the aforementioned large model security protection method are implemented.

[0036] In summary, the present invention includes the following technical innovations: (1) Financial Specific Security Framework: The first dedicated private domain knowledge base for the financial sector, which not only contains general security field data sets, but also covers data sets on security aspects of the financial sector.

[0037] (2) Multi-language security support: A cross-language security matching mechanism, including minority languages, has been implemented, filling a gap in the industry.

[0038] (3) Multi-level security audit: An innovative multi-level security audit mechanism is proposed, which takes into account the advantages of both semantic matching and model matching.

[0039] (4) Multi-dimensional expert alignment: For the first time, the multi-dimensional alignment of supervision, ethics, risk and culture is integrated into a unified system to ensure all-round compliance output.

[0040] (5) Streaming security audit technology: It solves the real-time security audit problem of streaming output, ensuring content security while maintaining a smooth user experience.

[0041] (6) Cross-round dialogue security detection: It realizes the recognition of distributed command injection and progressive deception in multi-round dialogues, effectively preventing users from breaking through the security protection of large models through circuitous means.

[0042] Compared with existing large-scale model security protection technologies, this invention has the following advantages: High efficiency: A multi-level matching algorithm achieves highly efficient security detection, minimizing latency while ensuring security.

[0043] High accuracy: The measured interception rate is close to 98%, far higher than the industry average.

[0044] User experience optimization: The streaming security audit mechanism ensures security without affecting user experience, solving industry problems.

[0045] Scalability: The system is modular in design and can be easily expanded to other verticals such as healthcare, law, etc.

[0046] Adaptability: Through continuous learning mechanisms, the system can automatically adapt to new attack patterns and maintain long-term effectiveness.

[0047] Other features and advantages of this application will be described in detail in the following description, or will be understood through implementation of the relevant technical solutions of this application. The objectives and other advantages of this application can be achieved through the technical features and technical means clearly indicated in the description, claims, and drawings, and obtained through the implementation of these technical contents. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings involved in the description of the embodiments. It should be noted that the drawings only illustrate some embodiments of the present application. For those skilled in the art, other relevant drawings can be derived from these drawings without engaging in creative work.

[0049] Figure 1 This is the overall design architecture and workflow diagram of the large-scale model safety shield system of the present invention.

[0050] Figure 2 This is a flowchart of the post-training of the expert model in the large model security shield system of the present invention.

[0051] Figure 3 This is an overall implementation flow chart of the large model safety protection method of the present invention.

[0052] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0054] In this document, the term "including" and any variations thereof (such as "including," "comprising," etc.) are open-ended expressions and should be understood as meaning "including but not limited to," meaning that the listed contents are not exhaustive and may include other contents not explicitly mentioned. The term "based on" should be understood as meaning "based at least in part on," meaning that the basis or condition referred to may not be the only factor and may also involve other relevant factors. The term "one embodiment" should be understood as meaning "at least one embodiment," meaning that the described embodiment is not the only possible implementation method and that other similar embodiments may exist.

[0055] In this application, the terms "a" and "a plurality" are used to modify related elements or features in an illustrative, non-restrictive manner. Unless the context clearly indicates otherwise, "a" should be understood as meaning "at least one," and "a plurality" should be understood as meaning "at least two." Those skilled in the art should interpret these terms appropriately based on the semantics and logical relationships of the context to ensure that they encompass the possibility of "one or more."

[0056] In order to more clearly illustrate the technical solution of the present application, the following will further illustrate it through embodiments of specific scenarios.

[0057] Figure 1 Shown is the composition structure and workflow of the large-scale model safety shield system of the present invention.

[0058] (1) Overall architecture of the large-scale safety shield system of the present invention This system adopts a multi-layered protection architecture, including input layer security protection, business scenario processing layer security alignment, and output layer security audit. Data is exchanged between modules through a secure communication protocol to ensure the security of the entire processing process. The functional modules and technical descriptions of this system are shown in the following table:

[0059] The core workflow of this system is as follows: 1. The user enters the prompt word to enter the system.

[0060] 2. The prompt word shield module performs security detection and matching.

[0061] 3. The prompt words that pass the security test are sent to the multi-dimensional alignment expert model.

[0062] 4. The model generates content that enters the streaming output pipeline.

[0063] 5. The streaming security audit module monitors and filters output content in real time.

[0064] 6. The final compliant content is presented to the user.

[0065] (2) Prompt word shield module The prompt word shield module is the first line of defense for this system, specifically designed to identify and filter harmful user input. Its technical implementation includes: 1. Financial Specific Security Framework A dedicated private domain knowledge base for the financial sector has been built, including: • General security audit set (including crime, bias, ethics, privacy, etc.) • Common financial fraud pattern library (contains more than 500 patterns) • Financial Sensitive Information Identification Rule Set • Cross-turn dialogue attack pattern library • Prompt word attack pattern library (including injection attacks, jailbreak attacks, deception attacks, etc.) This private domain knowledge base specifically focuses on security compliance requirements and risk identification in the financial sector, including advanced threats such as cue word attacks, political bias, fraud detection, and customer privacy breaches. It is constructed through a combination of domain expert annotation and automated learning, and is regularly updated to address new attack vectors. If a risk is identified, pre-set fallback security policies are immediately invoked, extracting or generating fallback scripts from the private domain knowledge base to ensure a secure downstream processing environment.

[0066] 2. Multilingual matching system To support global business, the system implements a multilingual matching mechanism: • Cross-language pattern recognition based on vector matching • Specialized word segmentation and data synthesis for minority languages ​​(including Arabic, Thai, etc.) • Expert review model supporting multiple languages In terms of technical implementation, the following steps are adopted: (1) Convert the input text into a general semantic vector.

[0067] (2) Calculate the similarity with the attack pattern vector in the pre-stored private domain knowledge base.

[0068] (3) When the similarity exceeds the threshold, security interception is triggered.

[0069] (4) When semantic matching fails, it is intercepted again through the expert review model.

[0070] 3. Matching algorithm This system uses an innovative multi-level matching algorithm to effectively intercept harmful input at the user input level:

[0071] The above algorithm specifically incorporates financial domain pattern detection capabilities, including: • Fraud pattern detection: Common fraud patterns • Deception pattern detection: Identify deception techniques such as indirect expressions, segmented instructions, metaphorical hints, etc. • Compliance testing: Compliance requirements of the State Financial Supervision and Administration Bureau Function Description: • ContextAnalysis(input): Analyze the context of user input to capture the relevance of cross-turn dialogues.

[0072] • SemanticMatch(input, knowledge_base, context_analysis): performs preliminary semantic matching and combines contextual information for precise matching.

[0073] • FraudPatternDetection(input, fraud_patterns, context_analysis): Detects whether the input contains financial fraud or sensitive information, taking into account contextual data to improve matching accuracy.

[0074] • SafetyAuditModel(input, context_analysis): Check the safety of the input through the security audit model and make further judgments based on context information.

[0075] Return value description: • If the number of decoy matches exceeds the threshold, the risk type is returned as “fraud” and the request fails.

[0076] • If the semantic match does not meet the minimum threshold, directly return “fail” and determine the risk type.

[0077] • If all checks pass, returns "Passed".

[0078] (3) Multi-dimensional alignment expert model In the business processing flow, the output of the large language model (LLM) contains harmful content and hallucination risks. The multi-dimensional alignment expert model considers multiple dimensions such as values, business needs, and compliance through a post-training process to ensure that the output content is safe, compliant, and meets business needs.

[0079] The main implementation methods include: 1. Post-training expert model Based on the pre-trained model, the expert model implements domain adaptation by introducing a trainable low-rank matrix (Lora) at a specific layer, including: • Design and optimization of scene prompt words • Manual annotation and synthesis of scene data • Supervised fine-tuning • Reinforcement learning based on human feedback • Expert model performance evaluation Figure 2 The following figure shows the detailed post-training process of the expert model, which includes the following steps: Scenario data and safety-related data are used as input data sources, and prompt words are generated by combining the scenario data and safety-related data; the generated prompt words are used to train the basic model, and the basic model is fine-tuned in a supervised manner to improve the performance of the model on specific tasks; the reinforcement learning process of the model is guided by human feedback to further improve the performance of the model; the model that has undergone reinforcement learning is evaluated to determine whether it meets specific requirements, and the model that meets the requirements is quantized to reduce the model's resource consumption and improve operational efficiency. The model that does not meet the requirements returns to the supervised fine-tuning step for further model optimization; the training process ends after the model quantization is completed.

[0080] 2. Multi-dimensional alignment evaluation mechanism Based on the subjective evaluation of large models, the model alignment ability is evaluated from multiple dimensions, including: • Business directive alignment: Ensure that model outputs meet business directive requirements • Instruction attack: Prevent the leakage of business instructions and induce the model to output harmful information • Compliance alignment: ensuring output complies with financial regulatory requirements of various countries • Risk alignment: Avoid risky financial behavior • Values ​​alignment: aligning with the values ​​of people in different regions On the one hand, the multi-dimensional alignment evaluation mechanism provides quantitative indicators for model optimization, helping developers accurately identify defects, make targeted adjustments to training data or algorithms, and promote continuous model evolution; on the other hand, multi-dimensional indicators allow customers to more intuitively judge the reliability of model outputs, improve the user experience, and at the same time, feedback data can feed back into model iteration, forming a virtuous cycle.

[0081]

[0082] Function Description: • GetEvaluationDimensions(evaluation_type): Gets the evaluation dimensions applicable to the evaluation type and their corresponding weights. The returned value may be a list or dictionary containing the dimension names and weights. For example: [("Factual Correctness", 0.25), ("Logical Coherence", 0.2), ("Creativity", 0.15), ...].

[0083] • EvaluationModel(input, reference_answer, alignment_dimensions): Calls the evaluation model to calculate the score for each dimension (1-10 points).

[0084] • overall_score: comprehensive score, which is obtained by weighted calculation of the scores of each dimension (dimensions and weights are dynamic), such as factual correctness, logical coherence, richness, creativity, completeness, clarity and other dimensions.

[0085] (IV) Streaming Security Audit Module (Output Supervision Module) One of the core innovations of the present invention is the implementation of real-time security auditing of streaming output.

[0086] 1. Streaming audit mechanism • Lightweight real-time layer Based on the private domain knowledge base combined with the rule engine, it can quickly intercept explicit risks (such as sensitive words, violent content, etc.).

[0087] • Deep analysis layer A sliding window mechanism is introduced to divide the streaming output into multiple semantic segments, and a high-precision audit model is asynchronously called to perform secondary verification for contextual coherence, implicit bias or complex ethical issues to avoid missed judgments.

[0088] Implement real-time monitoring and review during the large model generation phase to detect potential violations or harmful content in the output information, provide a final security barrier, intercept and correct non-compliant outputs, and ensure that the final output content is safe and compliant.

[0089] 2. Dynamic adjustment algorithm This system is able to adjust output in real time while maintaining a smooth experience:

[0090] Function Description: • PrivateKnowledgeBaseAndRuleEngine.Review(current_segment): Use the private knowledge base and rule engine to perform a lightweight review to determine whether the current segment is safe.

[0091] • CanBeFixed(issue): Determines whether the issue found in the lightweight audit can be fixed.

[0092] • SafeReplace(current_segment, issue): Repairs the current segment based on the issue to make it meet safety requirements.

[0093] • generator.AdjustDirection(feedback): Adjust the generation direction of the generator to solve security issues.

[0094] • HighPrecisionAuditModel.Evaluate(buffer): After all fragments are generated, perform a high-precision security audit on the contents of the buffer.

[0095] Return value description: • If a problem is detected at any stage and cannot be fixed, the build is terminated and the “fallback” is returned.

[0096] • If all fragments pass the review, the buffer content will be returned; otherwise, the “backup script” will be returned.

[0097] 3. Sensitive word repair technology For content marked as unsafe but repairable, this system implements an intelligent repair mechanism: • Use synonym replacement technology to replace sensitive words • Reconstruct sentences that may have unsafe implications • Insert security qualifiers while maintaining semantic consistency (V) Bottom-line guarantee mechanism As the last line of defense, this system has designed a safety net mechanism: 1. Emergency interruption: Immediately interrupt output when high-risk content is detected.

[0098] 2. Fallback script: Provides preset fallback scripts as alternative responses.

[0099] In addition, the above technical solutions of the present invention may be implemented in the following alternative ways in different embodiments: 1. Alternatives to the Prompt Word Shield (1) Rule-based filtering system: Uses a traditional rule engine instead of a knowledge graph to identify attack patterns by manually writing rules. This solution is simple to implement, but has high maintenance costs and poor adaptability to new attacks.

[0100] (2) Independent security model pre-filtering: Use an independent security classification model to pre-filter all inputs. This solution has high accuracy but increases system complexity and latency.

[0101] 2. Alternatives to Multi-dimensional Alignment Expert Models (1) Instruction fine-tuning method: Inject financial security knowledge by fine-tuning the basic model. This solution is relatively simple to implement, but the effect is limited and it is difficult to meet the requirements of multi-dimensional alignment.

[0102] (2) External rule constraint system: Instead of modifying the model itself, the output is constrained by an external rule system. This solution is highly flexible, but may result in unnatural output content and a poor user experience.

[0103] 3. Streaming Security Audit Alternatives (1) Pre-generation and post-review method: First, the content is fully generated, and then it is presented to the user in a streaming manner after it is reviewed and approved. This solution is simple to implement, but it will significantly increase the time consumption and affect the user experience.

[0104] Figure 3 The following is the overall implementation process of the large model security protection method provided by this application, including the following steps: S1. The prompt word shield module receives user input and performs preliminary security checks. This includes real-time detection of user input, identifying and filtering potentially harmful or illegal content. If risks are detected, it immediately invokes pre-set safety strategies and private domain knowledge bases to extract or generate safety-critical phrases, ensuring a secure downstream processing environment. S2. The Multi-Dimensional Alignment Expert Model module performs multi-dimensional alignment on inputs that pass preliminary security checks during business scenario processing. This involves integrating business scenario expertise with adversarial protection technologies. It uses supervised fine-tuning and reinforcement learning to enhance the model's instruction alignment capabilities, ensuring that the generated content is compliant and meets value requirements. S3. The output supervision module monitors and reviews the output content of the business scenario processing layer in real time. This includes real-time monitoring and review during the large model generation phase, detecting potential violations or harmful content in the output information, intercepting and correcting non-compliant outputs, and ensuring the security and compliance of the final output content.

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate possible implementations of the apparatus, methods, and computer program products according to various embodiments of the present application, including architecture, functions, and operations. In these figures, each box may represent a module, a program segment, or a portion of a code, which contains one or more executable instructions for implementing a specified logical function. It should be noted that each box in the block diagram and / or flowchart, and the combination of these boxes, can be implemented using a dedicated hardware-based system to implement the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0106] like Figure 4As shown, an embodiment of the present application further discloses an electronic device, comprising: a processor 310, a communication interface 320, a memory 330 for storing a computer program executable by the processor, and a communication bus 340. The processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 executes the executable computer program to implement the steps of the large model security protection method described above.

[0107] It is understood that, in addition to the memory and processor, the electronic device may also include an input device (e.g., a keyboard), an output device (e.g., a display), and other communication modules. These input devices, output devices, and other communication modules all communicate with the processor via an I / O interface (i.e., an input / output interface).

[0108] The operation of the present application can be implemented by writing computer program code using one or more programming languages ​​or a combination thereof. The programming languages ​​include but are not limited to the following types: Object-oriented programming languages, such as Java, Smalltalk, C++, etc.; A conventional procedural programming language, such as "C" or a similar programming language.

[0109] The execution methods of the program code include but are not limited to: Executes entirely on the user's computer; Partially executed on the user's computer and partially on a remote computer; Executed as a standalone software package; Executes entirely on the remote computer or server.

[0110] In scenarios involving a remote computer, the remote computer can be connected to the user's computer via any type of network, including but not limited to a local area network (LAN) or a wide area network (WAN). Additionally, the remote computer can be connected to an external computer via an Internet service provider, such as the Internet.

[0111] Furthermore, the present application also discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute each step of the large model security protection method disclosed in the present application.

[0112] In the context of this application, computer-readable storage media refers to tangible media that can store computer program code and related data. Specific examples include, but are not limited to, the following: (1) Portable computer disk: A removable magnetic storage medium such as a floppy disk.

[0113] (2) Hard disk: includes fixed storage devices such as mechanical hard disks and solid-state hard disks.

[0114] (3) Random Access Memory (RAM): Volatile storage medium used for temporary storage of data and program code.

[0115] (4) Read-only memory (ROM): A non-volatile storage medium used to store fixed programs and data.

[0116] (5) Erasable Programmable Read-Only Memory (EPROM) or Flash Memory: A non-volatile storage medium that supports multiple erasing and programming.

[0117] (6) Fiber optic storage device: storage medium based on fiber optic technology.

[0118] (7) Compact Disc Read-Only Memory (CD-ROM): A read-only medium that stores data in the form of an optical disc.

[0119] (8) Optical storage devices: storage media based on optical principles, such as DVDs and Blu-ray discs.

[0120] (9) Magnetic storage devices: storage media based on magnetic principles, such as magnetic tapes and disks.

[0121] (10) Any suitable combination of the above: for example, combining multiple storage media to meet different storage requirements.

[0122] These computer-readable storage media can be used to store the program code and related data described in this application to support the operation of the program and the persistent storage of data.

[0123] In particular, according to embodiments of the present application, the processes described in the flowcharts can be implemented as computer software programs. For example, embodiments of the present application relate to a computer program product comprising a computer program carried on a non-transitory computer-readable medium. The computer program includes program code for executing the large model security protection method disclosed in the present application. When the computer program is executed by a processing device, it can implement the above-mentioned functions defined in the embodiments of the present application.

[0124] Although the above discussion contains several specific implementation details, these details should not be interpreted as limiting the scope of this application. The above description is only a preferred embodiment of the present application and an illustration of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features. At the same time, this application should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concepts.

[0125] Those skilled in the art should also understand that they may modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein with equivalents, without departing from the spirit and scope of the technical solutions of the embodiments of the present application. Such modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the core spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A large model safety shield system, characterized in that: The safety shield system includes: The input layer security protection module is used to receive user input and perform preliminary security checks. The input layer security protection module includes a prompt word shield module to detect user input in real time and identify and filter potentially harmful or illegal content; A business scenario processing layer security alignment module is used to perform multi-dimensional alignment processing on inputs that pass preliminary security checks during the business scenario processing process. The business scenario processing layer security alignment module includes a multi-dimensional alignment expert model module, which is used to integrate business scenario expertise with adversarial protection technology, and improve the model instruction alignment capability through supervised fine-tuning and reinforcement learning to ensure that the generated content is compliant and meets value requirements; The output layer security audit module is used to monitor and audit the content output by the business scenario processing layer in real time. The output layer security audit module includes an output supervision module, which is used to implement real-time monitoring and review in the large model generation stage, detect potential violations or harmful content in the output information, intercept and correct non-compliant outputs, and ensure that the final output content is safe and compliant.

2. The safety shield system according to claim 1, characterized in that: The prompt word shield module includes a financial-specific security framework for building a private domain knowledge base dedicated to the financial sector, including a general security audit set, a common financial fraud pattern library, a financial sensitive information identification rule set, and a cross-round dialogue attack pattern library; When the prompt word shield module discovers a risk, it immediately calls the preset safety strategy and private domain knowledge base to extract or generate safety words to ensure the safety of the downstream processing environment.

3. The safety shield system according to claim 1, characterized in that: The input layer security protection module also includes: A multilingual matching system module, designed to support multilingual matching mechanisms for global businesses, includes cross-lingual pattern recognition based on vector matching, specialized word segmentation for minority languages, data synthesis, and an expert review model that supports multiple languages. The matching algorithm module is used to effectively intercept harmful input at the user input layer, including a cross-turn dialogue attack pattern library and a prompt word attack pattern library. The matching algorithm module also includes financial field pattern detection functions, including fraud pattern detection, deception pattern detection, and compliance detection.

4. The safety shield system according to claim 1, characterized in that: The business scenario processing layer security alignment module also includes: The post-training expert model module is used to achieve domain adaptation by introducing trainable low-rank matrices at specific layers based on the pre-trained model. This module includes scene cue word design and optimization, manual annotation and synthesis of scene data, supervised fine-tuning, reinforcement learning based on human feedback, and expert model performance evaluation. The multi-dimensional alignment assessment module is used to evaluate the alignment capabilities of the model from multiple dimensions, including business instruction alignment, instruction attack, compliance alignment, risk alignment, and value alignment.

5. The safety shield system according to claim 4, characterized in that: The working process of the post-training expert model module is as follows: Scenario data and safety-related data are used as input data sources, and prompt words are generated by combining the scenario data and safety-related data; Use the generated prompt words to train the base model and perform supervised fine-tuning on the base model to improve the performance of the model on specific tasks; Use human feedback to guide the model’s reinforcement learning process and further improve the model’s performance; The model that has undergone reinforcement learning is evaluated to determine whether it meets specific requirements. Models that meet the requirements are quantized to reduce the model's resource consumption and improve operating efficiency. Models that do not meet the requirements return to the supervised fine-tuning step for further model optimization. After the model quantization is completed, the training process ends.

6. The safety shield system according to claim 1, characterized in that: The output layer security audit module also includes: Dynamic adjustment algorithm module, used to adjust output in real time while maintaining a smooth experience; The sensitive word repair module is used to intelligently repair content marked as unsafe but repairable. The intelligent repair includes: replacing sensitive words using synonym replacement technology, reconstructing sentences that may have implicit unsafe meanings, or inserting safe qualifiers while maintaining semantic coherence.

7. A large model security protection method, characterized in that: The safety protection method comprises the following steps: S1. The prompt word shield module receives user input and performs preliminary security checks. This includes real-time detection of user input, identifying and filtering potentially harmful or illegal content. If risks are detected, it immediately invokes pre-set safety strategies and private domain knowledge bases to extract or generate safety-critical phrases, ensuring a secure downstream processing environment. S2. The Multi-Dimensional Alignment Expert Model module performs multi-dimensional alignment on inputs that pass preliminary security checks during business scenario processing. This involves integrating business scenario expertise with adversarial protection technologies. It uses supervised fine-tuning and reinforcement learning to enhance the model's instruction alignment capabilities, ensuring that the generated content is compliant and meets value requirements. S3. The output supervision module monitors and reviews the output content of the business scenario processing layer in real time. This includes real-time monitoring and review during the large model generation phase, detecting potential violations or harmful content in the output information, intercepting and correcting non-compliant outputs, and ensuring the security and compliance of the final output content.

8. The safety protection method according to claim 7, characterized in that: Step S1 also includes: Build a dedicated private domain knowledge base for the financial sector, including a general security audit set, a common financial fraud pattern library, a set of financial sensitive information identification rules, and a cross-round dialogue attack pattern library; Implement a multilingual matching mechanism to support global business, including cross-language pattern recognition based on vector matching, specialized word segmentation for minority languages, data synthesis, and an expert review model that supports multiple languages. A multi-level matching algorithm is used to effectively intercept harmful input at the user input layer, including a cross-turn dialogue attack pattern library and a prompt word attack pattern library. The multi-level matching algorithm also includes financial field pattern detection functions, including fraud pattern detection, deception pattern detection and compliance detection.

9. The safety protection method according to claim 7, characterized in that: Step S2 also includes: Based on the pre-trained model, domain adaptation is achieved by introducing trainable low-rank matrices at specific layers, including scene cue word design and optimization, manual annotation and synthesis of scene data, supervised fine-tuning, reinforcement learning based on human feedback, and expert model performance evaluation. Evaluate the model's alignment capabilities from multiple dimensions, including business directive alignment, directive attack, compliance alignment, risk alignment, and value alignment.

10. The safety protection method according to claim 7, characterized in that: Step S3 also includes: Implement real-time monitoring and review during the large model generation phase, including a lightweight real-time layer and a deep analysis layer; Dynamically adjust the algorithm to adjust the output in real time while maintaining a smooth experience; Intelligently repair content marked as unsafe but repairable, including: replacing sensitive words with synonym replacement technology, reconstructing sentences that may have implicit unsafe meanings, or inserting safe qualifiers while maintaining semantic coherence.

Citation Information

Patent Citations

  • System and method for identifying and filtering improper content based on large language model

    CN118484539A

Cited By

  • Safety control method, system and equipment for interaction data of large language model and medium

    CN121051738A

  • Personal information inspection and protection method based on large language model

    CN121093388A

  • A personal information inspection protection method based on a large language model

    CN121093388B

  • Detection method for large model safety protection fence product content value filtering function and related equipment

    CN121212263A

  • Domain large model content security detection method, system, equipment and medium

    CN121580255A