Industrial large model content security capability construction method, apparatus and device, and medium

By embedding and encoding the security knowledge base entries and storing vector databases, combined with the risk judgment of the black and white list library, the problem of prompt word attack recognition of the big model generated content is solved, and the optimization of user interaction experience and the reliability of content security is achieved.

CN120179786APending Publication Date: 2025-06-20SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510337496.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and intercept prompt word attacks without affecting the user's interactive experience, ensuring the safety and reliability of the content generated by the big model.

Method used

Through the bidirectional encoder model, the embedded encoding operation is performed on the knowledge base entries in the security knowledge base, and vector data is generated and stored in the preset vector database; the blacklist and whitelist libraries are created based on regular matching technology, and the risk level of the user input problem is judged, and the answer method is selected based on the risk level.

Benefits of technology

It realizes accurate identification and intercepting prompt word attacks without affecting the user's interactive experience, ensuring the safety and reliability of the content generated by the big model, and improving user satisfaction and trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179786A_ABST
    Figure CN120179786A_ABST
Patent Text Reader

Abstract

The invention discloses an industry large model content security capability construction method and device, equipment and a medium, and relates to the technical field of computers. Comprising the following steps: executing an embedded coding operation on knowledge base entries in a safe knowledge base through a bidirectional encoder model to obtain vector data containing original text literal meaning and semantic features of the knowledge base entries, and storing the vector data in a preset vector database; creating a blacklist library and a white list library based on a regular matching technology, and judging a risk level of a user input problem based on data stored in the blacklist library and the white list library; and selecting a corresponding answering mode based on the risk level of the question input by the user. Therefore, on the premise that user interaction experience is not affected, hint word attacks can be accurately recognized and intercepted, and it is ensured that large model generation content is safe and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a method, device, equipment and medium for constructing the content security capability of an industry large model. Background Art

[0002] In the context of the rapid development of artificial intelligence technology today, the application of industry large models has penetrated into various fields, providing efficient and intelligent services. However, with the increasingly powerful text generation capabilities of these models, content security issues have gradually emerged. Generative large models can not only bring great convenience, but also be used for improper purposes, such as generating inappropriate or harmful content through prompt attacks. Such behaviors not only damage the security of the content and the credibility of the model, but also pose challenges to the stable operation of the model and the user experience. Traditional security technologies mainly focus on the review of static content, and for dynamically generated content, especially the information generated in real time by large models, their detection and protection mechanisms are still insufficient. Current defense measures, such as manual review and keyword filtering, although can prevent the spread of improper content to a certain extent, are inefficient and difficult to cover comprehensively when facing complex and changeable threats. In addition, due to the complexity of the large model structure and the diversity of the generated content, this further increases the difficulty of content security detection.

[0003] As can be seen from the above, how to accurately identify and intercept prompt attacks without affecting the user interaction experience and ensure the security and reliability of the content generated by large models is an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for constructing the content security capability of an industry large model, which can accurately identify and intercept prompt attacks without affecting the user interaction experience and ensure the security and reliability of the content generated by large models. The specific solutions are as follows:

[0005] In a first aspect, the present application provides a method for constructing the content security capability of an industry large model, including:

[0006] Performing an embedding encoding operation on the knowledge base entries in the security knowledge base through a bidirectional encoder model to obtain vector data containing the literal meaning and semantic features of the original text of the knowledge base entries, and storing the vector data in a preset vector database;

[0007] Create a blacklist library and a whitelist library based on regular matching technology, and determine the risk level of the user input problem based on the data stored in the blacklist library and the whitelist library; the blacklist library is used to store sensitive word data with risk characteristics; the whitelist library is used to store user data that is incorrectly intercepted by the security model on the business side; the risk levels include a first risk level and a second risk level;

[0008] Select a corresponding answer method based on the risk level of the user input problem; among them, if the risk level of the user input problem is the first risk level, reject the user input problem and perform risk interception; if the risk level of the user input problem is the second risk level, generate an answer text that meets the content security conditions based on the data in the preset vector database.

[0009] Optionally, the method for constructing the content security capability of the industry large model further includes:

[0010] If a user input problem is received, perform an embedding encoding operation on the user input problem through the bidirectional encoder model to obtain target vector data corresponding to the user input problem;

[0011] Perform a similarity search operation based on the target vector data to determine a target knowledge base entry corresponding to the user input problem from the preset vector database;

[0012] Correspondingly, generating an answer text that meets the content security conditions based on the data in the preset vector database includes:

[0013] Generate an answer text that meets the content security conditions based on the target knowledge base entry;

[0014] Among them, the similarity search operation is an operation of searching for a knowledge base entry corresponding to the target vector data from the preset vector database based on a preset similarity threshold.

[0015] Optionally, determining the risk level of the user input problem based on the data stored in the blacklist library and the whitelist library includes:

[0016] If the user input problem contains data that matches the sensitive word data in the blacklist library, the risk level of the user input problem is the first risk level;

[0017] If the user input problem contains data that matches the user data in the whitelist library, the risk level of the user input problem is the second risk level;

[0018] If the user input question does not contain data that matches the sensitive word data in the blacklist library and the user data in the whitelist library, the risk level of the user input question is determined by a preset classification model.

[0019] Optionally, determining the risk level of the user input question by a preset classification model includes:

[0020] Converting the original text data of the user input question into a corresponding vocabulary unit sequence; the vocabulary unit sequence includes the coding identifier and position information corresponding to each vocabulary unit;

[0021] Mapping the vocabulary units and the position information corresponding thereto to a high-dimensional space to generate corresponding vector data;

[0022] The initial classification model is trained by positive and negative training samples that meet the preset positive and negative sample ratios to obtain a trained classification model; the positive and negative training samples include positive training samples composed of historical user report data, manually annotated data, and public data sets, and negative training samples composed of normal conversation data, news data, and real business data;

[0023] The trained classification model is used to analyze the vector data to determine the risk level of the user input question.

[0024] Optionally, the generating a reply text that meets the content security condition based on the data in the preset vector database includes:

[0025] The answer content generated based on the data in the preset vector database is limited by setting prompt words, and the answer content is fine-tuned and optimized by a preset fine-tuning data set to obtain an answer text that meets the content security conditions;

[0026] The preset fine-tuning data set includes historical case data that meets content security conditions.

[0027] Optionally, the fine-tuning and optimizing the answer content by using a preset fine-tuning data set to obtain an answer text that meets content security conditions includes:

[0028] Fine-tune the answer content using the Lora method based on the preset fine-tuning data set;

[0029] The answer content is optimized based on the DPO algorithm to obtain an answer text that meets content security requirements.

[0030] Optionally, the method for building industry large model content security capabilities further includes:

[0031] Receive a data modification instruction through a preset standardized data interface and verify the data modification instruction;

[0032] Parse the data modification instruction that passes the verification, and extract the dynamic operation instruction set therein. The dynamic operation instruction set includes at least one of addition, deletion, and update operations for the blacklist and / or whitelist;

[0033] Execute a data modification operation on the target list library based on the dynamic operation instruction set; the target list library is the blacklist library and / or whitelist library.

[0034] In a second aspect, the present application provides an industrial large model content security capability construction device, including:

[0035] A knowledge base entry storage module, configured to perform an embedding encoding operation on the knowledge base entries in the security knowledge base through a bidirectional encoder model to obtain vector data including the literal meaning and semantic features of the original text of the knowledge base entries, and store the vector data in a preset vector database;

[0036] A user question risk level judgment module, configured to create a blacklist library and a whitelist library based on regular matching technology, and judge the risk level of the user input question based on the data stored in the blacklist library and the whitelist library; the blacklist library is used to store sensitive word data with risk characteristics; the whitelist library is used to store user data that is incorrectly intercepted by the security model on the business side; the risk levels include a first risk level and a second risk level;

[0037] A user question response module, configured to select a corresponding answer method based on the risk level of the user input question; wherein, if the risk level of the user input question is the first risk level, reject the user input question and perform risk interception; if the risk level of the user input question is the second risk level, generate an answer text that meets the content security conditions based on the data in the preset vector database.

[0038] In a third aspect, the present application provides an electronic device, including:

[0039] A memory, configured to store a computer program;

[0040] A processor, configured to execute the computer program to implement the foregoing industrial large model content security capability construction method.

[0041] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, wherein the computer program, when executed by a processor, implements the foregoing industrial large model content security capability construction method.

[0042] The present application provides a method for constructing the content security capability of an industry large model. First, an embedding encoding operation is performed on the knowledge base entries in the security knowledge base through a bidirectional encoder model to obtain vector data containing the literal meaning and semantic features of the original text of the knowledge base entries, and the vector data is stored in a preset vector database; then, a blacklist library and a whitelist library are created based on regular matching technology, and the risk level of the user input question is judged based on the data stored in the blacklist library and the whitelist library; the blacklist library is used to store sensitive word data with risk features; the whitelist library is used to store user data that is wrongly intercepted by the security model on the business side; the risk levels include high risk level, medium risk level, and low risk level; finally, a corresponding answering method is selected based on the risk level of the user input question; wherein, if the risk level of the user input question is the high risk level, a refusal operation is performed on the user input question and risk interception is carried out; if the risk level of the user input question is the medium risk level or the low risk level, an answer text that meets the content security conditions is generated based on the data in the preset vector database.

[0043] As can be seen from the above, the present application stores the vector data containing the literal meaning and semantic features of the original text of the security knowledge base entries in a preset vector database, and can accurately determine the knowledge base entries related to the user input question through similarity retrieval; the risk level of the user input question is divided by the black and white lists, which improves the judgment ability of the industry large model for risk problems, and generates a more positive answer based on the risk level instead of a single refusal answer, greatly improving the user's satisfaction and trust. Thus, it is possible to accurately identify and intercept prompt word attacks without affecting the user interaction experience, ensuring the security and reliability of the content generated by the large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0045] Figure 1 It is a flowchart of a method for constructing the content security capability of an industry large model disclosed in the present application;

[0046] Figure 2 It is a flowchart of a specific method for constructing the content security capability of an industry large model disclosed in the present application;

[0047] Figure 3 It is a schematic diagram of a device for constructing the content security capability of an industry large model disclosed in the present application;

[0048] Figure 4 This is a structural diagram of an electronic device disclosed in the present application. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] In the context of the rapid development of artificial intelligence technology today, the application of industry large models has penetrated into various fields, providing efficient and intelligent services. However, as the text generation ability of these models becomes increasingly powerful, content security issues have gradually emerged. Generative large models can not only bring great convenience, but may also be used for improper purposes, such as generating inappropriate or harmful content through prompt attacks. Such behaviors not only damage the security of the content and the credibility of the model, but also pose challenges to the stable operation of the model and the user experience. Traditional security technologies mostly focus on the review of static content, and for dynamically generated content, especially the information generated in real time by large models, their detection and protection mechanisms are still insufficient. Current defense measures, such as manual review and keyword filtering, although can prevent the spread of improper content to a certain extent, are inefficient and difficult to cover comprehensively when facing complex and changing threats. In addition, due to the complexity of the large model structure and the diversity of the generated content, this further increases the difficulty of content security detection. Therefore, the present application provides a solution for constructing the content security capability of the industry large model, which can accurately identify and intercept prompt attacks without affecting the user interaction experience, and ensure the security and reliability of the content generated by the large model.

[0051] See Figure 1 As shown, the embodiments of the present application disclose a method for constructing the content security capability of an industry large model, including:

[0052] Step S11: Perform an embedding encoding operation on the knowledge base entries in the security knowledge base through a bidirectional encoder model to obtain vector data containing the literal meaning and semantic features of the knowledge base entries, and store the vector data in a preset vector database.

[0053] In this embodiment, the entries in the security knowledge base are encoded to understand the semantic information of the text and convert natural language into a representation in a low-dimensional vector space. By invoking the selected BGE (Bidirectional Encoder Representations from Transformers, an open-source Chinese-English semantic vector model) model, an embedding encoding operation is performed on all entries in the security knowledge base to convert them into vector data with a fixed length. These vector data not only contain the literal meaning of the original text but also retain semantic features, facilitating subsequent similarity calculations. When the user submits a query request, the question is first preprocessed, including removing extra spaces, standardizing the format, etc., to reduce noise interference and improve matching accuracy. Then, the BGE model used in the preprocessing stage is utilized to perform an embedding operation on the question input by the user to generate a corresponding vector representation. Specifically, if a user input question is received, an embedding encoding operation is performed on the user input question through the bidirectional encoder model to obtain target vector data corresponding to the user input question; a similarity search operation is performed based on the target vector data to determine a target knowledge base entry corresponding to the user input question from the preset vector database, and an answer text meeting the content security conditions is generated based on the target knowledge base entry. Among them, the similarity search operation is an operation of searching for a knowledge base entry corresponding to the target vector data from the preset vector database based on a preset similarity threshold. It is worth mentioning that by setting a reasonable similarity threshold, irrelevant entries can be effectively filtered out, and only those truly relevant and trustworthy answers are returned.

[0054] Step S12: Create a blacklist library and a whitelist library based on regular matching technology, and determine the risk level of the user input question based on the data stored in the blacklist library and the whitelist library.

[0055] In this embodiment, by creating a blacklist library of sensitive words, for known high-risk expression forms, high-risk words are intercepted through regular expression matching technology. At the same time, in order to avoid the subsequent large model from affecting normal business operations, a whitelist mechanism is added. The whitelist contains user data that has been misjudged by the security model on the business side. Hitting the whitelist will directly release it to ensure that they will not be wrongly intercepted. Specifically, the blacklist library is used to store sensitive word data with risk characteristics; the whitelist library is used to store user data that has been wrongly intercepted by the security model on the business side. Specifically, judging the risk level of the user input problem based on the data stored in the blacklist library and the whitelist library may include: if the user input problem contains data that matches the sensitive word data in the blacklist library, the risk level of the user input problem is the first risk level; if the user input problem contains data that matches the user data in the whitelist library, the risk level of the user input problem is the second risk level; if the user input problem does not contain data that matches the sensitive word data in the blacklist library and the user data in the whitelist library, the risk level of the user input problem is determined by a preset classification model. That is to say, this combination of black and white lists can not only effectively prevent security threats but also ensure the normal usage experience of legitimate users. It is worth mentioning that the black and white lists can be flexibly configured through a preset interface. Specifically, the method for constructing the content security capability of the industry large model may further include: receiving a data modification instruction through a preset standardized data interface and verifying the data modification instruction; parsing the data modification instruction that passes the verification and extracting the dynamic operation instruction set therein, the dynamic operation instruction set includes at least one of the operations of adding, deleting, and updating the blacklist and / or the whitelist; performing a data modification operation on the target list library based on the dynamic operation instruction set; the target list library is the blacklist library and / or the whitelist library. That is to say, the flexible configuration of the black and white lists can facilitate adjustment according to the actual situation and ensure the efficient operation of the system.

[0056] Furthermore, to address the poor user interaction experience that may result from rejecting all responses to risk questions, an intelligent response mechanism based on risk subtype classification is proposed. This mechanism can further classify the risk types of the questions input by users into high-risk (should be rejected) and medium-low-risk (should be positively replied). To more precisely handle different types of potential security risk questions, the present invention further classifies the risk type risk_type and the risk subtype risk_sub_type. Specifically, determining the risk level of the user input question through a preset classification model may include: converting the original text data of the user input question into a corresponding sequence of lexical units; the sequence of lexical units contains the coding identifiers and position information corresponding to each lexical unit; mapping the lexical units and their corresponding position information to a high-dimensional space to generate corresponding vector data; training an initial classification model with positive and negative training samples that meet a preset positive-negative sample ratio to obtain a trained classification model; the positive and negative training samples include positive training samples composed of historical user report data, manually labeled data, and public data sets, and negative training samples composed of normal dialogue data, news data, and real data on the business side; using the trained classification model to analyze the vector data to determine the risk level of the user input question. Among them, a risk subtype of -N indicates high risk and should be rejected; a risk subtype of -A indicates medium-low risk and a positive guidance model should be called for a reply.

[0057] Step S13: Select a corresponding response method based on the risk level of the user input question.

[0058] In this embodiment, if the risk level of the user input question is the first risk level, a rejection operation is performed on the user input question and risk interception is carried out; if the risk level of the user input question is the second risk level, a response text that meets the content security conditions is generated based on the data in the preset vector database. Among them, a positive guidance model is separately trained for the above medium and low risk data to generate responses, rather than directly calling the industry large model, to ensure the correctness of the answers. Specifically, the generating of the response text that meets the content security conditions based on the data in the preset vector database may include: limiting the response content generated based on the data in the preset vector database by setting prompt words, and fine-tuning and optimizing the response content through a preset fine-tuning data set to obtain a response text that meets the content security conditions; wherein, the preset fine-tuning data set includes historical case data that meets the content security conditions. That is, by selecting high-quality historical textbooks, official websites, authoritative news reports and other authoritative materials as the fine-tuning data set, to ensure that the model learns correct historical and cultural knowledge, as well as official policies and positions. These data sources not only cover a wide range of historical backgrounds, but also include the latest policy interpretations and official statements, ensuring the timeliness and accuracy of the model's answers. Among them, the data set also contains some positively guided cases reviewed by experts to help the model better understand how to respond appropriately to sensitive topics. These cases include the correction of common misunderstandings, the objective interpretation of historical events, and the supportive statements of official policies.

[0059] Furthermore, the Lora fine-tuning technology (Lora, i.e., Low-Rank Adaptation) and DPO reinforcement learning technology (DPO, i.e., Direct Preference Optimization) are used to enhance the discrimination capability of the forward guidance model. Specifically, the fine-tuning and optimization of the answer content by presetting the fine-tuning data set to obtain the answer text that meets the content security conditions may include: fine-tuning using the Lora method, updating the decomposed low-rank matrix, and reducing the training time and video memory requirements. Compared with full parameter fine-tuning, Lora can significantly reduce resource consumption while maintaining model performance, and is particularly suitable for rapid iteration and deployment of large-scale models. The deepspeed framework is used for training to optimize training efficiency and resource utilization. deepspeed provides a variety of optimization techniques, such as mixed precision training, gradient accumulation, distributed training, etc., which can significantly accelerate the training process and improve the convergence speed of the model. After fine-tuning, the direct preference optimization (DPO, Direct Preference Optimization) algorithm based on deep reinforcement learning is further introduced to achieve the value alignment of the model. The DPO algorithm directly optimizes the model's preference for generating text by comparing the quality of different answers, making it more consistent with human values ​​and social norms.

[0060] As can be seen from the above, the embodiment of the present application stores the vector data containing the literal meaning and semantic features of the original text of the security knowledge base entry in a preset vector database, and can accurately determine the knowledge base entry related to the user input question by similarity retrieval; divide the risk level of the user input question by black and white lists, improve the industry big model's ability to judge risk issues, and generate more positive answers based on risk levels, rather than a single refusal to answer, which greatly improves user satisfaction and trust. In this way, it is possible to accurately identify and intercept prompt word attacks without affecting the user's interactive experience, ensuring the security and reliability of the content generated by the big model.

[0061] See also Figure 2 As shown, the embodiment of the present application discloses a specific method for building content security capabilities of a large industry model, including:

[0062] In this embodiment, multiple core modules are included: a security knowledge base retrieval module based on similarity matching, a black / white list retrieval module based on regular matching, a rejection / rejection module based on a generative large model, and a positive guidance response module based on a generative large model.

[0063] In this embodiment, the security knowledge base retrieval module selects a Bi-Encoder model with excellent performance to encode the entries in the security knowledge base. This model can understand the semantic information of the text and convert natural language into a representation in a low-dimensional vector space. By invoking the selected BGE model, perform embedding encoding operations on all entries in the security knowledge base and convert them into vector data with a fixed length. These vector data not only contain the literal meaning of the original text but also retain semantic features, facilitating subsequent similarity calculations. Store the encoded knowledge base entries in an efficient vector database such as ChromaDB. Quickly find the closest vector, that is, the most similar knowledge entry, in a large dataset. When the user submits a query request, first preprocess the question, including removing extra spaces, standardizing the format, etc., to reduce noise interference and improve matching accuracy. Then, use the same BGE model as in the preprocessing stage to perform an embedding operation on the user's input question to generate the corresponding vector representation. Next, perform a similarity search in ChromaDB to find the knowledge base entry closest to the user's question vector. By setting a reasonable similarity threshold, irrelevant entries can be effectively filtered out, and only those truly relevant and trustworthy answers are returned.

[0064] In this embodiment, the black / white list retrieval module based on regular matching establishes a special blacklist library of sensitive words and intercepts these words through regular expression matching technology for known high-risk expression forms. This method not only significantly reduces latency but also ensures the efficiency and scalability of the system. At the same time, to avoid the subsequent large model affecting normal business operations, this module adds a whitelist mechanism. The whitelist contains user data misjudged by the security model on the business side, and hitting the whitelist will directly release them to ensure that they will not be wrongly intercepted. This combination of black and white lists can both effectively prevent security threats and guarantee the normal usage experience of legitimate users.

[0065] In this embodiment, the rejection / denial module based on the generative large model selects Qwen1.5-1.8B-Intruct as the base model, including the following core components:

[0066] Input processing module: Responsible for converting the original text data into a sequence of lexical units through a tokenization algorithm, with each unit attached with the corresponding encoding identifier and position information.

[0067] Embedding layer: Maps the lexical units and their position information to a high-dimensional space to form a multi-dimensional vector representation. This process is completed using a pre-trained word embedding matrix, and position encoding is added to retain the position information in the sequence.

[0068] Multi-layer transformer encoder: It consists of a series of sub-modules with consistent structures, each of which integrates self-attention mechanism, multi-head attention, hierarchical normalization, residual connection and feedforward neural network. These mechanisms work together to improve the model's ability to understand complex patterns in text and the efficiency of feature extraction. In particular, the self-attention mechanism enables the model to dynamically focus on key elements when processing long sequences, while multi-head attention enhances the model's ability to capture diverse features. It should be noted that in order to balance the learning process of the model and avoid overfitting or underfitting, we set a reasonable positive and negative sample ratio. Among them, the sources of positive samples include: User report data: Violation content reported by users and confirmed by manual review is collected from multiple online platforms. These data provide common violation cases in the real world, making the model closer to practical applications. Specialized labeled data: Specialized personnel annotate specific types of illegal content after cleaning to ensure the accuracy and authority of the data. Public datasets: Use existing public datasets, such as official lists of illegal content, labeled data in academic research, etc. For questions that do have risks, in addition to marking the risk type, as mentioned above, they should also be marked as the type of positive guidance for rejection / non-rejection. Used to improve user interactivity. Negative sample sources include: Normal conversation data: A large amount of normal user conversation data is collected from social media, forums, chat records and other channels. These data have been cleaned and preprocessed to remove possible sensitive information to ensure that they are suitable for training. News articles: Reports from authoritative news media are selected, covering a variety of topics to ensure that the model can understand a wide range of topic backgrounds and improve its generalization ability. Real data on the business side: Real business data provided by the business side.

[0069] Furthermore, the Low-Rank Adaptation (Lora) technique is adopted for local parameter updates. Compared with the traditional full-parameter fine-tuning method, Lora only needs to adjust a small number of parameters, significantly reducing the training time and video memory requirements. Specifically, the deepspeed framework is used to accelerate the training process, and key hyperparameters are set according to industry best practices, such as the cutoff length, lr_scheduler_type, number of epochs, Lora rank, and Lora scaling factor. Among them, a suitable cutoff length is set according to the actual application scenario and data characteristics, usually set to 512 or 1024, to balance model performance and computing resources. A longer cutoff length can capture more context information but also increase the computational complexity; select a learning rate scheduling strategy suitable for the fine-tuning task, such as linear decay or cosine annealing. Linear decay is suitable for most fine-tuning tasks, while cosine annealing can gradually reduce the learning rate in the later stage of training to prevent overfitting; adjust the number of training epochs according to the experimental results, usually set to 3-5 epochs, to ensure that the model achieves the best performance without overfitting. Too many training epochs may lead to model overfitting and affect its generalization ability; a lower rank value can effectively reduce the computational overhead while still maintaining sufficient expressive power; a higher scaling factor can enhance the effect of Lora and make the model better adapt to new tasks; set a suitable learning rate to ensure stable convergence of the model during fine-tuning. A lower learning rate helps prevent drastic changes in the model during fine-tuning and maintains its original knowledge structure. For example, in a specific implementation, the learning rate is set to one-tenth of the base model's pre-training value (i.e., 5e^-5), and the Lora rank is set to 8 to ensure that the model can still fully absorb new knowledge while maintaining high efficiency.

[0070] In this embodiment, the forward guidance response module based on the generative large model selects the qwen2.5-3B-Instruct model as the base model and sets specific prompts to ensure that the generated content conforms to the official stance. High-quality historical textbooks, official websites, authoritative news reports and other authoritative materials are selected as the fine-tuning data set to ensure that the model learns correct historical and cultural knowledge, as well as official policies and stances. These data sources not only cover a wide range of historical backgrounds, but also include the latest policy interpretations and official statements, ensuring the timeliness and accuracy of the model's answers. The data set also contains some positively guided cases reviewed by experts to help the model better understand how to respond appropriately to sensitive topics. These cases include corrections of common misunderstandings, objective explanations of historical events, and supportive statements of official policies.

[0071] Furthermore, this embodiment introduces the Direct Preference Optimization (DPO) algorithm based on deep reinforcement learning to achieve the alignment of the model's values. The DPO algorithm directly optimizes the preference of the model to generate text by comparing the quality of different answers, making it more in line with human values and social norms. The preference pair data format required for DPO training is <prompt, chosen, rejected>. Prompt is the user input, and chosen and rejected are two different answers, with chosen being considered the more appropriate and better answer.

[0072] As can be seen from the above, the embodiment of the present application can more accurately identify and intercept security-induced attacks and other forms of risk inputs by introducing a security knowledge base retrieval module based on similarity matching and a rejection classification module based on a generative large model, effectively preventing the generation of inappropriate or harmful content, and ensuring the security and reliability of the output content of the large model. The application of the positive guidance reply module enables the system to not only correctly reject inappropriate requests, but also provide valuable suggestions and support to users, thereby enhancing the friendliness and humanization of the system without affecting the user's interactive experience. Especially for medium and low-risk issues, the system can give more positive answers rather than a single rejection speech, which greatly improves user satisfaction and trust. Combined with the vector database ChromaDB and the BGE model, it is possible to quickly find the closest knowledge entry in a large-scale data set, improve the accuracy and response speed of the query results, and achieve correct responses to special questions to avoid errors that may be introduced by the randomness of the content generated by the large model. By adopting the black and white list mechanism, Lora fine-tuning technology, and DPO reinforcement learning technology, the discrimination ability of the security model is improved. It reduces unnecessary computing resource consumption and reduces the maintenance and operation costs of the system. At the same time, the flexible configuration of the black and white lists also facilitates adjustments based on actual conditions, ensuring the efficient operation of the system.

[0073] Accordingly, see Figure 3 As shown, the embodiment of the present application discloses a device for building content security capabilities of an industry large model, including:

[0074] A knowledge base entry storage module 11 is used to perform an embedding encoding operation on the knowledge base entries in the security knowledge base through a bidirectional encoder model to obtain vector data containing the original text literal meaning and semantic features of the knowledge base entries, and store the vector data in a preset vector database;

[0075] The user question risk level judgment module 12 is used to create a blacklist library and a whitelist library based on regular matching technology, and judge the risk level of the user input question based on the data stored in the blacklist library and the whitelist library; the blacklist library is used to store sensitive word data with risk characteristics; the whitelist library is used to store user data that is incorrectly intercepted by the security model on the business side; the risk level includes a first risk level and a second risk level;

[0076] The user question response module 13 is used to select a corresponding answer method based on the risk level of the user input question; if the risk level of the user input question is the first risk level, the user input question is rejected and risk interception is performed; if the risk level of the user input question is the second risk level, an answer text that meets the content security conditions is generated based on the data in the preset vector database.

[0077] As can be seen from the above, the embodiment of the present application stores the vector data containing the literal meaning and semantic features of the original text of the security knowledge base entry in a preset vector database, and can accurately determine the knowledge base entry related to the user input question by similarity retrieval; divide the risk level of the user input question by black and white lists, improve the industry big model's ability to judge risk issues, and generate more positive answers based on risk levels, rather than a single refusal to answer, which greatly improves user satisfaction and trust. In this way, it is possible to accurately identify and intercept prompt word attacks without affecting the user's interactive experience, ensuring the security and reliability of the content generated by the big model.

[0078] In some specific implementations, the user problem risk level determination module 12 may specifically include:

[0079] A first risk level judgment unit, configured to determine that if the question input by the user contains data matching the sensitive word data in the blacklist library, the risk level of the question input by the user is a first risk level;

[0080] A second risk level judgment unit, configured to determine that the risk level of the question input by the user is a second risk level if the question input by the user contains data that matches the user data in the whitelist library;

[0081] The risk level judgment submodule is used to determine the risk level of the user input question through a preset classification model if the user input question does not contain data matching the sensitive word data in the blacklist library and the user data in the whitelist library.

[0082] Furthermore, in some specific implementations, the risk level determination submodule may specifically include:

[0083] A vocabulary unit sequence conversion unit, used to convert the original text data of the user input question into a corresponding vocabulary unit sequence; the vocabulary unit sequence includes a coding identifier and position information corresponding to each vocabulary unit;

[0084] A vector data generating unit, used for mapping the vocabulary units and the position information corresponding thereto into a high-dimensional space to generate corresponding vector data;

[0085] A classification model training unit is used to train the initial classification model through positive and negative training samples that meet a preset positive and negative sample ratio to obtain a trained classification model; the positive and negative training samples include positive training samples composed of historical user report data, manually annotated data, and public data sets, and negative training samples composed of normal conversation data, news data, and real business data;

[0086] A risk level determination unit for analyzing the vector data using the trained classification model to determine the risk level of the user input question.

[0087] In some specific embodiments, the user question response module 13 may specifically include:

[0088] An answer text generation sub-module for limiting the answer content generated based on the data in the preset vector database by setting prompt words, and fine-tuning and optimizing the answer content through a preset fine-tuning data set to obtain an answer text that meets the content security conditions; wherein the preset fine-tuning data set includes historical case data that meets the content security conditions.

[0089] Further, in some specific embodiments, the answer text generation sub-module may specifically include:

[0090] An answer text fine-tuning unit for fine-tuning the answer content based on the preset fine-tuning data set using the Lora method;

[0091] An answer text optimization unit for optimizing the answer content based on the DPO algorithm to obtain an answer text that meets the content security conditions.

[0092] In some specific embodiments, the industry large model content security capability construction device may further include:

[0093] A target vector data generation unit for, if a user input question is received, performing an embedding encoding operation on the user input question through the bidirectional encoder model to obtain target vector data corresponding to the user input question;

[0094] A target knowledge base entry determination unit for performing a similarity search operation based on the target vector data to determine a target knowledge base entry corresponding to the user input question from the preset vector database;

[0095] A data modification instruction receiving unit for receiving a data modification instruction through a preset standardized data interface and verifying the data modification instruction;

[0096] A dynamic operation instruction set extraction unit for parsing the verified data modification instruction and extracting the dynamic operation instruction set therefrom, the dynamic operation instruction set including at least one of addition, deletion, and update operations for the blacklist and / or whitelist;

[0097] A data modification operation execution unit for performing a data modification operation on the target list library based on the dynamic operation instruction set; the target list library is the blacklist library and / or the whitelist library;

[0098] Correspondingly, the user question response module 13 may further include:

[0099] An answer text generation unit, configured to generate an answer text that meets content security conditions based on the target knowledge base entry.

[0100] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation to the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the industry large model content security capability construction method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0101] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0102] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.

[0103] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the industry large model content security capability construction method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks.

[0104] Further, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the method for constructing the content security capability of the industrial large model disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0105] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts.

[0106] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0107] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the technical field.

[0108] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0109] The above has introduced the technical solution provided by the present application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for building content security capabilities of an industry large model, characterized in that: include: Performing an embedding encoding operation on knowledge base entries in the security knowledge base through a bidirectional encoder model to obtain vector data containing the original text literal meaning and semantic features of the knowledge base entries, and storing the vector data in a preset vector database; Creating a blacklist library and a whitelist library based on regular matching technology, and judging the risk level of user input problems based on the data stored in the blacklist library and the whitelist library; The blacklist database is used to store sensitive word data with risk characteristics; The whitelist library is used to store user data that is erroneously intercepted by the security model on the business side; the risk level includes a first risk level and a second risk level; Selecting a corresponding answering method based on the risk level of the question input by the user; wherein, if the risk level of the question input by the user is the first risk level, refusing to answer the question input by the user and performing risk interception; If the risk level of the question input by the user is the second risk level, an answer text that meets the content security condition is generated based on the data in the preset vector database.

2. The method for building industry large model content security capabilities according to claim 1 is characterized in that: Also includes: If a user input question is received, embedding encoding is performed on the user input question through the bidirectional encoder model to obtain target vector data corresponding to the user input question; Performing a similarity search operation based on the target vector data to determine a target knowledge base entry corresponding to the user input question from the preset vector database; Accordingly, the step of generating a reply text that meets the content security condition based on the data in the preset vector database includes: Generate a response text that meets content security conditions based on the target knowledge base entry; The similarity search operation is an operation of searching the preset vector database for knowledge base entries corresponding to the target vector data based on a preset similarity threshold.

3. The method for building industry large model content security capabilities according to claim 1 is characterized in that: The determining the risk level of the user input problem based on the data stored in the blacklist library and the whitelist library includes: If the question input by the user contains data matching the sensitive word data in the blacklist library, the risk level of the question input by the user is the first risk level; If the question input by the user contains data that matches the user data in the whitelist library, the risk level of the question input by the user is the second risk level; If the user input question does not contain data that matches the sensitive word data in the blacklist library and the user data in the whitelist library, the risk level of the user input question is determined by a preset classification model.

4. The method for building industry large model content security capabilities according to claim 3 is characterized in that: Determining the risk level of the user input problem by a preset classification model includes: Converting the original text data of the user input question into a corresponding vocabulary unit sequence; the vocabulary unit sequence includes the coding identifier and position information corresponding to each vocabulary unit; Mapping the vocabulary units and the position information corresponding thereto to a high-dimensional space to generate corresponding vector data; The initial classification model is trained by positive and negative training samples that meet the preset positive and negative sample ratios to obtain a trained classification model; the positive and negative training samples include positive training samples composed of historical user report data, manually annotated data, and public data sets, and negative training samples composed of normal conversation data, news data, and real business data; The trained classification model is used to analyze the vector data to determine the risk level of the user input question.

5. The method for building industry large model content security capabilities according to claim 1 is characterized in that: The step of generating a reply text that meets the content security condition based on the data in the preset vector database includes: The answer content generated based on the data in the preset vector database is limited by setting prompt words, and the answer content is fine-tuned and optimized by a preset fine-tuning data set to obtain an answer text that meets the content security conditions; The preset fine-tuning data set includes historical case data that meets content security conditions.

6. The method for building industry large model content security capabilities according to claim 5 is characterized in that: The fine-tuning and optimizing the answer content by using a preset fine-tuning data set to obtain an answer text that meets the content security conditions includes: Fine-tune the answer content using the Lora method based on the preset fine-tuning data set; The answer content is optimized based on the DPO algorithm to obtain an answer text that meets content security requirements.

7. The method for building industry large model content security capabilities according to any one of claims 1 to 6, characterized in that: Also includes: Receiving a data modification instruction through a preset standardized data interface and verifying the data modification instruction; Parsing the verified data modification instructions, extracting a dynamic operation instruction set therein, wherein the dynamic operation instruction set includes at least one of adding, deleting, and updating operations for the blacklist and / or the whitelist; Based on the dynamic operation instruction set, a data modification operation is performed on a target list library; the target list library is the blacklist library and / or the whitelist library.

8. An industry large model content security capability building device, characterized in that: include: A knowledge base entry storage module, used to perform an embedding encoding operation on the knowledge base entries in the security knowledge base through a bidirectional encoder model to obtain vector data containing the original text literal meaning and semantic features of the knowledge base entries, and store the vector data in a preset vector database; A user question risk level judgment module is used to create a blacklist library and a whitelist library based on regular matching technology, and judge the risk level of the user input question based on the data stored in the blacklist library and the whitelist library; The blacklist database is used to store sensitive word data with risk characteristics; The whitelist library is used to store user data that is erroneously intercepted by the security model on the business side; The risk level includes a first risk level and a second risk level; A user question response module, configured to select a corresponding answer method based on the risk level of the question input by the user; wherein, if the risk level of the question input by the user is the first risk level, the question input by the user is rejected and risk interception is performed; If the risk level of the question input by the user is the second risk level, an answer text that meets the content security condition is generated based on the data in the preset vector database.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the method for building industry big model content security capabilities as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, it implements the method for building industry big model content security capabilities as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Method and system for determining voice reply strategy based on dynamic risk assessment

    CN120748402A

  • A method and system for determining voice response strategies based on dynamic risk assessment

    CN120748402B