Evaluation method and equipment for security of large-model long-text task

By building a multi-agent detector, the security performance of large models in long text tasks is systematically evaluated, which solves the problem that the existing technology cannot effectively evaluate the security of large models in long text tasks, and realizes a comprehensive and effective evaluation of the security of large models in long text tasks.

CN120106052APending Publication Date: 2025-06-06BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510187045.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing evaluation benchmarks cannot effectively evaluate the security performance of large models in long text tasks and cannot detect hidden security risks in long text scenarios.

Method used

Systematically evaluate the security performance of large models in long text tasks by building multi-agent detectors, including risk analysts, context summarizers, and security referees. This method collects long text contexts, constructs test instructions for testing security issues, enters the large model to be evaluated and obtains generated responses, and finally performs security evaluation through a multi-agent detector.

Benefits of technology

It realizes a comprehensive and effective assessment of the security of large models in long text tasks, and can detect hidden security risks in complex and real scenarios, improving the reliability of large models in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106052A_ABST
    Figure CN120106052A_ABST
Patent Text Reader

Abstract

The invention relates to an assessment method and device for large-model long-text task security, and belongs to the technical field of artificial intelligence. The method comprises the steps of collecting a long text context; constructing a test instruction for testing the security problem according to the content of the long text context; inputting the test instruction and the long text context into the to-be-evaluated large model, and obtaining a generated reply generated by the to-be-evaluated large model; and inputting the long text context, the test instruction and the generated reply into a multi-agent detector to obtain a security evaluation result of the to-be-evaluated large model. According to the method, the security performance of a large model in a more complex and real scene can be evaluated for a long text task, and the blank of the existing evaluation benchmark is filled. Through division of labor and cooperation of a risk analyst, a context summary staff and a safety referee, misleading or neglecting of hidden unsafe content generated by the large model is avoided, the comprehensiveness and accuracy of the evaluation process are ensured, and the reliability of the large model in practical application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a method and device for evaluating the security of large-model long text tasks. Background Art

[0002] With the continuous development of long text processing technology, large models (LLMs) are becoming more and more capable of understanding and generating long texts. However, the introduction of long texts has led to some new security issues in these models, such as introducing additional harmful content or causing cognitive burden on the model. These security issues may cause LLMs to produce unsafe outputs in practical applications, causing serious security risks. Therefore, it is of great significance to evaluate the security performance of LLMs on various long text tasks.

[0003] However, existing evaluation benchmarks cannot be well applied to the safety performance evaluation of LLM in long text tasks. Existing long text task evaluation benchmarks mainly focus on the general performance of LLM, such as the accuracy of question answers, whether the replies are helpful to the questions, etc., while ignoring the safety performance of LLM. The current safety evaluation benchmarks for LLM are mainly composed of shorter questions, which are often less than a few hundred words in length and do not contain a long context. Therefore, it is difficult to conduct a safety evaluation of LLM in long text scenarios. Although the recently proposed LongSafetyBench conducted a preliminary evaluation of the long text safety performance of LLM, the benchmark is evaluated in the form of multiple-choice questions, which mainly examines the ability of LLM to identify harmful content, and does not evaluate the safety of LLM-generated content, which deviates from the current actual application scenarios of LLM, which is mainly based on generation.

[0004] Therefore, how to achieve a comprehensive and effective assessment of the security of large-model long text tasks is a topic worthy of study for technical personnel in this field. Summary of the invention

[0005] In view of the above analysis, an embodiment of the present invention aims to provide a method and device for evaluating the security of large-model long text tasks, aiming to comprehensively and effectively evaluate the security of content generated by LLM when performing long text tasks.

[0006] In a first aspect of the present application, a method for evaluating the security of a large model long text task is provided, comprising:

[0007] Gathering long text context;

[0008] Constructing a test instruction for testing security issues according to the content of the long text context;

[0009] Inputting the test instruction and the long text context into the large model to be evaluated, and obtaining a generated response generated by the large model to be evaluated;

[0010] Inputting the long text context, the test instruction and the generated response into a multi-agent detector to obtain a security evaluation result of the large model to be evaluated;

[0011] Wherein, the multi-agent detector includes: a risk analyst, a context summarizer, and a safety referee;

[0012] The risk analyst is used to evaluate whether the test instruction has potential security risks, and list possible safe and unsafe generated responses for the test instruction;

[0013] The context summarizer is used to extract key information from the long text context;

[0014] The safety referee is used to analyze the input test instructions and generated responses in combination with the information of the risk analyst and the context summarizer to obtain a safety assessment result of whether the generated responses are safe.

[0015] Optionally, constructing a test instruction for testing security issues according to the content of the long text context includes:

[0016] Determine a security scenario based on the content of the long text context, and extract security keywords from the long text context;

[0017] Based on the security scenario, construct a plurality of test instructions respectively belonging to different task types;

[0018] From multiple test instructions belonging to different task types, the one with the highest probability of inducing the large model to generate an unsafe response is selected as the final test instruction.

[0019] Optionally, after the final test instruction, the following is further included:

[0020] A consistency check is performed on the test instruction, the long text context, and the security keyword to ensure that the test instruction meets the security assessment requirements of the long text task.

[0021] Optionally, collecting the long text context includes:

[0022] For multiple preset security scenarios, expand and generate security keywords corresponding to each security scenario;

[0023] Using the security keywords, collect long documents related to the security scenario;

[0024] A plain text portion of the long document is extracted as a long text context.

[0025] Optionally, after extracting the plain text portion of the long document as the long text context, the method further includes:

[0026] Merge multiple long text contexts with the same security scenario from different sources to generate one long text context.

[0027] Optionally, the risk analyst is used to evaluate whether the test instruction has potential security risks, and list possible safe generation responses and unsafe generation responses of the test instruction, including:

[0028] Calling a predefined security keyword thesaurus, matching the text in the test instruction with the security keywords in the security keyword thesaurus, and detecting whether the test instruction has a potential security risk;

[0029] Using a pre-trained text classification model, perform multi-label classification on the test instruction to determine whether it belongs to an unsafe category;

[0030] By prompting the engineering to receive the large model simulation answer, possible safe generation responses and unsafe generation responses of the test instruction are listed to obtain instruction risk information.

[0031] Optionally, the implementation process of the context summarizer for extracting key information from the long text context includes:

[0032] Use a pre-trained text summarization model to extract text summaries from the context of long texts;

[0033] Extracting text keywords from the long text context;

[0034] The text summary and the text keywords are used as key information.

[0035] Optionally, the safety referee is used to analyze the input test instructions and generated responses in combination with the information of the risk analyst and the context summarizer, and the implementation process of obtaining the safety evaluation result of whether the generated response is safe includes:

[0036] Inputs instruction risk information from risk analysts, key information from context summarizers, and generates responses into the safety referee;

[0037] A rule matching method is used to determine whether the generated response contains any illegal content, and a violation detection evaluation is obtained;

[0038] Use a large model to evaluate the logical consistency between the generated response and the long text context, and obtain a logical consistency detection evaluation;

[0039] Use a large model to evaluate whether the generated response contains unsafe content and obtain a safety detection evaluation;

[0040] Based on the violation detection evaluation, the logic consistency detection evaluation, the security detection evaluation, and the instruction risk information, a security evaluation result of whether the generated response is safe is obtained.

[0041] In a second aspect of the present application, a device for evaluating the security of a large model long text task is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, any one of the above-described methods for evaluating the security of a large model long text task is implemented.

[0042] According to a third aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for evaluating the security of a large model long text task according to any one of the above-mentioned methods is implemented.

[0043] The evaluation method for the security of large-model long-text tasks provided in this application systematically evaluates the security performance of large models in long-text tasks by constructing a multi-agent detector. Compared with the existing large-model security evaluation methods that mainly target short-text tasks and cannot fully detect the hidden security risks in long-text tasks, this application targets long-text tasks and can evaluate the security performance of large models in more complex and realistic scenarios, filling the gap in existing evaluation benchmarks. Through the division of labor and cooperation among risk analysts, context summarizers and security referees, risk analysts evaluate the security risks of instructions in advance, context summarizers ensure that the evaluation is based on key information, and security referees make final judgments based on information from multiple parties to avoid misleading or hidden unsafe content generated by large models from being ignored, ensuring the comprehensiveness and accuracy of the evaluation process, and improving the reliability of large models in practical applications. In addition, this application also provides an evaluation device for the security of large-model long-text tasks with the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0045] Figure 1 A flowchart of a specific implementation method of the method for evaluating the security of large model long text tasks provided in this application;

[0046] Figure 2This is a schematic diagram of the specific implementation process of the risk analyst in this application;

[0047] Figure 3 A schematic diagram of the specific implementation process of the context summarizer for this application;

[0048] Figure 4 This is a schematic diagram of the specific implementation process of the safety referee application;

[0049] Figure 5 A schematic diagram of the overall framework of the evaluation method for the security of large-model long text tasks provided in this application;

[0050] Figure 6 This is a structural block diagram of the evaluation device for the security of large model long text tasks provided in this application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. It should be noted that, in the absence of conflict, the embodiments in the present disclosure and the features in the embodiments can be combined, separated, interchanged and / or rearranged with each other. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0052] The terms used here are for the purpose of describing specific embodiments, and are not intended to be restrictive. As used here, unless the context clearly indicates otherwise, the singular forms "one (kind, person)" and "said (the)" are also intended to include plural forms. In addition, when the terms "comprise" and / or "include" and their variations are used in this specification, it is explained that there are stated features, integral bodies, steps, operations, parts, assemblies and / or their groups, but it is not excluded that there are or add one or more other features, integral bodies, steps, operations, parts, assemblies and / or their groups. It should also be noted that, as used here, the terms "substantially", "approximately" and other similar terms are used as approximate terms and not as degree terms, so that they are used to explain the inherent deviations of the measured values, calculated values ​​and / or the values ​​provided that will be recognized by those of ordinary skill in the art.

[0053] The flowchart of a specific implementation method of the security assessment method for large model long text tasks provided in this application is as follows: Figure 1 As shown, the method specifically includes:

[0054] S101: Collect long text context.

[0055] This step can be done by searching the Internet or obtaining real long text contexts from open databases to ensure the authenticity and applicability of the benchmark. The collected long text contexts can cover different types of security scenarios and target multiple different task types, which are not limited here.

[0056] As a specific implementation method, the safety scenarios may specifically include but are not limited to: harmful content, prejudice and discrimination, physical and mental harm, illegal activities, moral violations, privacy and property, and sensitive topics. Task types specifically include but are not limited to: question-and-answer tasks, generation tasks, brainstorming, summary tasks, rewriting tasks, and role-playing.

[0057] Considering that the existing security-related long text context data is small, it is necessary to collect long text context data from scratch. This application expands and generates security keywords corresponding to each security scenario for multiple different preset security scenarios; uses the security keywords to collect long documents related to the security scenario; and extracts the plain text part of the long document as the long text context.

[0058] After extracting the plain text portion of the long document as the long text context, the method further includes: merging a plurality of long text contexts from different sources with the same security scenario to generate one long text context, so as to cover the case of multi-document tasks.

[0059] S102: Constructing a test instruction for testing security issues according to the content of the long text context.

[0060] Specifically, according to the content of the long text context, a security scenario is determined, and security keywords of the long text context are extracted; based on the security scenario, a plurality of test instructions belonging to different task types are constructed.

[0061] After the long text context is collected in the above step S101, this step is used to construct a test instruction for each long text context.

[0062] The test instructions constructed as above further include a screening step: from a plurality of test instructions belonging to different task types, the one with the highest probability of inducing the large model to generate an unsafe response is screened out as the final test instruction.

[0063] After the final test instruction, the method further includes: performing a consistency check on the test instruction, the long text context and the security keyword, and filtering out inconsistent data among the three, so that the test instruction meets the security assessment requirements of the long text task.

[0064] After the screening is completed, the final long text safety assessment data set, also known as LongSafety, can be obtained.

[0065] S103: Input the test instruction and the long text context into the large model to be evaluated, and obtain a generated response generated by the large model to be evaluated.

[0066] Specifically, GPT-4, Claude, and LLaMA large models can be used to generate responses.

[0067] S104: Input the long text context, the test instruction and the generated response into the multi-agent detector to obtain the security evaluation result of the large model to be evaluated.

[0068] The multi-agent detector includes: risk analysts, context summarizers, and safety referees. The three work together to split and analyze the input from different angles, and finally determine whether the model response is safe.

[0069] The risk analyst is used to evaluate whether the test instruction has potential security risks and list possible safe and unsafe generated responses for the test instruction. The obtained information is unified and integrated into a reference analysis of input risks and provided to the security referee to assist him in making judgments.

[0070] The context summarizer is used to extract key information from the long text context. The context summarizer summarizes the core content of the long context, removes irrelevant interference content, and extracts information related to the instruction to ensure that the safety referee can obtain the most critical information in the long text.

[0071] The safety referee is used to analyze the input test instructions and generated responses in combination with the information of the risk analyst and the context summarizer to obtain a safety assessment result of whether the generated response is safe. After integrating the information provided by the risk analyst and the context summarizer, the safety referee conducts a comprehensive analysis of the input instructions and responses, and makes a final judgment on whether the responses are safe, thereby completing the entire safety testing process.

[0072] This framework makes full use of the cooperation between the three agents and performs efficient and accurate analysis on each part of the input. Compared with existing security detectors, it can better utilize the information contained in long texts, thereby improving the accuracy of the final judgment.

[0073] Reference Figure 2 As shown, the implementation process for the risk analyst to evaluate whether the test instruction has potential security risks and list possible safe generation responses and unsafe generation responses of the test instruction specifically includes:

[0074] S201: calling a predefined security keyword thesaurus, matching the text in the test instruction with the security keywords in the security keyword thesaurus, and detecting whether the test instruction has a potential security risk.

[0075] The security keyword thesaurus can be constructed based on manual annotation and automatic expansion. Use the predefined security keyword thesaurus to quickly screen whether the test instructions contain potential security risks (such as illegality, privacy leakage, discriminatory content, etc.)

[0076] S202: Use a pre-trained text classification model to perform multi-label classification on the test instruction to determine whether it belongs to an unsafe category.

[0077] A large-scale security instruction dataset can be used to pre-train a text classification model to classify and predict the input instructions and output the security category of the test instruction.

[0078] S203: By prompting the engineering to receive the large model simulation answer, possible safe generation responses and unsafe generation responses of the test instruction are listed to obtain instruction risk information.

[0079] Enter the above test instructions into the LLM big model, let the big model simulate its possible answers, predict possible safe and unsafe responses, and judge potential risks to obtain instruction risk information. If the unsafe response contains sensitive information, it means that there is a risk in the response generated by the big model.

[0080] Reference Figure 3 As shown, the implementation process of the context summarizer for extracting key information from the long text context specifically includes:

[0081] S301: Use a pre-trained text summary model to extract a text summary of the long text context.

[0082] Long text contexts are usually long (5,424 words on average in the LongSafety dataset). Using the original text directly for safety evaluation will increase computational complexity and may contain a lot of irrelevant information. The pre-trained text summary model can extract the core content from it, helping safety referees focus on the information most relevant to the test instructions.

[0083] S302: Extracting text keywords from the long text context.

[0084] The summary alone may miss some key information, so this step requires further extraction of text keywords to help security referees quickly understand the core content of the text.

[0085] S303: Using the text summary and the text keywords as key information.

[0086] Reference Figure 4 As shown, the implementation process of the safety referee for analyzing the input test instructions and generated responses in combination with the information of the risk analyst and the context summarizer to obtain the safety evaluation result of whether the generated responses are safe specifically includes:

[0087] S401: Input the instruction risk information from the risk analyst, the key information from the context summarizer, and the generated response into the security referee.

[0088] S402: Using a rule matching method to determine whether the generated reply includes any illegal content, and obtaining a violation detection evaluation.

[0089] Specifically, keyword matching can be used to detect whether sensitive words are included, regular matching or NLP rules can be used to identify illegal structures, and BERT / Transformer classification models can be used to detect illegal categories. This step provides violation detection evaluation by quickly identifying high-risk responses.

[0090] S403: Use the large model to evaluate the logical consistency of the generated response and the long text context to obtain a logical consistency detection evaluation.

[0091] Use the large model to evaluate the logical consistency of the generated response with the long text context, so as to detect whether the generated response of the large model is consistent with the context of the long text context and avoid misleading or off-topic answers.

[0092] S404: Use the large model to evaluate whether the generated response includes unsafe content, and obtain a safety detection evaluation.

[0093] Use the reasoning power of large models to self-assess whether generated responses have potential security risks.

[0094] S405: Based on the violation detection evaluation, the logic consistency detection evaluation, the security detection evaluation, and the instruction risk information, a security evaluation result of whether the generated response is safe is obtained.

[0095] Figure 5 The figure shows the overall framework of the evaluation method for the safety of large-model long text tasks provided by the present application. In this specific example, LongSafety includes 7 types of safety scenarios or safety issues and 6 types of tasks, with a total of 1,543 test cases, and the average length of the test cases is 5,424 words.

[0096] Among them, the seven types of security scenarios or security issues are: harmful content, prejudice and discrimination, physical and mental harm, illegal activities, moral violations, privacy and property, and sensitive topics. The six types of tasks are: question-and-answer tasks, generation tasks, brainstorming, summary tasks, rewriting tasks, and role-playing. By testing multiple task types, LongSafety is ensured to be suitable for different application scenarios. By classifying security issues, different types of security risks of large models can be evaluated in a targeted manner.

[0097] The overall data collection process includes three steps: context collection, instruction construction, and data screening.

[0098] The context collection process includes: first, further refinement based on the above 7 security scenarios, and expansion in each security scenario. The expansion process can be done through manual expert annotation and combined with NLP keyword expansion to generate corresponding security keywords. Then, the data annotator uses the given security keywords to crawl relevant data through Internet search engines or open data sets, and uses the text classification model to filter out irrelevant documents and find relevant long documents. Select the long documents whose content is consistent with the security keywords, and extract the plain text part of the long documents as the long text context.

[0099] Instruction construction: After completing context collection, data annotators can construct three test instructions. The test instructions are consistent with the security scenarios reflected by the long text context and its related security keywords, and belong to different task types to increase the diversity of instructions. At the same time, these instructions need to be able to induce certain security risks to test whether the model will produce unsafe responses. Before constructing test instructions, data annotators need to read the corresponding long text context and check the security scenarios therein to ensure that the test instructions constructed subsequently are consistent with the long text context in the security scenario.

[0100] Data screening: After the data annotator completes the data collection in the first two steps, the obtained data is screened. For the three test instructions constructed in the previous step, the one that is most likely to induce unsafe responses can be retained, and the remaining two can be discarded. After that, the consistency between the context, safety instructions, and safety keywords can be further checked, and inconsistent data between the three can be screened out. After the data screening is completed, the final long text safety evaluation data set LongSafety can be obtained.

[0101] The multi-agent evaluation framework is a core component of LongSafety, which is used to automatically evaluate the safety of responses generated by large models.

[0102] Send the test instructions and long text context to the large model to be evaluated to obtain the generated response.

[0103] Risk analysts assess the potential risks of test instructions (such as whether they induce illegal behavior); predict possible safe and unsafe responses to prevent model errors from being generated.

[0104] Contextual summarizers extract key information from long texts and remove irrelevant content; ensuring that security assessments are based on complete context to prevent misjudgments.

[0105] The safety referee integrates the evaluation results of the risk analyst and the context summarizer to ultimately determine whether the response is safe and generate a safety assessment report, including but not limited to safety scores, unsafe categories, and improvement suggestions.

[0106] This application proposes an evaluation benchmark for comprehensive security evaluation of LLM on long text generation tasks, covering a variety of security issues and long text task types, and can fully detect the security risks of LLM on long text tasks. At the same time, this application also proposes a multi-agent security detector framework for security detection of the model's replies in long text scenarios, which has a higher detection accuracy in long text scenarios than existing detectors.

[0107] Compared with existing technologies, LongSafety has longer test cases, with an average text length of 5,424 words per test case, which is suitable for safety assessment of long text tasks. In addition, LongSafety also has a complete safety classification system and a variety of instruction forms to ensure that it can cover a wide range of task types, thereby ensuring the comprehensiveness of the assessment.

[0108] In addition, this application also provides an evaluation device for the security of large model long text tasks, such as Figure 6 As shown in the structural block diagram of the large model long text task security assessment device provided in the present application, the device includes a memory 61 and a processor 62, and the memory 61 stores a computer program. When the computer program is executed by the processor 62, it implements the large model long text task security assessment method described in any one of the above.

[0109] In addition, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for evaluating the security of large model long text tasks according to any of the above-mentioned methods is implemented.

[0110] Computer readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0111] The professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0112] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0113] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for evaluating the security of large-model long text tasks, characterized in that: include: Gathering long text context; Constructing a test instruction for testing security issues according to the content of the long text context; Inputting the test instruction and the long text context into the large model to be evaluated, and obtaining a generated response generated by the large model to be evaluated; Inputting the long text context, the test instruction and the generated response into a multi-agent detector to obtain a security evaluation result of the large model to be evaluated; Wherein, the multi-agent detector includes: a risk analyst, a context summarizer, and a safety referee; The risk analyst is used to evaluate whether the test instruction has potential security risks, and list possible safe and unsafe generated responses for the test instruction; The context summarizer is used to extract key information from the long text context; The safety referee is used to analyze the input test instructions and generated responses in combination with the information of the risk analyst and the context summarizer to obtain a safety assessment result of whether the generated responses are safe.

2. The method for evaluating the security of large model long text tasks according to claim 1 is characterized in that: The constructing of a test instruction for testing security issues according to the content of the long text context comprises: Determine a security scenario based on the content of the long text context, and extract security keywords from the long text context; Based on the security scenario, construct a plurality of test instructions respectively belonging to different task types; From multiple test instructions belonging to different task types, the one with the highest probability of inducing the large model to generate an unsafe response is selected as the final test instruction.

3. The method for evaluating the security of large model long text tasks according to claim 2 is characterized in that: Following the final test instruction, there is also: A consistency check is performed on the test instruction, the long text context, and the security keyword to ensure that the test instruction meets the security assessment requirements of the long text task.

4. The method for evaluating the security of large model long text tasks according to claim 1 is characterized in that: The collecting of long text context includes: For multiple preset security scenarios, expand and generate security keywords corresponding to each security scenario; Using the security keywords, collect long documents related to the security scenario; A plain text portion of the long document is extracted as a long text context.

5. The method for evaluating the security of large model long text tasks according to claim 4 is characterized in that: After extracting the plain text portion of the long document as the long text context, the method further includes: Merge multiple long text contexts with the same security scenario from different sources to generate one long text context.

6. The method for evaluating the security of a large model long text task according to any one of claims 1 to 5, characterized in that: The risk analyst is used to evaluate whether the test instruction has potential security risks, and list possible safe generation responses and unsafe generation responses of the test instruction. The implementation process includes: Calling a predefined security keyword thesaurus, matching the text in the test instruction with the security keywords in the security keyword thesaurus, and detecting whether the test instruction has a potential security risk; Using a pre-trained text classification model, perform multi-label classification on the test instruction to determine whether it belongs to an unsafe category; By prompting the engineering to receive the large model simulation answer, possible safe generation responses and unsafe generation responses of the test instruction are listed to obtain instruction risk information.

7. The method for evaluating the security of a large model long text task according to any one of claims 1 to 5, characterized in that: The implementation process of the context summarizer for extracting key information from the long text context includes: Use a pre-trained text summarization model to extract text summaries from the context of long texts; Extracting text keywords from the long text context; The text summary and the text keywords are used as key information.

8. The method for evaluating the security of a large model long text task according to any one of claims 1 to 5, characterized in that: The safety referee is used to analyze the input test instructions and generated responses in combination with the information of the risk analyst and the context summarizer to obtain a safety evaluation result of whether the generated response is safe, including: Inputs instruction risk information from risk analysts, key information from context summarizers, and generates responses into the safety referee; A rule matching method is used to determine whether the generated response contains any illegal content, and a violation detection evaluation is obtained; Use a large model to evaluate the logical consistency between the generated response and the long text context, and obtain a logical consistency detection evaluation; Use a large model to evaluate whether the generated response contains unsafe content and obtain a safety detection evaluation; Based on the violation detection evaluation, the logic consistency detection evaluation, the security detection evaluation, and the instruction risk information, a security evaluation result of whether the generated response is safe is obtained.

9. A device for evaluating the security of large-model long text tasks, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method for evaluating the security of a large model long text task according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the method for evaluating the security of a large model long text task according to any one of claims 1 to 8 is implemented.