A risk detection method, device, equipment, medium, and product for a language model
By performing intention and manual separation detection of the input data of the language model, combining output data analysis, differentiated detectors and rule bases are built, the detection problem of Prompt injection risks in the language model is solved, and safety and adaptability are improved.
Patent Information
- Application Number
- CN202410841701.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-06-26
AI Technical Summary
The existing language models have the risk of Prompt injection during application, which is difficult to effectively detect and defend, resulting in security problems.
By separating the intent description content and the intent description content from the model input data, conducting intention risk detection and technical risk detection respectively, combining the model output data for comprehensive analysis, identifying potential malicious intentions and attack methods, and building a differentiated detector and rule base for risk assessment.
It improves the security of the language model, can effectively detect existing and new types of attack requests, has good scalability, generalization and timeliness, and adapts to the risk detection needs of data of different sources in different scenarios.
Smart Images

Figure CN118656822B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information security technology, and in particular, to a method, device, equipment, medium, and product for risk detection of a language model. Background Art
[0002] With the development of language model technology, the application scenarios of language models are increasing, especially the large language model (LLM) is more and more widely used. For example, language models can be applied to dialogue scenarios or video generation scenarios, etc.
[0003] In addition, in some application processes, it is found that the model may have some risks, such as Prompt injection risk, etc., making the security defense requirements for the model urgent. Summary of the Invention
[0004] This application provides a method, device, equipment, medium, and product for risk detection of a language model, which can improve security.
[0005] To achieve the above object, the technical solutions provided in this application are as follows:
[0006] This application provides a method for risk detection of a language model, and the method includes:
[0007] Obtain the model input data of the target language model;
[0008] Determine at least one of the intention description content and the technique description content from the model input data;
[0009] Perform intention risk detection processing on the determined intention description content to obtain an intention risk detection result, and / or perform technique risk detection processing on the determined technique description content to obtain a technique risk detection result;
[0010] Determine the risk detection result of the target language model applied to the model input data according to at least one of the obtained intention risk detection result and the technique risk detection result.
[0011] In a possible implementation manner, the model input data includes at least one prompt;
[0012] The determining at least one of the intention description content and the technique description content from the model input data includes:
[0013] Determine at least one of the intention description content and the technique description content from the at least one prompt according to the different source types of each of the at least one prompt.
[0014] In a possible implementation, the source types of the respective prompt words in the at least one prompt word include: system prompt words and user prompt words; the intent description content is determined based on the system prompt words and user prompt words in the model input data; the technique description content is determined based on the user prompt words in the model input data.
[0015] In a possible implementation, the method further includes:
[0016] Perform abnormal response detection processing on the response description information to obtain an abnormal response detection result; the response description information is determined based on the model input data and / or the model output data of the target language model; the model output data is obtained by the target language model processing the model input data;
[0017] The risk detection result of the target language model applied to the model input data is also determined based on the abnormal response detection result.
[0018] In a possible implementation, the model input data includes multiple data, and the source types of different data are different; the intent description content and the technique description content are determined based on the source types of the multiple data.
[0019] In a possible implementation, the multiple data includes a plugin response and at least one prompt word, the source type of the plugin response is different from the source types of the respective prompt words, and the source types of different prompt words are different; the intent description content and the technique description content are determined based on the at least one prompt word and the source types of the at least one prompt word; the response description information is determined based on the plugin response and / or the model output data.
[0020] In a possible implementation, the at least one prompt word includes system prompt words; the response description information is also determined based on the system prompt words.
[0021] In a possible implementation, the method further includes:
[0022] Obtain model detection constraint information;
[0023] Determine a detection execution device that matches the model detection constraint information from the detection execution devices corresponding to multiple candidate constraint information;
[0024] The matching detection execution device is used to perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result, and / or perform technique risk detection processing on the determined technique description content to obtain a technique risk detection result.
[0025] In a possible implementation, the model detection constraint information includes at least one of scenario constraint information, regional constraint information, service constraint information, language constraint information, and detection item constraint information.
[0026] In a possible implementation, the matching detection execution device includes a detector corresponding to the intent description content and a detector corresponding to the method description content; the detector corresponding to the intent description content is configured to perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result; the detector corresponding to the method description content is configured to perform method risk detection processing on the determined method description content to obtain a method risk detection result.
[0027] In a possible implementation, the intent description content includes intent content of at least two source types; the detector corresponding to the intent description content includes detectors corresponding to intent content of various source types; for any detector corresponding to the intent content of a source type, the detector is configured to perform intent risk detection processing on the intent content of that source type.
[0028] In a possible implementation, after determining the detection execution device that matches the model detection constraint information from the detection execution devices corresponding to multiple candidate constraint information, the method further includes:
[0029] Sending a detection request to the matching detection execution device; the matching detection execution device is configured to perform intent risk detection processing on the intent description content carried in the detection request to obtain an intent risk detection result, and / or perform method risk detection processing on the method description content carried in the detection request to obtain a method risk detection result;
[0030] Receiving feedback information from the matching detection execution device, where the feedback information includes the intent risk detection result and / or the method risk detection result.
[0031] In a possible implementation, the matching detection execution device is further configured to perform abnormal response detection processing on the response description information to obtain an abnormal response detection result; the response description information is determined based on the model input data and / or the model output data of the target language model; the model output data is obtained by the target language model processing the model input data.
[0032] In a possible implementation, the detection execution devices corresponding to the multiple candidate constraint information are constructed based on a pre-constructed detection set, and the detection set includes one or more of at least one detection rule, at least one detection vector, and at least one detection model;
[0033] After determining the risk detection result of the target language model applied to the model input data, the method further includes:
[0034] If the risk detection result of the target language model applied to the model input data indicates a risk, then update the detection set according to the model input data.
[0035] This application provides a risk detection device for a language model, including:
[0036] A data acquisition unit, configured to acquire model input data of a target language model;
[0037] A content determination unit, configured to determine at least one of intent description content and technique description content from the model input data;
[0038] A data detection unit, configured to perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result, and / or perform technique risk detection processing on the determined technique description content to obtain a technique risk detection result;
[0039] A result determination unit, configured to determine the risk detection result of the target language model applied to the model input data according to at least one of the obtained intent risk detection result and technique risk detection result.
[0040] This application provides an electronic device, and the device includes: a processor and a memory;
[0041] The memory is configured to store instructions or computer programs;
[0042] The processor is configured to execute the instructions or computer programs in the memory so that the electronic device executes the risk detection method of the language model provided by this application.
[0043] This application provides a computer-readable medium, and instructions or computer programs are stored in the computer-readable medium. When the instructions or computer programs run on a device, the device is caused to execute the risk detection method of the language model provided by this application.
[0044] This application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the risk detection method of the language model provided by this application.
[0045] Compared with the related technologies, the present application has at least the following advantages:
[0046] In the technical solution provided by the present application, after obtaining the model input data of the target language model, at least one of the intent description content and the technique description content is determined from the model input data first; then, intent risk detection processing is performed on the intent description content to obtain an intent risk detection result, and / or, technique risk detection processing is performed on the technique description content to obtain a technique risk detection result; finally, based on the intent risk detection result and / or the technique risk detection result, the risk detection result of the target language model under the model input data is determined, so that the risk detection result can indicate whether there is a risk, such as a Prompt injection risk, etc., thereby effectively avoiding the security problems caused by the risk, and further being beneficial to improving security.
[0047] It can be seen that since the present application performs risk detection processing on the model input data from the two dimensions of intent and technique, so that the present application can analyze whether there is a Prompt injection risk from the two features of malicious intent and attack technique respectively, thereby enabling the present application to not only detect the pre-learned attack samples, but also detect the new attack requests composed of the malicious intent and attack techniques that appear in these attack samples. Furthermore, the number of samples required to be learned by the present application does not increase exponentially with the increase in the number of risk types. In this way, the present application has good scalability, generalization, and timeliness. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or the related technologies. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0049] Figure 1 It is a flowchart of a method for detecting risks of a language model provided by an embodiment of the present application;
[0050] Figure 2 It is a schematic diagram of a detection processing logic provided by an embodiment of the present application;
[0051] Figure 3 It is a schematic diagram of a detection processing framework provided by an embodiment of the present application;
[0052] Figure 4 It is a schematic diagram of a detection processing flow provided by an embodiment of the present application;
[0053] Figure 5 A schematic diagram of vector recall provided by an embodiment of the present application;
[0054] Figure 6 A schematic diagram of a classification process provided by an embodiment of the present application;
[0055] Figure 7 Another schematic diagram of a classification process provided by an embodiment of the present application;
[0056] Figure 8 A schematic diagram of the structure of a risk detection device for a language model provided by an embodiment of the present application;
[0057] Figure 9 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0058] Next, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0059] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.
[0060] In this article, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately executed after "A", but may include one or more intermediate steps.
[0061] It can be understood that the data involved in the technical solution of the present application (including but not limited to the data itself, the acquisition, use, storage, or deletion of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0062] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to the relevant users in an appropriate manner and the authorization of the relevant users should be obtained according to the relevant laws and regulations, where the relevant users may include any type of right subject, such as individuals, enterprises, and groups.
[0063] For example, when receiving an active request from a user, a prompt message is sent to the relevant user to clearly prompt the relevant user that the operation requested by the user will require obtaining and using the information of the relevant user, so that the relevant user can autonomously choose whether to provide information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0064] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the relevant user in response to receiving an active request from the relevant user may be, for example, a pop-up window manner, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide information to the electronic device.
[0065] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0066] The following begins a detailed description of the solution of the present disclosure.
[0067] Through research by the inventors of the present disclosure, it is found that a security defense direction for the model is model detection and response, and the "model detection and response" specifically is: comprehensively determining whether there is a risk operation by detecting the input and output of the model, the internal calculation process of the model, and the actions of the downstream modules of the model, and intercepting or filtering the risk operation. It can be seen that for the security defense direction of "model detection and response", it can identify unexpected requests by detecting the input and output of the model. Based on this, a security defense idea for the model is: inserting a detection and filtering process in the input and output links of the model to ensure the security of the content entering the model and the content output by the model, and further protecting against risks such as Prompt injection risks.
[0068] The inventors of the present disclosure also found that for some related solutions for detecting model input and output, these solutions can combine a rule engine, a small-parameter model similar to Natural Language Processing (NLP), vector retrieval, etc., to directly perform risk detection processing on the overall input and output of the model to address risks such as direct injection attacks and language logic attacks. Among them, the rule engine refers to using a rule engine similar to Yara to determine whether the input or output hits a predefined rule, usually in regular form. The small-parameter model refers to implementing risk classification for the model input or output by fine-tuning a small-parameter model similar to Bert. The vector retrieval refers to judging whether the model input or output hits similar attack samples by vectorizing and storing known attack samples in the database.
[0069] The inventors of the present disclosure further found that the related solutions shown in the above paragraph have the following problems: Since model attacks can be composed of two features, malicious intent and attack techniques, as shown in the following Table 1, the free combination of different features can form new model attacks, resulting in the number of samples required for the related solutions implemented by performing risk detection on the entire Prompt increasing exponentially as the number of risk types increases. This leads to problems such as poor scalability, poor generalization, and poor timeliness of the related solutions.
[0070]
[0071] Table 1 Different Features of Model Attacks
[0072] Based on the above findings, in order to better improve security, the present application provides a risk detection method for a language model, which includes: after obtaining the model input data of the target language model, first determining at least one of the intent description content and the technique description content from the model input data; then, performing intent risk detection processing on the intent description content to obtain an intent risk detection result, and / or performing technique risk detection processing on the technique description content to obtain a technique risk detection result; finally, determining the risk detection result of the target language model under the model input data according to the intent risk detection result and / or the technique risk detection result, so that the risk detection result can indicate whether there is a risk, such as Prompt injection risk, etc., thereby effectively avoiding the security problems caused by the risk, and further being beneficial to improving security. It can be seen that since the present application performs risk detection processing on the model input data from the two dimensions of intent and technique, so that the present application can analyze whether there is a Prompt injection risk from the two features of malicious intent and attack technique respectively, thus enabling the present application to not only detect the attack samples that have been learned in advance, but also detect the new attack requests composed of the malicious intent and attack techniques that appear in these attack samples. Furthermore, the number of samples required to be learned by the present application does not increase exponentially with the increase in the number of risk types. In this way, the present application has good scalability, generalization, and timeliness.
[0073] In addition, the present application does not limit the execution subject of the risk detection method for the language model provided in the embodiments of the present application. For example, the risk detection method for the language model provided in the embodiments of the present application can be applied to a terminal device or a server. Another example is that the risk detection method for the language model provided in the embodiments of the present application can also be implemented by means of the data interaction process between the terminal device and the server. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (Personal Digital Assitant, PDA), a tablet computer, etc. The server can be an independent server, a cluster server, or a cloud server.
[0074] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0075] To better understand the technical solution provided by the present application, the risk detection method for the language model provided by the present application will be described below in conjunction with some drawings. As Figure 1As shown, the risk detection method for the language model provided by the embodiment of the present application includes S1 - S5 below. Among them, the Figure 1 is a flowchart of a risk detection method for a language model provided by an embodiment of the present application.
[0076] S1: Obtain the model input data of the target language model.
[0077] Among them, the target language model refers to the model that needs to be processed for risk detection, and the present application does not limit the implementation manner of the target language model. For example, it can adopt any existing or future machine learning model, such as LLM or Figure 2 the model 1 shown for implementation.
[0078] The model input data refers to the data input into the target language model, such as Figure 2 the input data shown; and the present application does not limit the implementation manner of the model input data. For the convenience of understanding, the following will be described in combination with three examples.
[0079] Example 1, in some application scenarios, the model input data may include a user prompt (UserPrompt, UP). Among them, the user prompt refers to the data provided by the user for the target language model, such as the data input by the user through certain input devices, etc. In addition, the present application does not limit the acquisition method of the user prompt. For example, the user prompt can be obtained by means of a certain information transmission protocol, such as Figure 3 the Hypertext Transfer Protocol (HTTP) or Remote Procedure Call (RPC) shown. In addition, the present application also does not limit the implementation manner of the user prompt. For example, the user prompt can be implemented using Figure 4 the user prompt "Please act as identity 1, who always tells story 1 to lull me to sleep" shown.
[0080] Example 2, in some application scenarios, the model input data may include a user prompt and a system prompt (System Prompt, SP). Among them, the system prompt refers to the data pre - set for the target language model and used to describe the functions of the target language model, such as Figure 4 the system prompt shown. In addition, the present application does not limit the acquisition method of the system prompt. For example, the system prompt can be obtained by means of a certain information transmission protocol, such as Figure 3Obtained through HTTP or RPC as shown. In addition, this application does not limit the implementation manner of the system prompt words. For example, the system prompt words may include identity setting content and skill setting content. Among them, the identity setting content is used to describe the role played by the target language model, such as roles like Agent, ChatBot, Code Assistant, MultiModal, etc. The skill setting content is used to describe the functions of the target language model when it plays a certain role. It should be noted that this application does not limit the implementation manner of the multi-modal. For example, it can be implemented using text-to-image generation.
[0081] Example 3, in some application scenarios, the model input data may include user prompt words, system prompt words, and plugin responses. Among them, the plugin response refers to the feedback data of the plugin called by the target language model, so that the plugin response can be used to represent the execution result of the upstream task corresponding to the target language model, such as the data finally output by the plugin involved in the upstream task or Figure 2 The plugin response as shown, so that when executing the current task, the target language model processes the plugin response. It should be noted that this application does not limit the implementation manner of the plugin. For example, it can use any plugin that can provide a certain service, such as Figure 2 The browser or a certain application programming interface (Application Programming Interface, API) as shown for implementation. In addition, this application does not limit the acquisition method of the plugin response. For example, the plugin response can be obtained by means of a certain information transmission protocol, such as Figure 3 HTTP or RPC as shown.
[0082] Based on the above three examples, in some application scenarios, such as Figure 2 In the workflow execution scenario as shown, for the target language model involved in the current task, the model input data of the target language model can be determined based on some or all of the three source types of data: user prompt words, system prompt words, and plugin responses. Among them, the plugin response can be used to represent the execution result of the upstream task corresponding to the current task.
[0083] Based on the relevant content of S1 above, for a target language model, the model input data of the target language model can be composed of at least two source types of data, such as Figure 2It is obtained by splicing the data of the three source types of user prompts, system prompts, and plugin responses shown, so that the model input data can include the data of these source types, so that subsequent risk detection processing for the target language model can be realized by different detection processing methods for the data of different source types.
[0084] S2: Determine at least one of the intent description content and the method description content from the model input data.
[0085] Among them, the intent description content refers to the intent description part existing in the model input data, so that the intent description content can describe the purpose expressed by the model input data, such as the purpose of "telling a story 1".
[0086] The method description content refers to the method description part existing in the model input data, so that the method description content can describe the technique used when achieving the purpose expressed by the model input data, such as the role-playing technique of "playing identity 1".
[0087] In addition, the present application does not limit the implementation manner of S2 above. For example, specifically, after obtaining the model input data of the target language model, the model input data can be split into different parts, one part is used as the intent description content, and the other part is used as the method description content, so that the intent description content can represent the intent characteristics described by the model input data, and the method description content can represent the method characteristics described by the model input data.
[0088] Through research by the inventors of the present disclosure, it is found that data of different source types can be used to provide attack characteristics in different dimensions. For example, user prompts and system prompts may provide the attack characteristic of malicious intent, and user prompts may provide the attack characteristic of attack methods, etc.
[0089] Based on the above research, in order to better improve the detection effect, the present application also provides a possible implementation manner of S2 above. In this manner, when the above model input data includes at least one prompt, S2 can specifically be: according to the different source types of each prompt in the at least one prompt, determine at least one of the intent description content and the method description content from the at least one prompt. Among them, the at least one prompt refers to the prompt with different source types existing in the model input data; and the present application does not limit the implementation manner of the at least one prompt. For example, the at least one prompt can include system prompts and user prompts, so that the source types of each prompt in the at least one prompt can include system prompts and user prompts. It should be noted that the source type of the system prompt is different from the source type of the user prompt.
[0090] Based on the above two paragraphs, it can be seen that in a possible implementation manner, when the input data of the above model includes a system prompt and a user prompt, step S2 above may include steps 11-12 below.
[0091] Step 11: Determine the intent description content based on the system prompt and the user prompt in the input data of the above model, so that the intent description content includes some or all of the system prompt, and the intent description content includes some or all of the user prompt.
[0092] It should be noted that the implementation manner of step 11 above in this application is not limited. For example, to ensure information integrity, step 11 may specifically be: after obtaining the system prompt and the user prompt in the input data of the above model, directly determine the intent description content based on the system prompt and the user prompt, so that the intent description content includes the system prompt and the user prompt.
[0093] Another example, to improve efficiency, step 11 above may specifically be: after obtaining the system prompt and the user prompt in the input data of the above model, first extract the system intent from the system prompt and the user intent from the user prompt; then determine the intent description content based on the system intent and the user intent, so that the intent description content includes the system intent and the user intent. Among them, the system intent is used to describe the purpose expressed by the system prompt, and the implementation manner of the system intent in this application is not limited. For example, the system intent may include the intent description part existing in the system prompt. The user intent is used to describe the purpose expressed by the user prompt, and the implementation manner of the user intent in this application is not limited. For example, the user intent may include the intent description part existing in the user prompt.
[0094] Based on the relevant content of step 11 above, in some application scenarios, when the input data of the above model includes prompt words of at least two source types, according to the intention determination methods corresponding to different source types, the intention description content is determined from these prompt words, so that the intention description content includes the intention content of the at least two source types, thereby enabling the intention description content to better describe the purpose expressed by the model input data. Among them, the prompt words of the at least two source types are used to represent the prompt words with different source types existing in the model input data. In addition, for any one of the at least two source types, the intention content of the source type is determined from the prompt words of the source type through the intention determination method corresponding to the source type, so that the intention content of the source type can describe the purpose expressed by the prompt words of the source type; moreover, the present application does not limit the implementation manner of the intention content of the source type. For example, the intention content of the source type may include some or all of the content in the prompt words of the source type. In addition, the present application does not limit the implementation manner of the intention determination method corresponding to the source type. For example, the intention determination method corresponding to the source type is determined according to the intention determination requirements corresponding to the source type, so that the intention content determined based on the intention determination method corresponding to the source type meets the intention determination requirements, which is beneficial to improving the intention determination effect and thus beneficial to improving the malicious intention detection effect.
[0095] Step 12: Determine the method description content according to the user prompt words in the input data of the above model, so that the method description content includes some or all of the user prompt words.
[0096] It should be noted that the present application does not limit the implementation manner of step 12 above. For example, in order to ensure information integrity, step 12 may specifically be: after obtaining the user prompt words in the input data of the above model, directly determine the user prompt words as the method description content, so that the method description content includes the user prompt words.
[0097] For another example, in order to improve efficiency, step 12 may specifically be: after obtaining the user prompt words in the input data of the above model, extract the method description part from the user prompt words and use the method description part as the method description content, so that the method description content includes the method description part, thereby enabling the method description content to include some of the user prompt words, and further enabling the time consumed in the subsequent processing of the method description content to be relatively less, which is beneficial to improving efficiency.
[0098] It should be noted that the present application does not limit the correlation between the execution time of step 12 above and the execution time of step 11 above. For example, the former is earlier than the latter. For another example, the former is later than the latter. For still another example, the two are the same.
[0099] Based on the relevant content of the above steps 11 to 12, in a possible implementation, when the above model input data includes a system prompt and a user prompt, the intent description content can be determined based on the system prompt and the user prompt, and the method description content is determined based on the user prompt, so that subsequent corresponding detection processing can be performed on the intent description content and the method description content respectively to determine whether there is a malicious intent and an attack method, so that subsequent whether there is a Prompt injection risk can be comprehensively determined based on these results, so as to effectively avoid the defects caused by detecting the entire model input data, thereby facilitating improving the detection effect.
[0100] Based on the above content, in a possible implementation, when the above model input data includes at least one prompt, and the source types of the prompts in the at least one prompt include a system prompt and a user prompt, the above intent description content can be determined based on the system prompt and the user prompt in the model input data; the above method description content can be determined based on the user prompt in the model input data.
[0101] Based on the relevant content of the above S2, for some application scenarios, such as Figure 2 For the scenario shown, after obtaining the model input data of the target language model, if the model input data is obtained by splicing data of multiple source types, then the content of at least two dimensions of intent description content and method description content can be split from the model input data according to the source types of different parts of the model input data, so that subsequent whether there is a Prompt injection risk of the model input data can be comprehensively determined based on the risk detection results of these two dimensions.
[0102] S3: Perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result.
[0103] Among them, the intent risk detection result is used to indicate whether there is a malicious intent in the intent description content, such as the malicious intent shown in Table 2 below.
[0104] In addition, the present application does not limit the implementation manner of the above intent risk detection processing. For example, the intent risk detection processing can be implemented by using Figures 2 - 4 any one of the malicious intent detections shown. For another example, the intent risk detection processing can be implemented by means of some pre-constructed rules and models, such as Figure 3 the rules and models shown.
[0105]
[0106] Table 2 Examples of Malicious Intentions
[0107] It can be seen that in a possible implementation, in order to better improve the detection effect, S3 above can specifically be: using multiple intent detection means to perform intent risk detection processing on the above-mentioned intent description content to obtain an intent risk detection result. It should be noted that this application does not limit these multiple intent detection means. For example, these multiple intent detection means may include one or more of rule checking, vector recall, and classification methods. Among them, the rule checking is used to determine whether the intent description content hits a predefined rule, such as a regular expression, etc.; and this application does not limit the implementation manner of the rule checking. For example, the rule checking can be implemented using a rule engine or Figure 4 the rule checking 1 shown. The vector recall is used to determine whether the intent description content hits these attack samples based on the similarity between the vectorized features of the intent description content and the vectorized features of some attack samples stored in a pre-constructed vector library; and this application does not limit the implementation manner of the vector recall. For example, the vector recall can be implemented using Figure 4 the vector recall 1 shown or Figure 5 the vector recall shown. The classification method is used to classify the intent description content with the help of some classifiers, such as a small-parameter model or a large-parameter model, such as Figure 4 the classification processing 1 shown, Figure 6 or Figure 7 the classification processing shown, etc. It should be noted that this application does not limit the implementation manner of the small-parameter model. For example, it can be implemented using a Natural Language Processing (NLP) model or Figure 6 the model shown. In addition, this application does not limit the implementation manner of the large-parameter model either. For example, it can be implemented using a FewShot large language model (LLM) or a fine-tuned LLM.
[0108] In addition, this application does not limit the execution condition of S3 above. For example, specifically, it can be: after determining that the intent description content is obtained, execute S3.
[0109] S4: Perform technique risk detection processing on the determined technique description content to obtain a technique risk detection result.
[0110] Among them, the technique risk detection result is used to indicate whether there is content for describing attack techniques in the technique description content, such as the content shown in Table 3 below.
[0111] In addition, this application does not limit the implementation manner of the above-mentioned technique risk detection processing. For example, the technique risk detection processing can be implemented using Figures 2 - 4Implement the attack method detection shown in any one of them. For another example, this method risk detection process can be carried out with the help of some pre - constructed rules and models, such as Figure 3 The rules and models shown are implemented.
[0112] It can be seen that in a possible implementation, in order to better improve the detection effect, S4 above can specifically be: performing a method risk detection process on the method description content above by using a variety of method detection means to obtain a method risk detection result. It should be noted that this application does not limit these various method detection means. For example, these various method detection means can include one or more of rule checking, vector recall, and classification methods. Among them, this rule checking is used to determine whether the method description content hits a predefined rule, such as a regular expression, etc.; and this application does not limit the implementation manner of this rule checking. For example, this rule checking can be implemented by using a rule engine or Figure 4 The rule checking 2 shown is implemented. This vector recall is used to determine whether the method description content hits these attack samples based on the similarity between the vectorized features of the method description content and the vectorized features of some attack samples stored in a pre - constructed vector library; and this application does not limit the implementation manner of this vector recall. For example, this vector recall can be implemented by using Figure 4 The vector recall 2 shown or Figure 5 The vector recall shown is implemented. This classification method is used to perform a classification process on the method description content by using some classifiers, such as a small - parameter model or a large - parameter model, such as Figure 4 The classification process 2 shown, Figure 6 Or Figure 7 The classification process shown, etc.
[0113]
[0114] Table 3 Examples of attack methods
[0115] In addition, this application does not limit the execution conditions of S4 above. For example, specifically, it can be: after determining that the method description content is obtained, execute this S4.
[0116] Furthermore, when the execution conditions of S4 above and the execution conditions of S3 above are simultaneously satisfied, this application does not limit the correlation between the execution time of S4 above and the execution time of S3 above. For example, the former is earlier than the latter. For another example, the former is later than the latter. Also, for example, the two are the same.
[0117] S5: Determine the risk detection result of the target language model applied to the model input data based on at least one of the obtained intent risk detection result and method risk detection result.
[0118] Among them, the risk detection result of the target language model under the model input data is used to indicate whether there is a Prompt injection risk in the model input data.
[0119] In addition, this application does not limit the implementation manner of S5 above. For example, specifically, it can be: after obtaining the intent risk detection result and the technique risk detection result, some summary processing can be performed on these results, such as Figure 2 the combined analysis shown, Figure 3 the aggregation analysis shown, or Figure 4 the aggregation analysis shown, etc. for summary processing to obtain the risk detection result of the target language model under the model input data.
[0120] It can be seen that in a possible implementation manner, S5 above can specifically be: if the above intent risk detection result indicates the existence of a malicious intent or the above technique risk detection result indicates the existence of an attack technique, it can be determined that there is a risk in the model input data. Therefore, a pre-set risk characterization value, such as the value "1", can be determined as the risk detection result of the target language model under the model input data, so that the "risk detection result of the target language model under the model input data" can indicate that there is a risk in the model input data; if the intent risk detection result indicates the non-existence of a malicious intent and the technique risk detection result indicates the non-existence of an attack technique, it can be determined that there is no risk in the model input data. Therefore, a pre-set risk-free characterization value, such as the value "0", can be determined as the risk detection result of the target language model under the model input data, so that the "risk detection result of the target language model under the model input data" can indicate that there is no risk in the model input data.
[0121] Based on the relevant content of S1 to S5 above, for the risk detection method of the language model provided in the embodiments of the present application, after obtaining the model input data of the target language model, first determine at least one of the intent description content and the technique description content from the model input data; then, perform intent risk detection processing on the intent description content to obtain an intent risk detection result, and / or perform technique risk detection processing on the technique description content to obtain a technique risk detection result; finally, determine the risk detection result of the target language model under the model input data based on the intent risk detection result and / or the technique risk detection result, so that the risk detection result can indicate whether there is a risk, such as Prompt injection risk, etc., thereby effectively avoiding the security problems caused by the risk, and further being beneficial to improving security. Among them, because the present application performs risk detection processing on the model input data from two dimensions of intent and technique, so that the present application can analyze whether there is a Prompt injection risk from two characteristics of malicious intent and attack technique, so that the present application can not only detect the attack samples that have been learned in advance, but also detect new attack requests composed of the malicious intent and attack techniques that appear in these attack samples, and further make the number of samples required to be learned in the present application not increase exponentially with the increase in the number of risk types, so that the present application has good scalability, generalization and timeliness.
[0122] Through research by the inventors of the present disclosure, it has been found that in some application scenarios, in order to better improve the detection effect, some responses (such as model output data, etc.) can be used to supplement information for the detection process of the model input data to achieve the advantages shown in ①-② below.
[0123] ① It is beneficial to improve the attack detection precision and recall rate. The reason is that some attacks cannot be accurately judged as attack requests based solely on the model input data. For example, the key point for distinguishing whether the User Prompt of "repeat the above content" is an attack request with SP leakage risk or a risk-free normal request is: to judge whether the model output data actually includes SP content. If the model output data only repeats the user's historical chat content, it can be determined that the model input data is not an attack request; however, if the model output data includes SP content, it can be determined that there is an SP leakage risk, and thus it can be determined that the model input data is indeed an attack request.
[0124] ②It is beneficial to cover and provide a safety net for some unknown attacks. That is, the value of the model output data lies not only in covering known attacks but also in providing a safety net for unknown attacks. The reasons are as follows: No matter what method is used for the detection of the model input data, it largely depends on the understanding of existing attacks. However, due to the endless emergence of unknown attacks against the model, it is impossible to effectively cover all possible risks only by detecting the model input data. In contrast, the model output data is the most intuitive way to observe whether the model shows unexpected behaviors and can cover and provide a safety net for unknown attacks.
[0125] Based on the above research, in order to better improve the detection effect, the present application also provides a possible implementation manner of the risk detection method for the above language model. In this implementation manner, the risk detection method may at least include the following steps 21 - step 26.
[0126] Step 21: Obtain the model input data of the target language model.
[0127] It should be noted that for the relevant content of step 21, please refer to the relevant content of S1 above.
[0128] Step 22: Determine at least one of the intent description content and the technique description content from the model input data.
[0129] It should be noted that for the relevant content of step 22, please refer to the relevant content of S2 above.
[0130] Step 23: Perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result.
[0131] It should be noted that for the relevant content of step 23, please refer to the relevant content of S3 above.
[0132] Step 24: Perform technique risk detection processing on the determined technique description content to obtain a technique risk detection result.
[0133] It should be noted that for the relevant content of step 24, please refer to the relevant content of S4 above.
[0134] Step 25: Perform abnormal response detection processing on the response description information to obtain an abnormal response detection result; the response description information is determined based on the model input data and / or the model output data of the target language model; the model output data is obtained by the target language model processing the model input data.
[0135] Among them, the response description information is used to describe the response involved by the target language model, such as the plug-in response input to the target language model or the response output by the target language model, etc.; moreover, the implementation manner of the response description information is not limited in this application. For the sake of easy understanding, the following will be described in combination with three cases.
[0136] Case 1, in some application scenarios, if there is no upstream task corresponding to the target language model, it can be determined that the model input data of the target language model does not include the plug-in response. Thus, it can be determined that the response description information corresponding to the target language model can be determined based on the model output data of the target language model, so that the response description information includes the model output data. Among them, the model output data refers to the data obtained by the target language model processing the model input data, such as Figure 2 the output data shown, so that the model output data can represent the response given by the target language model for the model input data, thereby enabling the model output data to provide some supplementary information for the detection process of the model input data to a certain extent.
[0137] Case 2, in some application scenarios, if there is an upstream task corresponding to the target language model, it can be determined that the model input data of the target language model includes the plug-in response. Thus, it can be determined that the response description information corresponding to the target language model can be determined based on the plug-in response and the model output data of the target language model, so that the response description information includes the plug-in response and the model output data, so as to be able to utilize these two responses to assist the detection process of the model input data subsequently.
[0138] Case 3, in some application scenarios, if there is an upstream task corresponding to the target language model, but the model output data of the target language model cannot provide supplementary information for the risk detection process, then in order to better improve the efficiency, the response description information corresponding to the target language model can be determined based on the plug-in response, so that the response description information includes the plug-in response, so as to be able to utilize this one response to assist the detection process of the model input data subsequently.
[0139] Based on the relevant content of the above response description information, in a possible implementation manner, the response description information corresponding to the above target language model can be determined based on the plug-in response and / or the model output data of the target language model, so that the response description information can represent the response required for reference when performing risk detection on the model input data of the target language model.
[0140] In fact, in some application scenarios, such as the system prompt leakage detection scenario, in order to better improve the detection effect, the present application also provides a possible implementation manner of the above response description information. In this manner, when the above model input data includes at least the system prompt, the response description information is also determined based on the system prompt, so that the system prompt can participate in the subsequent abnormal response detection process, which is beneficial to improving the abnormal response detection effect. It should be noted that the present application does not limit this determination process. For example, in some scenarios, the response description information can be determined based on the system prompt and the plugin response, so that subsequent abnormal response detection processing can be performed based on the system prompt and the plugin response. Another example is that in some scenarios, the response description information can be determined based on the system prompt and the model output data, so that subsequent abnormal response detection processing can be performed based on the system prompt and the model output data. Still another example is that in some scenarios, the response description information can be determined based on the system prompt, the plugin response, and the model output data, so that subsequent abnormal response detection processing can be performed based on the system prompt, the plugin response, and the model output data.
[0141] The abnormal response detection result is used to indicate whether there is abnormal content in the response involved by the target language model, as shown in Table 4 below.
[0142]
[0143] Table 4 Examples of Abnormal Responses
[0144] In addition, the abnormal response detection result is obtained by performing abnormal response detection processing on the response description information corresponding to the target language model. It should be noted that the present application does not limit the implementation manner of this abnormal response detection processing. For example, this abnormal response detection processing can be implemented using Figures 2 - 4 any of the abnormal response detections shown. Another example is that this abnormal response detection processing can be implemented with the help of some pre-constructed rules and models, such as Figure 3 the rules and models shown.
[0145] It can be seen that in a possible implementation manner, in order to better improve the detection effect, step 25 above can specifically be: performing abnormal response detection processing on the above response description information using a variety of response detection means to obtain an abnormal response detection result. It should be noted that the present application does not limit these various response detection means. For example, these various response detection means can include one or more of rule checking, vector recall, and classification methods. Among them, the rule checking is used to determine whether the response description information hits a predefined rule, such as a regular expression, etc.; and the present application does not limit the implementation manner of this rule checking. For example, this rule checking can be implemented using a rule engine or Figure 4The rule check 3 shown is implemented. This vector recall is used to determine whether the response description information hits these attack samples based on the similarity between the vectorized features of the response description information and the vectorized features of some attack samples stored in the pre-constructed vector library; moreover, the implementation manner of this vector recall is not limited in this application. For example, this vector recall can be implemented using Figure 4 the vector recall 3 shown or Figure 5 the vector recall shown. This classification method is used to classify the response description information by means of some classifiers, such as a small-parameter model or a large-parameter model, such as Figure 4 the classification process 3 shown, Figure 6 or Figure 7 the classification process shown, etc.
[0146] In addition, the correlation relationship among the execution time of step 25, the execution time of step 24, and the execution time of step 23 is not limited in this application. For example, the three are the same. Another example is that the three satisfy a certain permutation order. Also, after steps 24 and 23 are executed, if the intent risk detection result determined by step 23 indicates no risk and the technique risk detection result determined by step 24 indicates no risk, then step 25 is executed to achieve a fallback, which is beneficial to saving detection time and thus beneficial to improving efficiency.
[0147] Step 26: Determine the risk detection result of the target language model applied to the model input data based on at least two of the above intent risk detection result, the above technique risk detection result, and the above abnormal response detection result.
[0148] It should be noted that the implementation manner of step 26 above is not limited in this application. For example, specifically, after obtaining the intent risk detection result, the technique risk detection result, and the abnormal response detection result, some summary processing can be performed on these results, such as Figure 2 the combined analysis shown, Figure 3 the aggregation analysis shown, or Figure 4 the aggregation analysis shown, etc., to obtain the risk detection result of the target language model applied to the model input data.
[0149] It can be seen that in a possible implementation manner, the above "risk detection result of the target language model applied to the model input data" is determined not only based on at least one of the above intent risk detection result and the above technique risk detection result, but also based on the above abnormal response detection result.
[0150] Actually, in some application scenarios, such as Figure 4In the scenario shown, in order to better improve the detection effect, the fallback can be carried out by means of the abnormal response detection result. Based on this, the present application also provides a possible implementation manner of step 26 above. In this implementation manner, step 26 can specifically be: first, based on the above-mentioned intent risk detection result and the above-mentioned technique risk detection result, determine whether a risk is hit, such as determining whether risks such as malicious intent and attack techniques are hit; if a risk is hit, a pre-set risk characterization value, such as the value "1", can be determined as the risk detection result of the target language model applied to the model input data; if no risk is hit, then based on the above-mentioned abnormal response detection result, determine whether an abnormal response is hit. If an abnormal response is hit, this risk characterization value can be determined as the risk detection result of the target language model applied to the model input data; if no abnormal response is hit, a pre-set risk-free characterization value, such as the value "0", can be determined as the risk detection result of the target language model applied to the model input data.
[0151] Based on the relevant content of steps 21 to 26 above, it can be seen that in some application scenarios, the information in the three dimensions of intent description content, technique description content, and response description information can be determined first based on the model input data and the model output data of the target language model; then, based on the risk detection results in these three dimensions, the risk detection result of the target language model applied to the model input data can be comprehensively determined, which is beneficial to improving the risk detection effect.
[0152] In addition, the present application does not limit the determination methods of the above-mentioned intent description content, technique description content, and response description information. For example, in some application scenarios, in order to better improve the detection effect, when the above-mentioned model input data includes multiple data with different source types, the intent description content, the technique description content, and the response description information are determined based on the source types of the multiple data. This is beneficial to implementing different detection processes for data of different source types, thereby being beneficial to improving the detection effect.
[0153] It can be seen that in a possible implementation, when the above-mentioned model input data includes a plugin response and at least one previous prompt word, the source type of the plugin response is different from the source types of the respective prompt words, and the source types of different prompt words are different, the response description information can be determined based on the plugin response and / or the above-mentioned model output data, and the above-mentioned intention description content and the above-mentioned technique description content are determined based on the at least one prompt word and the source type of the at least one prompt word. It should be noted that for the determination process of the intention description content and the technique description content, please refer to the above. In addition, the present application does not limit the implementation manner of the response description information. For example, in some application scenarios, such as the scenario of paying attention to the risk of SP leakage, if the at least one prompt word includes a system prompt word, the response description information can be determined based on the plugin response, the system prompt word, and the model output data, so as to be able to perform abnormal response detection based on these three types of data subsequently, such as detecting whether there is a risk of SP leakage, etc., which is beneficial to improving the detection effect.
[0154] Based on the above two paragraphs and related content, it can be seen that in a possible implementation, after obtaining the model input data and the model output data of the target language model, the model input data and the model output data can be concatenated into one data; so that subsequently, based on the source types of each part in the concatenated data, the concatenated data can be split into intention description content, technique description content, and response description information, so that subsequently, malicious intention detection processing can be performed based on the intention description content, attack technique detection processing can be performed based on the technique description content, and abnormal response detection processing can be performed based on the response description information, so that subsequently, the final risk detection result can be determined based on these three detection results.
[0155] It can be seen that in a possible implementation, when the data to be detected includes user prompt words, system prompt words, plugin responses, and model output data, the user prompt words and the system prompt words can participate in intention risk detection processing to determine whether there is a malicious intention; the user prompt words can participate in technique risk detection processing to determine whether there is an attack technique detection; the system prompt words, the plugin responses, and the model output data can participate in abnormal response detection processing to determine whether there is an abnormal response, so that subsequently, the final risk detection result can be determined based on these three detection results.
[0156] The inventors of the present disclosure have found through research that when the data to be detected includes data of multiple source types, there are differences in the risk boundaries of data of different source types in different situations. For example, in the scenario of chatbot review, it is necessary to detect whether there is a risk in the system prompt words, but it is not necessary to detect whether there is a risk in the user prompt words. Another example is that in some cases, if the feature of "introducing role setting" appears in the user prompt words, it can be determined that there is a risk, but if the feature of "introducing role setting" appears in the system prompt words, it can be determined that there is no risk. Still another example is that in some cases, if the feature of "enabling the model to call the component to write information" appears in the plugin response, it can be determined that there is a risk, but if the feature of "enabling the model to call the component to write information" appears in the user prompt words, it can be determined that there is no risk. It can be seen that in order to better improve the detection effect, differential discrimination can be made according to different situations to effectively avoid problems such as false positives / negatives in different situations.
[0157] Based on the above research, it can be known that in order to better meet the risk detection requirements in different situations, different detection items can be configured for data of different source types, so as to effectively avoid the defects caused by using the same set of detection items for different situations or using the same set of detection items for data of different source types in the same situation, thereby being beneficial to improving the detection effect. Based on this, the present application also provides a possible implementation manner of the risk detection method of the above language model. In this manner, the risk detection method may at least include the following steps 31-step 32.
[0158] Step 31: Obtain model detection constraint information.
[0159] Among them, the model detection constraint information is used to constrain the risk detection process of the target language model, so that the model detection constraint information can describe the risk detection requirements that need to be met by the risk detection process, such as which features need to be detected for risk or what kind of detection needs to be performed for each feature.
[0160] In addition, the present application does not limit the implementation manner of the model detection constraint information. For example, the model detection constraint information may include at least one of scenario constraint information, region constraint information, service constraint information, language constraint information, and detection item constraint information. For the sake of easy understanding, these constraint information will be introduced separately below.
[0161] For the above-mentioned scenario constraint information, the scenario constraint information is used to describe the scenario requirements for the risk detection process of the target language model, such as requirements in aspects such as application scenarios and / or usage scenarios; moreover, this application does not limit the implementation manner of the scenario constraint information. For example, the scenario constraint information may include application scenario description information and / or usage scenario description information. Among them, the application scenario description information is used to describe the application scenario of the target language model; moreover, this application does not limit the implementation manner of the application scenario description information. For example, it may be implemented using application scenario parameters pre-configured for the target language model and / or application scenario parameters specified by the user. In addition, this application does not limit the implementation manner of the application scenario. For example, the application scenario may be implemented using a chatbot, an agent, a code assistant, or multimodality. The usage scenario description information is used to describe the usage scenario of the target language model so that the usage scenario description information can represent some usage situations of the target language model; moreover, this application does not limit the implementation manner of the usage scenario description information. For example, it may be implemented using usage scenario parameters pre-configured for the target language model and / or usage scenario parameters specified by the user. In addition, this application does not limit the implementation manner of the usage scenario. For example, the usage scenario may be implemented using Runtime or SpaAudit.
[0162] For the above-mentioned region constraint information, the region constraint information is used to describe the region requirements for the risk detection process of the target language model, such as requirements for a certain region, a certain country, or a certain continent, etc.; moreover, this application does not limit the implementation manner of the region constraint information. For example, it may be implemented using region parameters pre-configured for the target language model and / or region parameters specified by the user.
[0163] For the above-mentioned business constraint information, the business constraint information is used to describe the business requirements for the risk detection process of the target language model, such as requirements for Business 1, Business 2, or Business 3, etc.; moreover, this application does not limit the implementation manner of the business constraint information. For example, it may be implemented using business parameters pre-configured for the target language model and / or business parameters specified by the user.
[0164] For the above-mentioned language constraint information, the language constraint information is used to describe the language requirements for the risk detection process of the target language model, such as requirements for Chinese, English, or Japanese, etc.; moreover, this application does not limit the implementation manner of the language constraint information. For example, it may be implemented using language parameters pre-configured for the target language model and / or language parameters specified by the user.
[0165] For the above-mentioned detection item constraint information, the detection item constraint information is used to describe the requirements of the detection items for the risk detection process of the target language model, such as the requirement to perform SP leakage detection items, etc.; moreover, the present application does not limit the implementation manner of the detection item constraint information. For example, it can be implemented by using the detection item parameters pre-configured for the target language model and / or the detection item parameters specified by the user.
[0166] Based on the relevant content of the above-mentioned model detection constraint information, in a possible implementation manner, the model detection constraint information can be determined according to the user input and the configuration information pre-configured for the target language model, so that the model detection constraint information can represent some requirements that need to be met when performing risk detection processing on the target language model, such as Figure 3 the requirements in aspects such as the application scenario, usage scenario, region, business, language, and detection item shown, so as to better complete the risk detection processing of the target language model based on these requirements subsequently.
[0167] In addition, the present application does not limit the execution time of step 31 above, as long as it is ensured that the execution time of step 31 is earlier than the execution time of step 32 below.
[0168] Step 32: Determine a detection execution device that matches the above-mentioned model detection constraint information from the detection execution devices corresponding to multiple candidate constraint information, so that the matching detection execution device is at least used to perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result, and / or perform technique risk detection processing on the determined technique description content to obtain a technique risk detection result.
[0169] Among them, the multiple candidate constraint information refers to some pre-set optional constraint information, such as multiple application scenarios, multiple usage scenarios, multiple regions, multiple businesses, multiple languages, multiple detection items, etc. It can be seen that in a possible implementation manner, the multiple candidate constraint information may include some or all of at least one candidate application scenario, at least one candidate usage scenario, at least one candidate region, at least one candidate business, at least one candidate language, and at least one candidate detection item.
[0170] In addition, for the i-th candidate constraint information, the detection execution device corresponding to the i-th candidate constraint information refers to a device with a detection function that is pre-configured for the i-th candidate constraint information, so that the detection execution device corresponding to the i-th candidate constraint information is used to perform detection processing that meets the i-th candidate constraint information, such as intent risk detection processing, technique risk detection processing, and abnormal response detection processing, etc.; moreover, the present application does not limit the implementation manner of the detection execution device corresponding to the i-th candidate constraint information. For example, the detection execution device corresponding to the i-th candidate constraint information may include some or all of a malicious intent detector, an attack technique detector, and an abnormal response detector. Among them, the malicious intent detector is used to perform intent risk detection processing that meets the i-th candidate constraint information to identify malicious intent in the situation described by the i-th candidate constraint information. The attack technique detector is used to perform technique risk detection processing that meets the i-th candidate constraint information to identify attack techniques in the situation described by the i-th candidate constraint information. The abnormal response detector is used to perform abnormal response detection processing that meets the i-th candidate constraint information to identify abnormal responses in the situation described by the i-th candidate constraint information. Wherein, i is a positive integer, and i ≤ the number of pieces of information in the above-mentioned multiple candidate constraint information.
[0171] In addition, in order to better improve the detection effect, the present application also provides a possible implementation manner of the above-mentioned malicious intent detector. In this implementation manner, the malicious intent detector may include intent detectors corresponding to different source types, such as an intent detector corresponding to user prompt words and an intent detector corresponding to system prompt words, etc., to implement different intent risk detection processing for data of different source types, which is beneficial to better meeting the intent detection requirements in the situation described by the i-th candidate constraint information, thereby being beneficial to improving the detection effect. Among them, the intent detector corresponding to the user prompt words is used to perform intent risk detection processing on the intent content determined according to the user prompt words to identify whether there is malicious intent in the intent content, so as to meet the intent detection requirements for the user prompt words in this situation. The intent detector corresponding to the system prompt words is used to perform intent risk detection processing on the intent content determined according to the system prompt words to identify whether there is malicious intent in the intent content, so as to meet the intent detection requirements for the system prompt words in this situation.
[0172] In addition, to better improve the detection effect, the present application also provides a possible implementation manner of the above abnormal response detector. In this implementation manner, the abnormal response detector may include response detectors corresponding to different source types, such as a response detector corresponding to a plug-in response and a response detector corresponding to model output data, etc., so as to implement different abnormal response detection processes for data of different source types. This is beneficial to better meet the response detection requirements in the situation described by the ith candidate constraint information, thereby facilitating the improvement of the detection effect. Among them, the response detector corresponding to the plug-in response is used to perform abnormal response detection processing on the plug-in response to identify whether there is abnormal content in the plug-in response, so as to meet the response detection requirements for the plug-in response in the model input data in this situation. The response detector corresponding to the model output data is used to perform abnormal response detection processing on the model output data to identify whether there is abnormal content in the model output data, so as to meet the response detection requirements for the model output data in this situation.
[0173] Based on the relevant content of the above ith candidate constraint information, in some application scenarios, the detection execution device corresponding to the ith candidate constraint information may include detectors corresponding to data of different source types, so that these detectors can meet the detection requirements in the situation described by the ith candidate constraint information, such as malicious intention detection requirements and / or abnormal response detection requirements, etc. Among them, for any data of a source type, the detector corresponding to this data is constructed according to the detection requirements set for this source type in this situation, such as being constructed using Figure 3 the rule library, vector library, small model library, and large model library shown, etc., so that the detection process implemented by the detector corresponding to this data can better meet the detection requirements set for this source type in this situation. In this way, it can effectively meet the different detection requirements for data of different source types in different situations, thereby facilitating the improvement of the detection effect.
[0174] In addition, the present application does not limit the implementation manner of the detection execution device corresponding to the i-th candidate constraint information above. For example, it can be implemented by means of one or more of rule checking, vector recall, and classification. It can be seen that in a possible implementation manner, the construction method of the detection execution device corresponding to the i-th candidate constraint information is as follows: combining some or all of the rules in the rule library, some or all of the vectors in the vector library, some or all of the small-parameter models in the small model library, and some or all of the large-parameter models in the large model library to obtain the detection execution device corresponding to the i-th candidate constraint information, so as to meet the detection requirements in the situation described by the i-th candidate constraint information. Among them, the rule library is used to record the rules required for rule checking. The vector library is used to record the vectorized features of the attack samples required for vector recall. The small model library is used to record the small-parameter models with classification functions required for classification. The large model library is used to record the large-parameter models with classification functions required for classification.
[0175] In addition, the present application does not limit the construction methods of the rule library, vector library, small model library, and large model library in the above paragraph. For example, specifically, it can be as follows: First, collect business data through methods such as log embedding, offline data (Hive / ClickHouse), etc. to construct negative samples including data such as UserPrompt, SystemPrompt, and Tool, such as Figure 3 the negative samples shown, so that the negative samples can represent samples without attack risks, and collect attack data through open-source data, online attacks, automatic generation, etc. to construct positive samples, such as Figure 3 the positive samples shown, so that the positive samples can represent samples with attack risks; then, perform supervised fine-tuning (SFT), FewShot, vectorized storage, rule construction, etc. on these negative samples and positive samples to form rules and models for detecting different types of attack features, such as malicious intent, attack techniques, abnormal responses, etc.; finally, store these rules and these models in a database manner to obtain Figure 3 the rule library, vector library, small model library, and large model library shown, so that detectors for detecting data of different source types in different situations can be constructed using these libraries later.
[0176] Furthermore, for the above-mentioned target language model, after obtaining the model detection constraint information corresponding to the target language model, a detection execution device that matches the model detection constraint information can be determined from the detection execution devices corresponding to multiple candidate constraint information, such as a detection execution device that matches the scenario constraint information, a detection execution device that matches the region constraint information, a detection execution device that matches the service constraint information, a detection execution device that matches the language constraint information, and a detection execution device that matches the detection item constraint information, etc., so that the "detection execution device that matches the model detection constraint information" can meet the detection requirements described by the model detection constraint information, so as to subsequently implement the detection and processing of the model input data and / or model data of the target language model with the help of the "detection execution device that matches the model detection constraint information", such as intention risk detection processing, technique risk detection processing, and abnormal response detection processing, etc.
[0177] It can be seen that for the above-mentioned "detection execution device that matches the model detection constraint information", in some scenarios, such as the prompt word detection scenario, the "detection execution device that matches the model detection constraint information" can be used to perform intention risk detection processing on the intention description content to obtain an intention risk detection result, and perform technique risk detection processing on the technique description content to obtain a technique risk detection result. In other scenarios, such as the prompt word and response detection scenario, the "detection execution device that matches the model detection constraint information" can be used to perform intention risk detection processing on the intention description content to obtain an intention risk detection result, perform technique risk detection processing on the technique description content to obtain a technique risk detection result, and perform abnormal response detection processing on the response description information to obtain an abnormal response detection result.
[0178] In addition, the present application does not limit the implementation method of the above-mentioned "detection execution device that matches the model detection constraint information". For example, in order to better improve the detection effect, when the "detection execution device that matches the model detection constraint information" is at least used to perform intent risk detection processing on the intent description content to obtain intent risk detection results, and to perform technique risk detection processing on the technique description content to obtain technique risk detection results, the "detection execution device that matches the model detection constraint information" may at least include: a detector corresponding to the intent description content and a detector corresponding to the technique description content. Among them, the detector corresponding to the intent description content is used to represent the malicious intent detector that matches the model detection constraint information, so that the detector corresponding to the intent description content can be used to perform intent risk detection processing on the intent description content to obtain intention risk detection results; the detector corresponding to the technique description content is used to represent the attack technique detector that matches the model detection constraint information, so that the detector corresponding to the technique description content can be used to perform technique risk detection processing on the technique description content to obtain technique risk detection results.
[0179] In addition, based on the relevant content of the malicious intent detector above, in order to better improve the detection effect, the present application also provides a possible implementation method of the above "detector corresponding to the intent description content". In this way, when the intent description content includes at least two source types of intent content, the "detector corresponding to the intent description content" may include detectors corresponding to various source types of intent content, such as intent detectors corresponding to user prompt words and intent detectors corresponding to system prompt words. Among them, for the detector corresponding to the intent content of any source type, the detector is used to perform intent risk detection processing on the intent content of the source type, so that the aforementioned "detector corresponding to the intent description content" can perform different malicious intent detection processing on intent content of different source types, which can better meet the different malicious intent detection requirements for intent content of different source types in the current situation, thereby helping to improve the detection effect.
[0180] Furthermore, the present application does not limit the implementation method of the above step 32. For example, it can be implemented by means of task distribution, such as Figure 3Implement it in the task distribution method shown. It can be seen that in a possible implementation, step 32 can be: first, determine the detection execution device that matches the above model detection constraint information from the detection execution devices corresponding to multiple candidate constraint information; then send a detection request to the matching detection execution device, so that the matching detection execution device is at least used to perform intent risk detection processing on the intent description content carried in the detection request to obtain an intent risk detection result, and / or perform technique risk detection processing on the technique description content carried in the detection request to obtain a technique risk detection result; then, receive the feedback information of the matching detection execution device, so that the feedback information includes the intent risk detection result and / or the technique risk detection result. Among them, the detection request is used to request the matching detection execution device to perform detection processing on the data carried in the detection request, such as the intent description content, the technique description content, and the above response description information.
[0181] Based on the relevant content of the above steps 31 to 32, it can be known that in some application scenarios, such as Figure 3 In the scenario shown, after obtaining the model input data and model output data of the target language model, the corresponding content in the model input data and model output data can be distributed according to factors such as the application scenario, usage scenario, and region in the user-specified parameters and / or configuration parameters, so that the corresponding device can use means such as rule checking, vector recall, and classification methods to detect features such as malicious intent, attack techniques, and abnormal responses, and then summarize these three results to obtain the risk detection result of the target language model applied to the model input data. In this way, different detection processes can be performed on different source types of data in different situations, so as to better meet the detection needs of different source types of data in different situations, which is conducive to improving the detection effect.
[0182] It should be noted that for the step of "the corresponding device can use means such as rule checking, vector recall, and classification methods to detect features such as malicious intent, attack techniques, and abnormal responses" in the above paragraph, in some application scenarios, such as the scenario where a large number of detection requests are concurrent, during the execution of this step, it may experience Figure 3 The rough screening and fine screening shown, so as to identify the requests with attack risks from these detection requests as quickly as possible, which is conducive to improving the detection efficiency.
[0183] In fact, to better improve the detection effect, the present application also provides a possible implementation manner of the risk detection method for the above language model. In this manner, when the detection execution devices corresponding to multiple candidate constraint information above are constructed based on a pre-constructed detection set, and the detection set includes one or more of at least one detection rule, at least one detection vector, and at least one detection model, the risk detection method may at least include steps 31-32 above and step 33 below.
[0184] Step 33: If the risk detection result of the above target language model applied to the model input data indicates a risk, update the detection set based on the model input data.
[0185] Among them, the detection set refers to the object required for constructing the detection execution devices corresponding to multiple candidate constraint information, such as Figure 3 the rules in the rule library, the vectors in the vector library, the models in the small model library, and the models in the large model library shown.
[0186] In addition, for the above detection set, the detection set includes one or more of at least one detection rule, at least one detection vector, and at least one detection model. Among them, the detection rule refers to the rule required for detection processing; and the present application does not limit the at least one detection rule. For example, the at least one detection rule may include some or all of the rules in the rule library. The detection vector refers to the vector required for detection processing; and the present application does not limit the at least one detection vector. For example, the at least one detection vector may include some or all of the vectors in the vector library. The detection model refers to the model required for detection processing, such as a small-parameter model or a large-parameter model with classification functions, etc.; and the present application does not limit the at least one detection model. For example, the at least one detection model may include some or all of the models in the small model library and some or all of the models in the large model library. It can be seen that in a possible implementation manner, the detection set may include at least one of the rule library, the vector library, the small model library, and the large model library.
[0187] In addition, the present application does not limit the implementation manner of step 33 above. For example, it may specifically be: If the risk detection result of the above target language model applied to the model input data indicates a risk, update the detection set based on the model input data as a positive sample, so that the subsequent detection processing based on the updated detection set can accurately identify that the model input data has an attack risk.
[0188] Based on the relevant content of step 33 above, after obtaining the risk detection result of the target language model applied to the model input data, if the risk detection result indicates a risk, the model input data can be used as an attack sample to update the rule library, vector library, small model library, large model library, etc., so as to improve the detection function of the detector constructed based on these libraries, so that the detection processing implemented based on these detectors can achieve a better detection effect, which is beneficial to improving the detection effect.
[0189] In fact, in order to better improve security, the present application also provides a possible implementation manner of the risk detection method of the above language model. In this manner, the risk detection method may at least include step 41 below.
[0190] Step 41: If the risk detection result of the above target language model applied to the model input data indicates a risk, the model output data of the target language model is processed, such as interception processing, alarm processing, or rewriting processing, etc., to improve security.
[0191] In the present application, after obtaining the risk detection result of the target language model applied to the model input data, if the risk detection result indicates a risk, interception processing, alarm processing, or rewriting processing, etc., can be performed on the model output data of the target language model to overcome the security problems caused by the risk, which is beneficial to improving security.
[0192] In addition, the present application does not limit the implementation manner of the risk detection method of the above language model. For example, the risk detection method can adopt a detection module, such as in the form of a plugin, Figure 2 the analysis engine shown, or Figure 3 the analysis engine shown, etc., to be implemented, so that during the business interaction process, the interface of the detection module can be called before the model input data is input into the target language model and after the target language model outputs the model output data, so that the business can perform corresponding processing on the model output data based on the detection result feedback by the detection module, such as interception / alarm / rewriting / normal use, etc., so as to achieve protection without changing the business process and without changing the target language model, which is beneficial to avoiding the impact of this protection on the business and the target language model, and further beneficial to improving the protection effect.
[0193] Based on the relevant content of the risk detection method of the above language model, it can be seen that the technical solution provided by the present application has the advantages shown in the following (1) to (3).
[0194] (1) This application performs risk detection processing on the prompt words from two dimensions: malicious intent and attack methods, and aggregates the detection results under these two dimensions through voting to obtain the final risk detection result. This detection method enables the technical solution of this application to not show exponential growth in the number of rules and the number of vector library samples as the number of covered attack types increases, and has good scalability, generalization, and timeliness.
[0195] (2) This application splits data of different source types such as System Prompt, User Prompt, and Plugin Response in the model input data, and uses different models and rules for differential detection to adapt to the risk boundaries of different scenarios, so as to be able to perform detection in different business scenarios.
[0196] (3) This application fully combines the model output data for detection fallback, and determines whether there is an attack risk from the actual output of the model, so as to effectively identify attack requests with unclear features and unknown attack requests.
[0197] Based on the risk detection method of the language model provided by the embodiments of this application, the embodiments of this application also provide a risk detection device for a language model. The following will be combined with Figure 8 for explanation and illustration. Among them, Figure 8 is a schematic structural diagram of a risk detection device for a language model provided by the embodiments of this application. It should be noted that for the technical details of the risk detection device for a language model provided by the embodiments of this application, please refer to the relevant content of the risk detection method of the language model above.
[0198] As Figure 8 shown, the risk detection device 800 for a language model provided by the embodiments of this application includes:
[0199] A data acquisition unit 801, configured to acquire the model input data of the target language model;
[0200] A content determination unit 802, configured to determine at least one of the intent description content and the method description content from the model input data;
[0201] A data detection unit 803, configured to perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result, and / or perform method risk detection processing on the determined method description content to obtain a method risk detection result;
[0202] A result determination unit 804, configured to determine the risk detection result of the target language model applied to the model input data according to at least one of the obtained intent risk detection result and method risk detection result.
[0203] In a possible implementation, the model input data includes at least one prompt word;
[0204] The content determination unit 802 is specifically configured to: determine at least one of the intent description content and the technique description content from the at least one prompt word according to different source types of the respective prompt words in the at least one prompt word.
[0205] In a possible implementation, the source types of the respective prompt words in the at least one prompt word include: system prompt words and user prompt words; the intent description content is determined according to the system prompt words and user prompt words in the model input data; the technique description content is determined according to the user prompt words in the model input data.
[0206] In a possible implementation, the data detection unit 803 is further configured to perform an abnormal response detection process on the response description information to obtain an abnormal response detection result; the response description information is determined according to the model input data and / or the model output data of the target language model; the model output data is obtained by the target language model processing the model input data;
[0207] The risk detection result of the target language model applied to the model input data is also determined according to the abnormal response detection result.
[0208] In a possible implementation, the model input data includes multiple data, and the source types of different data are different; the intent description content, the technique description content, and the response description information are determined according to the source types of the multiple data.
[0209] In a possible implementation, the multiple data includes a plugin response and at least one prompt word, the source type of the plugin response is different from the source types of the respective prompt words, and the source types of different prompt words are different; the intent description content and the technique description content are determined according to the at least one prompt word and the source types of the at least one prompt word; the response description information is determined according to the plugin response and / or the model output data.
[0210] In a possible implementation, the at least one prompt word includes system prompt words; the response description information is also determined according to the system prompt words.
[0211] In a possible implementation, the risk detection device 800 of the language model further includes:
[0212] A constraint acquisition unit, configured to acquire model detection constraint information;
[0213] An equipment determination unit, configured to determine a detection execution device that matches the model detection constraint information from among multiple candidate constraint information-corresponding detection execution devices;
[0214] The matched detection execution device is configured to perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result, and / or perform method risk detection processing on the determined method description content to obtain a method risk detection result.
[0215] In a possible implementation manner, the model detection constraint information includes at least one of scenario constraint information, regional constraint information, service constraint information, language constraint information, and detection item constraint information.
[0216] In a possible implementation manner, the matched detection execution device includes a detector corresponding to the intent description content and a detector corresponding to the method description content; the detector corresponding to the intent description content is configured to perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result; the detector corresponding to the method description content is configured to perform method risk detection processing on the determined method description content to obtain a method risk detection result.
[0217] In a possible implementation manner, the intent description content includes intent content of at least two source types; the detector corresponding to the intent description content includes detectors corresponding to intent content of various source types; for any detector corresponding to intent content of a source type, the detector is configured to perform intent risk detection processing on the intent content of that source type.
[0218] In a possible implementation manner, the data detection unit 803 is specifically configured to:
[0219] Send a detection request to the matched detection execution device; the matched detection execution device is configured to perform intent risk detection processing on the intent description content carried in the detection request to obtain an intent risk detection result, and / or perform method risk detection processing on the method description content carried in the detection request to obtain a method risk detection result;
[0220] Receive feedback information from the matched detection execution device, where the feedback information includes the intent risk detection result and / or the method risk detection result.
[0221] In a possible implementation, the matching detection execution device is further configured to perform abnormal response detection processing on the response description information to obtain an abnormal response detection result; the response description information is determined based on the model input data and / or the model output data of the target language model; the model output data is obtained by the target language model processing the model input data.
[0222] In a possible implementation, the detection execution devices corresponding to the multiple candidate constraint information are constructed based on a pre-constructed detection set, and the detection set includes one or more of at least one detection rule, at least one detection vector, and at least one detection model;
[0223] The risk detection device 800 of the language model further includes:
[0224] An update unit, configured to, after determining the risk detection result of the target language model applied to the model input data, if the risk detection result of the target language model applied to the model input data indicates a risk, update the detection set according to the model input data.
[0225] Based on the relevant content of the risk detection device 800 of the above language model, the working principle of the risk detection device 800 of the language model provided in this application is as follows: after obtaining the model input data of the target language model, first determine at least one of the intent description content and the technique description content from the model input data; then, perform intent risk detection processing on the intent description content to obtain an intent risk detection result, and / or perform technique risk detection processing on the technique description content to obtain a technique risk detection result; finally, determine the risk detection result of the target language model under the model input data according to the intent risk detection result and / or the technique risk detection result, so that the risk detection result can indicate whether there is a risk, such as Prompt injection risk, etc., thereby effectively avoiding the security problems caused by the risk, and further being beneficial to improving security. It can be seen that since the risk detection device 800 of the language model performs risk detection processing on the model input data from the two dimensions of intent and technique, so that the risk detection device 800 of the language model can analyze whether there is a Prompt injection risk from the two characteristics of malicious intent and attack technique, thereby enabling the risk detection device 800 of the language model to not only detect the pre-learned attack samples, but also detect the new attack requests composed of the malicious intent and attack techniques that appear in these attack samples. Furthermore, the number of samples required for the risk detection device 800 of the language model will not increase exponentially with the increase in the number of risk types. In this way, the risk detection device 800 of the language model has good scalability, generalization, and timeliness.
[0226] In addition, an embodiment of the present application further provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation manner of the risk detection method of the language model provided by the embodiment of the present application.
[0227] See Figure 9 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0228] As Figure 9 shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which may perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 902 or the programs loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.
[0229] Generally, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9 shows an electronic device 900 having various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0230] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, the above functions defined in the method of the embodiment of the present disclosure are performed.
[0231] The electronic device provided in the embodiment of the present disclosure and the method provided in the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0232] An embodiment of the present application further provides a computer-readable medium, in which instructions or a computer program are stored. When the instructions or the computer program run on a device, the device is caused to execute any implementation manner of the risk detection method of the language model provided in the embodiment of the present application.
[0233] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0234] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0235] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0236] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the above method.
[0237] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0238] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0239] The units involved in the embodiments described in this disclosure can be implemented in software or in hardware. Among them, the name of the unit / module does not constitute a limitation to the unit itself in some cases.
[0240] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0241] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0242] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0243] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0244] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0245] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the art.
[0246] The foregoing description of the disclosed embodiments enables those skilled in the art to make or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A risk detection method for a language model, characterized in that, The method includes: Obtaining model input data of a target language model; Determining at least one of intent description content and technique description content from the model input data; Performing intent risk detection processing on the determined intent description content to obtain an intent risk detection result, where the intent risk detection result indicates whether there is a malicious intent in the intent description content, and / or performing technique risk detection processing on the determined technique description content to obtain a technique risk detection result, where the technique risk detection result indicates whether there is content for describing an attack technique in the technique description content; Determining a risk detection result of the target language model applied to the model input data according to at least one of the obtained intent risk detection result and technique risk detection result, where the risk detection result indicates whether there is a prompt injection risk in the model input data, and the model attacks indicated by the prompt injection risk include the malicious intent and the attack technique.
2. The method according to claim 1, wherein The model input data includes at least one prompt; The determining at least one of intent description content and technique description content from the model input data includes: Determining at least one of intent description content and technique description content from the at least one prompt according to different source types of the respective prompts in the at least one prompt.
3. The method according to claim 2, wherein The source types of the respective prompts in the at least one prompt include: system prompts and user prompts; The intent description content is determined according to system prompts and user prompts in the model input data; The technique description content is determined according to user prompts in the model input data.
4. The method according to claim 1, wherein The method further includes: Performing abnormal response detection processing on response description information to obtain an abnormal response detection result; the response description information is determined according to the model input data and / or model output data of the target language model; the model output data is obtained by the target language model processing the model input data; The risk detection result of the target language model applied to the model input data is further determined according to the abnormal response detection result.
5. The method according to claim 4, wherein The model input data includes multiple data, and the source types of different data are different; The intent description content, the technique description content, and the response description information are determined according to the source types of the multiple data.
6. The method according to claim 5, wherein The multiple data include a plugin response and at least one prompt, the source type of the plugin response is different from the source types of the respective prompts, and the source types of different prompts are different; The intent description content and the technique description content are determined according to the at least one prompt and the source types of the at least one prompt; The response description information is determined according to the plugin response and / or the model output data.
7. The method according to claim 6, wherein The at least one prompt includes system prompts; The response description information is further determined according to the system prompts.
8. The method according to claim 1, characterized in that, The method further includes: Obtaining model detection constraint information; Determine a detection execution device that matches the model detection constraint information from the detection execution devices corresponding to multiple candidate constraint information; The matching detection execution device is used to perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result, and / or perform technique risk detection processing on the determined technique description content to obtain a technique risk detection result.
9. The method according to claim 8, wherein The model detection constraint information includes at least one of scenario constraint information, regional constraint information, service constraint information, language constraint information, and detection item constraint information.
10. The method according to claim 8, wherein The matching detection execution device includes a detector corresponding to the intent description content and a detector corresponding to the technique description content; The detector corresponding to the intent description content is used to perform intent risk detection processing on the determined intent description content to obtain an intent risk detection result; The detector corresponding to the technique description content is used to perform technique risk detection processing on the determined technique description content to obtain a technique risk detection result.
11. The method according to claim 10, wherein, The intent description content includes intent content of at least two source types; The detector corresponding to the intent description content includes detectors corresponding to intent content of various source types; For the detector corresponding to the intent content of any source type, the detector is used to perform intent risk detection processing on the intent content of this source type.
12. The method according to claim 8, wherein After determining the detection execution device that matches the model detection constraint information from the detection execution devices corresponding to multiple candidate constraint information, the method further includes: Send a detection request to the matching detection execution device; the matching detection execution device is used to perform intent risk detection processing on the intent description content carried in the detection request to obtain an intent risk detection result, and / or perform technique risk detection processing on the technique description content carried in the detection request to obtain a technique risk detection result; Receive feedback information from the matching detection execution device, where the feedback information includes the intent risk detection result and / or the technique risk detection result.
13. The method according to any one of claims 8 - 12, characterized in that, The matching detection execution device is further used to perform abnormal response detection processing on the response description information to obtain an abnormal response detection result; The response description information is determined based on the model input data and / or the model output data of the target language model; the model output data is obtained by the target language model processing the model input data.
14. The method according to claim 8, characterized in that The detection execution devices corresponding to the multiple candidate constraint information are constructed based on a pre-constructed detection set, and the detection set includes one or more of at least one detection rule, at least one detection vector, and at least one detection model; After determining the risk detection result of the target language model applied to the model input data, the method further includes: If the risk detection result of the target language model applied to the model input data indicates a risk, update the detection set based on the model input data.
15. A risk detection device for a language model, characterized in that, Includes: A data acquisition unit, configured to acquire the model input data of the target language model; A content determination unit for determining at least one of an intent description content and a technique description content from the model input data; A data detection unit for performing an intent risk detection process on the determined intent description content to obtain an intent risk detection result, the intent risk detection result indicating whether there is a malicious intent in the intent description content, and / or performing a technique risk detection process on the determined technique description content to obtain a technique risk detection result, the technique risk detection result indicating whether there is content for describing an attack technique in the technique description content; A result determination unit for determining a risk detection result of the target language model applied to the model input data based on at least one of the obtained intent risk detection result and technique risk detection result, the risk detection result indicating whether there is a prompt injection risk in the model input data, and the model attack indicated by the prompt injection risk includes the malicious intent and the attack technique.
16. An electronic device, characterized in that, The device includes: a processor and a memory; The memory for storing instructions or computer programs; The processor for executing the instructions or computer programs in the memory so that the electronic device executes the method according to any one of claims 1-14.
17. A computer-readable medium, characterized in that, Instructions or computer programs are stored in the computer-readable medium, and when the instructions or computer programs are run on the device, the device executes the method according to any one of claims 1-14.
18. A computer program product, characterized in that, It includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method according to any one of claims 1-14.
Citation Information
Patent Citations
Risk management method and device, equipment and storage medium
CN115099680A
Large model risk assessment method and device, storage medium and electronic equipment
CN117113339A
Method, device and equipment for risk identification and readable medium
CN118014363A