Method and device for applying large language model

Through methods of extracting and optimizing user problems, combined with security tips, the problem that large language models are difficult to detect and defend against harmful content when facing jailbreak attacks is solved, achieving more efficient and secure content generation.

CN118709195BActive Publication Date: 2025-05-16BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410866848.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-05-16
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

When facing jailbreak attacks, large language models are difficult to detect and defend against the generation of harmful content, resulting in security threats.

Method used

By inputting the original problem to be queried together with the preset prompts to a large language model, the problem that responds to the true intention is extracted, and non-critical information is deleted to obtain the optimized problem. Then, enter the question together with the security prompts to make sure its answers comply with the security policy.

Benefits of technology

Effectively reveal the user's true intentions, enhance defense capabilities against jailbreak attacks, ensure the security of content generated by large language models, reduce computing costs and improve response efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118709195B_ABST
    Figure CN118709195B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a method and device for applying a large language model. The specific implementation of the method includes: inputting the original question to be queried together with a preset prompt into the large language model, outputting a first question, wherein the prompt is used to instruct the large language model to extract a question that reflects the true intention from the original question; deleting non-critical information from the first question to obtain a second question; inputting the second question together with a preset security prompt into the large language model, and outputting an answer, wherein the security prompt is used to instruct the large language model to comply with security policies when answering questions. This implementation can effectively defend against jailbreak attacks that are confusing and have implicit intentions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a method and device for applying a large language model. Background Art

[0002] With the deep integration of large language models in real-world applications, potential security flaws inherent in these frameworks have surfaced. These vulnerabilities can be exploited for harmful purposes, such as creating harmful content and supporting illegal activities. One of the main challenges facing the security of LLMs (Large Language Models) is the threat of jailbreaking attacks, which can bypass the calibration mechanisms and security measures of LLMs by embedding malicious queries in carefully crafted prompts, resulting in the generation of harmful content. These attacks are difficult to detect and pose a significant obstacle to the widespread adoption of LLMs. Summary of the invention

[0003] The embodiments of the present disclosure propose a method and apparatus for applying a large language model.

[0004] In a first aspect, an embodiment of the present disclosure provides a method for applying a large language model, comprising: inputting an original question to be queried together with a preset prompt into the large language model, and outputting a first question, wherein the prompt is used to instruct the large language model to extract a question that reflects the true intention from the original question; deleting non-critical information from the first question to obtain a second question; inputting the second question together with a preset security prompt into the large language model, and outputting an answer, wherein the security prompt is used to instruct the large language model to comply with security policies when answering questions.

[0005] In some embodiments, the first question includes a target sentence; and the deleting non-critical information from the first question to obtain the second question includes: marking the words in the first question to obtain a first tag sequence; deleting a predetermined number of tags from the first tag sequence in any combination to obtain at least one second tag sequence; inputting the text corresponding to the at least one second tag sequence into the large language model together with the prompt to obtain a probability of generating the target sentence; and determining the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the second question.

[0006] In some embodiments, the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold is determined as the second problem, including: calculating the loss value of each second tag sequence based on the negative logarithm of the probability of each second tag sequence generating the target sentence; and determining the second tag sequence with the smallest loss value as the second problem.

[0007] In some embodiments, the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold is determined as the second problem, including: calculating the loss value of each second tag sequence according to the negative logarithm of the probability of each second tag sequence generating the target sentence; calculating the gradient of each tag based on the loss value of each second tag sequence; and deleting the tag with the smallest gradient from the first tag sequence to obtain the second problem.

[0008] In some embodiments, the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold is determined as the second question, including: marking the words in the second question to obtain a third tag sequence; deleting a predetermined number of tags from the third tag sequence in any combination to obtain at least one fourth tag sequence; inputting the text corresponding to at least one fourth tag sequence into the large language model together with the prompt to obtain the probability of generating the target sentence; and determining the fourth tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the updated second question.

[0009] In some embodiments, the method further includes: inputting the second question together with a preset prompt into a large language model, and outputting a third question; inputting the third question together with a preset security prompt into the large language model, and outputting an answer.

[0010] In some embodiments, outputting the answer includes: in response to detecting that the second question does not comply with the security policy, outputting information that the answer is rejected.

[0011] In a second aspect, an embodiment of the present disclosure provides a device for applying a large language model, comprising: an extraction unit, configured to input an original question to be queried together with a preset prompt into the large language model, and output a first question, wherein the prompt is used to instruct the large language model to extract a question that reflects the true intention from the original question; a deletion unit, configured to delete non-critical information from the first question to obtain a second question; and a questioning unit, configured to input the second question together with a preset security prompt into the large language model, and output an answer, wherein the security prompt is used to instruct the large language model to comply with security policies when answering questions.

[0012] In some embodiments, the first question includes a target sentence; and the deletion unit is further configured to: mark the words in the first question to obtain a first tag sequence; delete a predetermined number of tags from the first tag sequence in any combination to obtain at least one second tag sequence; input the text corresponding to the at least one second tag sequence into the large language model together with the prompt to obtain a probability of generating the target sentence; determine the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the second question.

[0013] In some embodiments, the deletion unit is further configured to: calculate the loss value of each second tag sequence according to the negative logarithm of the probability of each second tag sequence generating the target sentence; and determine the second tag sequence with the smallest loss value as the second question.

[0014] In some embodiments, the deletion unit is further configured to: calculate the loss value of each second tag sequence based on the negative logarithm of the probability of each second tag sequence generating the target sentence; calculate the gradient of each tag based on the loss value of each second tag sequence; delete the tag with the smallest gradient from the first tag sequence to obtain a second question.

[0015] In some embodiments, the deletion unit is further configured to: mark the words in the second question to obtain a third tag sequence; delete a predetermined number of tags from the third tag sequence in any combination to obtain at least one fourth tag sequence; input the text corresponding to at least one fourth tag sequence into the large language model together with the prompt to obtain the probability of generating the target sentence; determine the fourth tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the updated second question.

[0016] In some embodiments, the device also includes a loop unit configured to: input the second question together with a preset prompt into a large language model, and output a third question; input the third question together with a preset security prompt into the large language model, and output an answer.

[0017] In some embodiments, the questioning unit is further configured to: in response to detecting that the second question does not comply with the security policy, output information that the answer is rejected.

[0018] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: one or more processors; a storage device on which one or more computer programs are stored, and when the one or more computer programs are executed by the one or more processors, the one or more processors implement a method as described in any one of the first aspect or the second aspect.

[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method as described in any one of the first aspect or the second aspect is implemented.

[0020] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which implements the method as described in any one of the first aspect or the second aspect when executed by a processor.

[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Other features, objects and advantages of the present disclosure will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0023] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;

[0024] Figure 2 is a flow chart of an embodiment of a method for applying a large language model according to the present disclosure;

[0025] Figure 3 is a schematic diagram of an application scenario of the method for applying a large language model according to the present disclosure;

[0026] Figure 4 is a flowchart of another embodiment of a method for applying a large language model according to the present disclosure;

[0027] Figure 5 is a structural diagram of an embodiment of a device for applying a large language model according to the present disclosure;

[0028] Figure 6 It is a structural diagram of a computer system of an electronic device suitable for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] The present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0030] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0031] Figure 1 An exemplary system architecture 100 is shown to which an embodiment of a method for applying a large language model or an apparatus for applying a large language model of the present disclosure can be applied.

[0032] like Figure 1As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0033] The user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as large language model applications, 3D video players, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0034] Terminal devices 101, 102, 103 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Group Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Group Audio Layer 4) players, laptop computers and desktop computers, etc. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0035] The server 105 may be a server that has a large language model installed to provide large language services, such as a question-answering server that generates answers to questions submitted by the terminal devices 101, 102, and 103. The question-answering server may extract the real question from the question submitted by the user according to the preset prompts and safety prompts, simplify the question, generate an answer through the large language model, and return it to the terminal device.

[0036] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, multiple software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0037] It should be noted that the method for applying a large language model provided in the embodiments of the present disclosure is generally executed by the server 105 , and accordingly, the device for applying a large language model is generally disposed in the server 105 .

[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0039] Continue to refer Figure 2 , shows a process 200 of an embodiment of a method for applying a large language model according to the present disclosure. The method for applying a large language model includes the following steps:

[0040] Step 201: input the original question to be queried together with the preset prompt into the large language model, and output the first question.

[0041] In this embodiment, the execution subject (eg, Figure 1 The server shown in the figure) can receive the original question to be queried sent by the terminal device through a wired connection or a wireless connection.

[0042] The original question can be a question generated by the jailbreak attack method. The jailbreak attack on the securely trained model is performed by submitting a jailbreak prompt P JB is an attempt to elicit a topic-relevant response to the malicious question q. Let Q JB Represents the entire jailbreak query:

[0043] Q JB =P JB ⊕q, (1)

[0044] where ⊕ denotes a combination operation. Stages 1, 2, and 3 provide examples of how we apply formula (1).

[0045]

[0046]

[0047] This application presets a prompt (P Extract ) to extract the real problem q JB of the jailbreak query Q ′ . Given a jailbreak query Q JB , use Extract(·) to remove the irrelevant parts in the jailbreak query that will have an adverse impact on the output, with the goal of generating a q ′ that does not deviate from the semantics of the original problem. It can be expressed as:

[0048] q ′ ~Extract(Q JB ). (2) In a specific implementation, Extract can be implemented as an instruction through a prompt. Specifically,

[0049] Extract(Q JB ) = LLM(P Extract ⊕ Q JB ), (3)

[0050] where P Extract is a soft extraction prompt used to elicit the real problem. Stage 4 shows the specific application of P Extract :

[0051]

[0052] The prompt (P Extract ) may include target statements such as, "Question (excluding user bias):", which is used to indicate that the output of the large language model must include the target statement to meet the expectation.

[0053] Step 202, remove non-critical information from the first question to obtain a second question.

[0054] In this embodiment, after the soft extraction stage, the real intention of the jailbreak attack with obvious intention can be extracted, but soft extraction does not play a good role in jailbreak attacks with unclear intention and complex prompts, thus introducing a hard deletion part.

[0055] Non-critical information can be some words without actual meaning such as modal particles, e.g., "right", "right?", "wow".

[0056] Non-critical information can also be filtered out by setting a filter word list.

[0057] Optionally, a trained keyword extraction model (e.g., a named entity recognition model) can be used to extract keywords and filter out non-critical information.

[0058] Step 203, input the second question together with a preset security prompt into the large language model to output an answer.

[0059] In this embodiment, the security prompt is used to instruct the large language model to comply with the security policy when answering questions.

[0060] Use the regenerated real problem q ′ Instead of the original question, the final response y is generated from the Large Language Model (LLM),

[0061] y~LLM(P Security ⊕q ′ ), (4)

[0062] Safety Tips P Security It is used to ensure that the final response strictly adheres to the security policy, thereby ensuring that any unsafe information is excluded. Phase 5 shows the P Security Details of the content.

[0063]

[0064] The method provided by the above-mentioned embodiment of the present disclosure proposes a two-stage method for revealing the true intent, namely "true intent defense", which includes a "soft extraction" stage and a "hard deletion" stage. The former stage uses a large language model to extract unbiased and real questions through prompt engineering. The "hard deletion" stage removes the least important part of the sentence. Finally, the extracted real questions are input into the target LLM together with the security prompt to generate a response. Given that these questions usually contain fewer tags and have clear intentions, the target LLM can easily defend against them.

[0065] In some optional implementations of the present embodiment, the first question includes a target sentence; and the deleting non-critical information from the first question to obtain the second question includes: marking the words in the first question to obtain a first tag sequence; deleting a predetermined number of tags from the first tag sequence in any combination to obtain at least one second tag sequence; inputting the text corresponding to at least one second tag sequence into a large language model together with the prompt to obtain a probability of generating the target sentence; and determining a second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the second question.

[0066] Due to the inherent limitations of the large language model (LLM), when extracting the true intent from long-text jailbreak attack queries, the soft extraction framework often inadvertently includes irrelevant information and fails to reliably mine the actual intent. To alleviate this problem, this application extracts the true intent by removing the tokens in the jailbreak attack that have the least impact on the prediction results. In the soft extraction framework, when the jailbreak query Q JB When input into the LLM, the entire query Q can be expressed as:

[0067] Q=P Extract⊕Q JB (5)

[0068]

[0069] The target output T of the function LLM(Q) is the real question, i.e., “Question (excluding user bias): [real question]”. The question can be tokenized by a tokenization algorithm. LLM is a sequence of tokens x. 1:n (where each x i is a mapping from the elements of the set Q, n is the number of tokens in the entire query Q) to the probability distribution of possible subsequent tokens. Specifically, for any x n+1 ∈T, use the following notation to represent the token x in front of a given 1:n In this case, the next token is x n+1 Probability:

[0070] p(x n+1 |x 1:n ), (6)

[0071] Therefore, writing To represent the generation of the sequence x given all the tags up to that point n+1:n+t The probability of each individual tag in is:

[0072]

[0073] Where t represents the size of the target output T. This application is concerned with generating a target sentence (e.g., the phrase "Question (excluding user bias):") after deleting some words from the question. probability.

[0074] A predetermined number of tokens are deleted each time, and the predetermined number can be one or more. Each token can represent one or more words. For example, 3 tokens can be deleted each time, and there can be multiple combinations of deletion methods, such as deleting the first 3 tokens, deleting the last 3 tokens, etc. Each deletion method can obtain a second token sequence.

[0075] The second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold is determined as the second question. If there is still a high probability of outputting the target sentence "Question (excluding user bias):" after deleting some words, it means that the deleted words will not affect the final result, and these non-critical information can be deleted to optimize the problem.

[0076] In some optional implementations of the present embodiment, the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold is determined as the second problem, including: calculating the loss value of each second tag sequence based on the negative logarithm of the probability of each second tag sequence generating the target sentence; and determining the second tag sequence with the smallest loss value as the second problem.

[0077] The jailbreak query loss that this application is concerned with is some target tag sequence (i.e., the negative log probability of the phrase "Question (excluding user bias):"):

[0078]

[0079] Therefore, optimizing the jailbreak query task can be written as an optimization problem:

[0080]

[0081] where x i ∈Q JB Indicates that during the optimization process, only Q JB part is optimized, and P Extract Some are not optimized.

[0082] In some optional implementations of the present embodiment, the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold is determined as the second problem, including: calculating the loss value of each second tag sequence according to the negative logarithm of the probability of each second tag sequence generating the target sentence; calculating the gradient of each tag based on the loss value of each second tag sequence; and deleting the tag with the smallest gradient from the first tag sequence to obtain the second problem.

[0083] To optimize objective (9), we can perform the optimization over a series of discrete inputs. The motivation for this approach comes from the greedy gradient-based approach: if we can evaluate every token in the query, we can maximize the removal of the tokens that have the least impact on the query, which will allow us to use fewer tokens to represent similar semantics. Therefore, we can use the gradient associated with the one-hot token indicator to evaluate the least important token in the query. Specifically, we use forward propagation to compute a linearized approximation of the ith token in the hint, x i The gradient of is expressed as:

[0084]

[0085] in represents the one-hot vector (i.e. a vector with 1 in one position and 0 in other positions) representing the current value of the i-th token. Note that since large language models typically form embeddings for each token, they can be written as , so we can immediately take the gradient of this quantity. We select the first p negative values ​​with the largest gradients as markers to be deleted and remove them from the original query. We call this complete method greedy gradient deletion. Stages 7 and 8 illustrate the optimized query Q in this context. ′ And the real question in the end is q ′ .

[0086]

[0087] In some optional implementations of this embodiment, the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold is determined as the second question, including: marking the words in the second question to obtain a third tag sequence; deleting a predetermined number of tags from the third tag sequence in any combination to obtain at least one fourth tag sequence; inputting the text corresponding to at least one fourth tag sequence into the large language model together with the prompt to obtain the probability of generating the target sentence; and determining the fourth tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the updated second question. Unimportant words in the question can be deleted multiple times to further optimize the question and improve the accuracy of the large language model in answering the question.

[0088] In some optional implementations of this embodiment, the outputting of the answer includes: in response to detecting that the second question does not comply with the security policy, outputting information of refusing to answer. When it is detected that the actual intention of the user is to attack, refusing to answer the question can defend against the attack.

[0089] Continue to see Figure 3 , Figure 3 FIG. 1 is a schematic diagram of an application scenario of the large language model method according to this embodiment. Figure 3 In the application scenario, during the soft extraction phase, the user submits a jailbreak query Q to the server through the terminal device. JB The server will jailbreak query Q JB and soft extraction hint P Extract The two words are input into the big language model together, and the real question Q is extracted through the big language model. Then in the hard deletion stage, the unimportant words in the real question are deleted to obtain the optimized real question Q'. Then the optimized real question Q' and the security prompt P are combined. Security The large language model will input the real questions together, and the large language model will determine that the real questions do not meet the security policy and refuse to answer them. This can resist jailbreak attacks.

[0090] Further references Figure 4 , which shows a process 400 of another embodiment of the method for applying a large language model. The process 400 of the method for applying a large language model includes the following steps:

[0091] Step 401: input the original question to be queried together with the preset prompt into the large language model, and output the first question.

[0092] Step 402, deleting non-critical information from the first question to obtain a second question.

[0093] Steps 401-402 are substantially the same as steps 201-202, and thus will not be described in detail.

[0094] Step 403: input the second question together with the preset prompt into the large language model, and output the third question.

[0095] In this embodiment, the optimized second question is input into the large language model together with the preset prompt to extract the real question again, which can improve the accuracy of intent recognition and obtain a more accurate question.

[0096] Step 404: Input the third question together with the preset safety prompt into the large language model, and output the answer.

[0097] In this embodiment, the difference from step 203 is that the second question in step 203 is replaced by the third question, so it is not described again.

[0098] The solution of this application can fully reveal the user's true intention, making it difficult to bypass the defense mechanism, and the attacker cannot trigger the unsafe behavior of the model through carefully constructed input. At the same time, when processing complex and ambiguous inputs, a large amount of computing resources is not required for intent recognition and security checks, which not only reduces the computing cost, but also improves the response efficiency of the model.

[0099] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for applying a large language model. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0100] like Figure 5 As shown, the large language model application device 500 of this embodiment includes: an extraction unit 501, a deletion unit 502 and a questioning unit 503. The extraction unit 501 is configured to input the original question to be queried together with a preset prompt into the large language model, and output a first question, wherein the prompt is used to instruct the large language model to extract a question that reflects the true intention from the original question; the deletion unit 502 is configured to delete non-critical information from the first question to obtain a second question; the questioning unit 503 is configured to input the second question together with a preset security prompt into the large language model, and output an answer, wherein the security prompt is used to instruct the large language model to comply with security policies when answering questions.

[0101] In this embodiment, the specific processing of the extraction unit 501, the deletion unit 502 and the questioning unit 503 of the large language model application device 500 can be referred to. Figure 2 Corresponding to step 201, step 202, and step 203 in the embodiment.

[0102] In some optional implementations of this embodiment, the first question includes a target sentence; and the deletion unit 502 is further configured to: mark the words in the first question to obtain a first tag sequence; delete a predetermined number of tags from the first tag sequence in any combination to obtain at least one second tag sequence; input the text corresponding to the at least one second tag sequence into the large language model together with the prompt to obtain the probability of generating the target sentence; determine the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the second question.

[0103] In some optional implementations of this embodiment, the deletion unit 502 is further configured to: calculate the loss value of each second tag sequence according to the negative logarithm of the probability of each second tag sequence generating the target sentence; and determine the second tag sequence with the smallest loss value as the second question.

[0104] In some optional implementations of this embodiment, the deletion unit 502 is further configured to: calculate the loss value of each second tag sequence based on the negative logarithm of the probability of each second tag sequence generating the target sentence; calculate the gradient of each tag based on the loss value of each second tag sequence; delete the tag with the smallest gradient from the first tag sequence to obtain the second question.

[0105] In some optional implementations of this embodiment, the deletion unit 502 is further configured to: mark the words in the second question to obtain a third tag sequence; delete a predetermined number of tags from the third tag sequence in any combination to obtain at least one fourth tag sequence; input the text corresponding to at least one fourth tag sequence into the large language model together with the prompt to obtain the probability of generating the target sentence; determine the fourth tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the updated second question.

[0106] In some optional implementations of this embodiment, the device also includes a loop unit (not shown in the drawings), which is configured to: input the second question together with a preset prompt into a large language model, and output a third question; input the third question together with a preset security prompt into the large language model, and output an answer.

[0107] In some optional implementations of this embodiment, the questioning unit 503 is further configured to: in response to detecting that the second question does not comply with the security policy, output information of refusing to answer.

[0108] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, storage and other aspects of user personal information involved in the technical solution of this disclosure are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information, network security and national security.

[0109] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0110] An electronic device comprises: one or more processors; a storage device on which one or more computer programs are stored, and when the one or more computer programs are executed by the one or more processors, the one or more processors implement the method described in process 200 or 400.

[0111] A computer-readable medium stores a computer program, wherein the computer program implements the method described in process 200 or 400 when executed by a processor.

[0112] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0113] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0114] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0115] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the road area planning method. For example, in some embodiments, the road area planning method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the road area planning method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the road area planning method in any other appropriate manner (e.g., by means of firmware).

[0116] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0117] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0118] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0119] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0120] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0121] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a server of a distributed system, or a server combined with a blockchain. The server may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. The server may be a server of a distributed system, or a server combined with a blockchain. The server may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0122] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0123] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for applying a large language model, comprising: Inputting the original question to be queried and a preset prompt into the large language model, and outputting a first question, wherein the prompt is used to instruct the large language model to extract a question that reflects the true intention from the original question, wherein the first question includes a target sentence; Marking the words in the first question to obtain a first mark sequence; Deleting a predetermined number of tags from the first tag sequence in any combination to obtain at least one second tag sequence; Inputting text corresponding to at least one second tag sequence together with the prompt into the large language model to obtain a probability of generating the target sentence; Determining a second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as a second question; The second question and a preset security prompt are input into the large language model, and an answer is output, wherein the security prompt is used to instruct the large language model to comply with the security policy when answering the question.

2. The method according to claim 1, wherein: The step of determining the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the second question includes: Calculating the loss value of each second token sequence according to the negative logarithm of the probability of each second token sequence generating the target sentence; The second labeling sequence with the smallest loss value is determined as the second problem.

3. The method according to claim 1, wherein: The step of determining the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the second question includes: Calculating the loss value of each second token sequence according to the negative logarithm of the probability of each second token sequence generating the target sentence; Calculate the gradient of each mark based on the loss value of each second mark sequence; The mark with the smallest gradient is deleted from the first mark sequence to obtain the second problem.

4. The method according to claim 1, wherein: The step of determining the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the second question includes: Mark the words in the second question to obtain a third mark sequence; Deleting a predetermined number of tags from the third tag sequence in any combination to obtain at least one fourth tag sequence; Inputting text corresponding to at least one fourth tag sequence together with the prompt into the large language model to obtain a probability of generating the target sentence; A fourth token sequence whose probability of generating the target sentence is greater than a predetermined threshold is determined as an updated second question.

5. The method according to claim 1, wherein: The method further comprises: Input the second question and the preset prompt into the large language model, and output the third question; The third question and a preset safety prompt are input into the large language model, and an answer is output.

6. The method according to claim 1, wherein: The output answer includes: In response to detecting that the second question does not comply with the security policy, outputting information that the answer is rejected.

7. A device for applying a large language model, comprising: an extraction unit, configured to input an original question to be queried together with a preset prompt into the large language model, and output a first question, wherein the prompt is used to instruct the large language model to extract a question that reflects the true intention from the original question, wherein the first question includes a target sentence; a deleting unit configured to mark the words in the first question to obtain a first tag sequence; delete a predetermined number of tags from the first tag sequence in any combination to obtain at least one second tag sequence; input text corresponding to the at least one second tag sequence together with the prompt into the large language model to obtain a probability of generating the target sentence; determine the second tag sequence whose probability of generating the target sentence is greater than a predetermined threshold as the second question; The questioning unit is configured to input the second question together with a preset security prompt into the large language model and output an answer, wherein the security prompt is used to instruct the large language model to comply with the security policy when answering the question.

8. An electronic device comprising: one or more processors; a storage device having one or more computer programs stored thereon, When the one or more computer programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Request processing method and device, electronic equipment and storage medium

    CN116955760A

  • Fuzzy test-based large language model vulnerability detection method and device

    CN117370994A