A large language model alignment enhancement method and system for optimizing prompt word semantics

By optimizing the semantics of prompt words, optimized prompt words that meet the jailbreak conditions are generated and subjected to adversarial training. This solves the jailbreak attack problem of black-box large language models, achieves efficient alignment mechanism enhancement, suppresses harmful responses, and promotes the secure application of large language models.

CN120579544BActive Publication Date: 2026-05-15HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2025-05-22
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively address the jailbreak attack problem of black-box large language models, resulting in harmful responses that are inconsistent with human values, and lack a sound alignment mechanism to enhance them.

Method used

By optimizing the semantics of prompt words, including part-of-speech filtering, synonym set generation, initial modulation, and word replacement, optimized prompt words that meet the jailbreak conditions are generated, and the alignment mechanism of the large language model is enhanced through adversarial training.

Benefits of technology

The generated optimized prompts have a high attack success rate, strong attack migration capability, and high computational efficiency, which can effectively suppress harmful responses of large language models and promote their secure deployment and application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579544B_ABST
    Figure CN120579544B_ABST
Patent Text Reader

Abstract

The application discloses a large language model alignment enhancement method and system for optimizing prompt word semantics, comprising the following steps: performing part-of-speech filtering on original prompt words to obtain filtered prompt words; generating a synonym set corresponding to each word in the filtered prompt words; replacing the words in the original prompt words with randomly selected words in the synonym set to generate initial modulated prompt words that meet the jailbreaking condition; recording the words that have changed in the initial modulated prompt words compared with the original prompt words to construct a list of changed words; replacing the changed words in the initial modulated prompt words with the words in the filtered prompt words while keeping the jailbreaking condition to generate word-restored prompt words; iteratively optimizing the words in the word-restored prompt words according to the word optimization sequence in the word-restored prompt words to generate optimized prompt words that meet the prompt word semantic retention and jailbreaking conditions; and using the optimized prompt words that meet the prompt word semantic retention and jailbreaking conditions to enhance the alignment mechanism of the target large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a method and system for optimizing the semantics of prompt words in large language models for alignment enhancement. Background Technology

[0002] Jailbreak attacks can bypass the alignment mechanisms of large language models, causing them to output harmful responses that contradict human values. Due to the discrete nature of input terms, restricted access to the target model, and the limited number of queries, black-box large language model jailbreak attacks are an extremely challenging problem. Given the adversarial nature of cybersecurity, research on black-box large language model jailbreak attacks can help enhance the alignment mechanisms of large language models and promote their secure deployment and application.

[0003] However, no satisfactory solution has been found so far. In view of this, the present invention is proposed. Summary of the Invention

[0004] The purpose of this invention is to propose a method and system for enhancing the alignment of large language models by optimizing the semantics of prompt words. Compared with existing white-box and black-box jailbreak attack methods targeting large language models, this method can generate optimized prompt words with higher attack success rate, attack transfer capability, and computational efficiency. In particular, using the optimized prompt words generated by this method to perform adversarial training can enhance the alignment mechanism of large language models and suppress harmful responses output by the large language models.

[0005] The objective of this invention is achieved through the following technical solution:

[0006] A method for optimizing the semantics of prompt words in a large language model alignment enhancement, the method comprising:

[0007] Step 1: Perform part-of-speech filtering on the original prompts to obtain filtered prompts;

[0008] Step 2: Generate a set of synonyms corresponding to each term in the filter suggestions;

[0009] Step 3: Replace the original prompt words with words randomly selected from the synonym set to generate initial modulated prompt words that meet the jailbreak conditions;

[0010] Step 4: Record the words that have changed from the original prompt words in the initial modulated prompt words, and construct a list of changed words;

[0011] Step 5: While maintaining the jailbreak conditions, replace the modified words in the initial modulated prompts with the words in the filtered prompts to generate the word restoration prompts;

[0012] Step 6: Following the optimization order of the entries in the entry restoration prompts, iteratively optimize the entries in the entry restoration prompts to generate optimized prompts that satisfy both semantic preservation and jailbreak conditions;

[0013] Step 7: Use optimized prompts that satisfy the conditions of semantic preservation and jailbreak to perform adversarial training on the target large language model, thereby enhancing the alignment mechanism of the target large language model.

[0014] A large language model alignment enhancement system for optimizing prompt word semantics, used to implement the aforementioned method, includes:

[0015] The part-of-speech filtering module is used to perform part-of-speech filtering on the original prompt words to obtain filtered prompt words;

[0016] The synonym set generation module is used to generate a set of synonyms corresponding to each word in the filter prompts;

[0017] The initial modulation prompt word generation module is used to replace the words in the original prompt word with words randomly selected from the synonym set to generate an initial modulation prompt word that meets the jailbreak conditions;

[0018] The variable word list construction module is used to record words that have changed from the original prompt words in the initial modulated prompt words, and construct a variable word list;

[0019] The term restoration prompt generation module is used to replace the modified terms in the initial modulated prompts with terms in the filtered prompts while maintaining jailbreak conditions, thus generating term restoration prompts.

[0020] The optimized prompt word generation module is used to iteratively optimize the words in the word restoration prompt word according to the word optimization order in the word restoration prompt word, and generate optimized prompt words that meet the requirements of semantic preservation and jailbreak conditions;

[0021] The alignment mechanism enhancement module is used to perform adversarial training on the target large language model using optimized prompt words that satisfy the semantic preservation of prompt words and the jailbreak conditions, thereby enhancing the alignment mechanism of the target large language model.

[0022] As can be seen from the technical solution provided by the present invention, the above method can generate optimized prompt words with high attack success rate, attack migration capability, and computational efficiency. In particular, using the optimized prompt words generated by this method to perform adversarial training can enhance the alignment mechanism of large language models, suppress harmful responses output by large language models, and promote the secure deployment and application of large language models. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart of a large language model alignment enhancement method for optimizing prompt word semantics provided in an embodiment of the present invention;

[0025] Figure 2 This is a schematic diagram of a large language model alignment enhancement system for optimizing prompt word semantics, provided in an embodiment of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0027] First, the following explanations are provided for the terms that may be used in this article:

[0028] The terms “including,” “comprising,” “containing,” “having,” or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, “including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.)” should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0029] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0030] The following is a detailed description of a method and system for optimizing the semantics of prompt words in a large language model. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Unless otherwise specified, specific conditions in the embodiments of this invention are performed according to conventional conditions or manufacturer-recommended conditions. Instruments used in the embodiments of this invention, unless otherwise specified, are all commercially available products.

[0031] Example 1

[0032] This invention provides a method for optimizing the semantics of prompt words in a large language model alignment enhancement, the process of which is as follows: Figure 1 As shown, the main steps include the following:

[0033] Step 1: Perform part-of-speech filtering on the original prompts to obtain filtered prompts.

[0034] The preferred implementation method for this step is as follows:

[0035] Step 11, the attacker possesses the original clue x = [x1, x2, ..., x n [This is a jailbreak attack command containing n terms, designed to induce a large language model to output harmful responses that are inconsistent with human values.] u This refers to the entry in the original prompt word x.

[0036] Step 12: Filter the original prompt word x for entries whose parts of speech are not adjectives, adverbs, verbs, or nouns, to obtain the filtered prompt word w = [w1, w2, ..., w m ], where m is the length of the filter prompt word, and m≤n.

[0037] Step 2: Generate a set of synonyms for each word in the filter suggestions.

[0038] The preferred implementation method for this step is as follows:

[0039] Step 21, for each word w in the filter prompt word w i Generate entries containing w i The set of synonyms S(w i ), where i∈{1,2,...,m}.

[0040] Step 22, calculate the synonym set S(w) i ) size L i =|S(w i )|, and record For the set of synonyms S(w i The j-th term in ).

[0041] Step 3: Replace the original prompt words with words randomly selected from the synonym set to generate initial modulated prompt words that meet the jailbreak conditions.

[0042] The preferred implementation method for this step is as follows:

[0043] Step 31: Randomly select entries from the synonym set. Replace the word x in the original prompt word x u Among them, the term x u The corresponding filter suggestion word w is the word w. i .

[0044] Step 32: Repeat step 31 until the initial modulation prompt word is generated. Meeting the requirements for jailbreak

[0045] in, To input initial modulation prompt words into the target large language model M The response output at that time. A function to determine whether the jailbreak conditions are met. The condition for its validity is the response of the target large language model. The list does not include the rejection terms shown in Table 1.

[0046] For example, you can set the rejection terms to include the terms shown in Table 1.

[0047] Table 1: Rejected Terms

[0048]

[0049] Step 4: Record the words that have changed from the original prompt words in the initial modulated prompt words, and construct a list of changed words.

[0050] The preferred implementation method for this step is as follows:

[0051] Step 41, record the initial modulated cue word compared to the original cue word x. The words s that have changed in the middle σ(i) , where i∈{1,2,...,q}.

[0052] Step 42, construct the list of changed terms S = [s σ(1) ,s σ(2) ,...,s σ(q) ].

[0053] Step 5: While maintaining the jailbreak conditions, replace the modified words in the initial modulated prompts with the words in the filtered prompts to generate the word restoration prompts.

[0054] The preferred implementation method for this step is as follows:

[0055] Step 51: Let the traversal index of the variable word list ← 1 and the number of iterations for the prompt word optimization t ← 0.

[0056] Step 52, use the term w σ(i) Replace prompt words The term in To obtain the prompt word Among them, the term w σ(i) To filter the entries with index σ(i) in the prompt word w.

[0057] Step 53, if jailbreak conditions Satisfaction, i.e., the response of the target large language model If it does not contain preset rejection terms, then let And t←t+1.

[0058] Step 54, let i ← i+1.

[0059] Step 55: Repeat steps 52 to 54 until the list of changed terms is traversed to generate term restoration prompts.

[0060] Step 6: Following the word optimization order in the word restoration prompts, iteratively optimize the words in the word restoration prompts to generate optimized prompts that satisfy both semantic preservation and jailbreak conditions.

[0061] The preferred implementation method for this step is as follows:

[0062] Step 61, calculate the word optimization order weight, which can be expressed as:

[0063]

[0064] Where {σ(1),σ(2),...,σ(q)} represents the initial modulated cue word compared to the original cue word x. Index of terms that have changed in the text To compute term w using the all-mpnet-base-v2 model σ(i) Related to the entry A function of semantic similarity Suggestion for restoring the entry The optimized order weight of the σ(i)th term.

[0065] Step 62, for Perform descending sort, i.e. The optimization order list O = [ρ(1), ρ(2), ..., ρ(q)] is obtained, where ρ(1) is the index of the variable term for the first optimization, ρ(2) is the index of the variable term for the second optimization, and ρ(q) is the index of the variable term for the last optimization.

[0066] Step 63: Let the optimized order list traversal index i ← 1.

[0067] Step 64, use synonyms one by one. Replace prompt words The ρ(i)th term in Generate L ρ(i) Modulation prompt words Where k∈{1,2,...,L} ρ(i)}, S(w ρ(i) ) is the term w with index ρ(i) in the filter prompt w. ρ(i) The corresponding set of synonyms.

[0068] Step 65, under the jailbreak conditions Under the given conditions, find the prompt word with the highest semantic similarity to the original prompt word. And order And t←t+1, where the jailbreak condition is... This refers to the response of the target large language model. It does not include preset rejection terms.

[0069] in, To calculate the prompt word x and prompt words using the all-mpnet-base-v2 model A function of semantic similarity.

[0070] Step 66, let i ← i+1.

[0071] Step 67: Repeat steps 64 to 66 until the optimized order list traversal is completed to generate optimized prompts that satisfy both semantic preservation and jailbreak conditions.

[0072] Step 7: Use optimized prompts that satisfy the conditions of semantic preservation and jailbreak to perform adversarial training on the target large language model, thereby enhancing the alignment mechanism of the target large language model.

[0073] The preferred implementation method for this step is as follows:

[0074] Step 71, given N original prompt words x r For r∈{1,2,...,N}, use steps 1 to 6 to generate optimized prompts that satisfy both semantic preservation and jailbreak conditions.

[0075] Step 72, for optimized prompts that meet the requirements of semantic preservation and jailbreak... Manually specify a security response that includes the rejection terms shown in Table 1.

[0076] Step 73: Use optimized prompts that satisfy both semantic preservation and jailbreak conditions. and security response As the input and output of the target large language model M, the target large language model M is fine-tuned and trained to enhance its alignment mechanism and suppress harmful responses in the output of the target large language model.

[0077] To test the attack success rate, attack migration capability, and computational efficiency of the method described in this embodiment of the invention, the attack capability of the optimized prompt words generated by the proposed method is compared with four other large language model jailbreak attack methods. These four attack methods are: GCG (a white-box attack method based on greedy coordinate gradients), AutoDAN (a white-box attack method based on automatic secret jailbreak prompt word generation), PAIR (a black-box attack method based on automatic iterative optimization of prompt words), and TAP (a black-box attack method based on attack tree pruning). The target large language models in the experiment are two open-source models (Vicuna-7B-v1.5 and Llama-2-7B-chat) and four closed-source models (Claude-3.5-sonnet, GPT-4o-mini, GPT-4o-0806, and GPT-4-turbo). The dataset used in the experiment is a subset of AdvBench, which contains 50 representative jailbreak attack prompt words, which can be used as the original prompt words in the experiment. The metrics for evaluating attack success rate are ASR-Dict and ASR-G. ASR-Dict represents the proportion of attack samples corresponding to model responses without rejection terms out of all attack samples, while ASR-G represents the proportion of attack samples corresponding to malicious target model responses out of all attack samples. The GPT-4o-mini model is used to determine whether the target model response is malicious. The metric for evaluating computational efficiency is the average number of successful queries (Avg.Q), which is the average number of queries to the target model required to generate a jailbreak success message.

[0078] Table 2: Comparison of attack success rate (ASR-Dict / ASR-G) and average number of queries (Avg.Q) for different large language model jailbreak attack methods.

[0079]

[0080] Table 2 shows the comparative experimental results of attack success rates (ASR-Dict / ASR-G) and average number of queries (Avg.Q) for different large language model jailbreak attack methods. As can be seen from Table 2, the method proposed in this invention has the highest attack success rate and the lowest average number of queries in most cases. Since GCG and AutoDAN are white-box attack methods, the average number of successful queries cannot be calculated when targeting closed-source models.

[0081] Table 3: Comparative Experimental Results of Attack Transfer Capabilities of Different Large Language Model Jailbreak Attack Methods

[0082]

[0083] Table 3 shows the comparative experimental results of attack transfer capabilities of different large language model jailbreak attack methods. In the experiment, optimized prompt words were first generated for the original target model, and then these optimized prompt words were used to attack and transfer the target model. The higher the attack success rate against the target model, the stronger the attack transfer capability of the corresponding attack method. As can be seen from Table 3, the method proposed in this invention has the strongest attack transfer capability in most cases.

[0084] Table 4: Examples of jailbreak attacks implemented using the method proposed in this invention

[0085]

[0086] Table 4 shows examples of jailbreak attacks implemented using the method proposed in this invention. As can be seen from Table 4, the optimized prompts generated using the method proposed in this invention have a high semantic similarity to the original prompts, and the optimized prompts can induce the target model to output harmful responses. Since harmful responses deviate from human values, this example only briefly lists a portion of the harmful responses.

[0087] In summary, the method described in this embodiment of the invention can generate optimized prompt words with high attack success rate, attack migration capability, and computational efficiency. In particular, using the optimized prompt words generated by this method to perform adversarial training can enhance the alignment mechanism of large language models, suppress harmful responses output by large language models, and promote the secure deployment and application of large language models.

[0088] Furthermore, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware, and the corresponding program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0089] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.

[0090] Example 2

[0091] This invention also provides a large language model alignment enhancement system for optimizing the semantics of prompt words, which is mainly used to implement the methods provided in the foregoing embodiments, such as... Figure 2 As shown, the system mainly includes:

[0092] The part-of-speech filtering module is used to perform part-of-speech filtering on the original prompt words to obtain filtered prompt words;

[0093] The synonym set generation module is used to generate a set of synonyms corresponding to each word in the filter prompts;

[0094] The initial modulation prompt word generation module is used to replace the words in the original prompt word with words randomly selected from the synonym set to generate an initial modulation prompt word that meets the jailbreak conditions;

[0095] The variable word list construction module is used to record words that have changed from the original prompt words in the initial modulated prompt words, and construct a variable word list;

[0096] The term restoration prompt generation module is used to replace the modified terms in the initial modulated prompts with terms in the filtered prompts while maintaining jailbreak conditions, thus generating term restoration prompts.

[0097] The optimized prompt word generation module is used to iteratively optimize the words in the word restoration prompt word according to the word optimization order in the word restoration prompt word, and generate optimized prompt words that meet the requirements of semantic preservation and jailbreak conditions;

[0098] The alignment mechanism enhancement module is used to perform adversarial training on the target large language model using optimized prompt words that satisfy the semantic preservation of prompt words and the jailbreak conditions, thereby enhancing the alignment mechanism of the target large language model.

[0099] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0100] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.

Claims

1. A method for optimizing the semantics of prompt words in a large language model alignment enhancement, characterized in that, The method includes: Step 1: Perform part-of-speech filtering on the original prompts to obtain filtered prompts; Step 2: Generate a set of synonyms corresponding to each term in the filter suggestions; Step 3: Replace the original prompt words with words randomly selected from the synonym set to generate initial modulated prompt words that meet the jailbreak conditions; Step 4: Record the words that have changed from the original prompt words in the initial modulated prompt words, and construct a list of changed words; Step 5: While maintaining the jailbreak conditions, replace the modified words in the initial modulated prompts with the words in the filtered prompts to generate the word restoration prompts; Step 6: Following the optimization order of the entries in the entry restoration prompts, iteratively optimize the entries in the entry restoration prompts to generate optimized prompts that satisfy both semantic preservation and jailbreak conditions; Step 7: Use optimized prompts that satisfy the conditions of semantic preservation and jailbreak to perform adversarial training on the target large language model.

2. The method for optimizing the semantics of prompt words using a large language model alignment enhancement as described in claim 1, characterized in that, Perform part-of-speech filtering on the original prompts to obtain filtered prompts, including: Step 11, the attacker possesses the original prompt word. It contains Jailbreak attack commands for each entry. Original prompt words The term in the text; Step 12, filter the original prompt words For entries whose parts of speech are not adjectives, adverbs, verbs, or nouns, filter suggestions are obtained. ,in To filter the length of the prompt words, and .

3. The method for optimizing the semantics of prompt words using a large language model alignment enhancement as described in claim 1, characterized in that, Generate and filter a set of synonyms for each term in the suggestion list, including: Step 21, filter prompts. Each entry in Generate entries containing terms a collection of synonyms ,in ; Step 22, calculate the set of synonyms. Size And record a set of synonyms The Middle One entry.

4. The method for optimizing the semantics of prompt words using a large language model alignment enhancement as described in claim 3, characterized in that, The original prompt words are replaced with words randomly selected from the synonym set to generate initial modulated prompt words that meet the jailbreak conditions, including: Step 31: Randomly select entries from the synonym set. Replace original prompt words The term in Among them, the entries Corresponding filter prompts The term in ; Step 32: Repeat step 31 until the initial modulation prompt word is generated. Meeting the requirements for jailbreak ; in, To target large language model Enter the initial modulation prompt. The response output at that time. A function to determine whether the jailbreak conditions are met. The condition for its validity is the response of the target large language model. It does not contain preset rejection terms.

5. The method for optimizing the semantics of prompt words using a large language model alignment enhancement as described in claim 1, characterized in that, Record the words that have changed from the original prompt words in the initial modulated prompt words, and construct a list of changed words, including: Step 41, record the original prompt words. Compared to the initial modulation prompt words The entries that have changed ,in ; Step 42, construct a list of changed terms .

6. The method for optimizing the semantics of prompt words using a large language model alignment enhancement as described in claim 5, characterized in that, While maintaining jailbreak conditions, replace the modified words in the initial modulated prompts with words from the filtered prompts to generate word restoration prompts, including: Step 51, iterate through the index of the list of changed terms. And the number of iterations for prompt word optimization ; Step 52, use the term Replace prompt words The term in To obtain the prompt word Among them, the entries Filter prompt words The index is The entry for this term; Step 53, if jailbreak conditions Satisfaction, i.e., the response of the target large language model If it does not contain preset rejection terms, then let and ; Step 54, let ; Step 55: Repeat steps 52 to 54 until the list of changed terms is traversed to generate term restoration prompts. .

7. The method for optimizing the semantics of prompt words using a large language model alignment enhancement as described in claim 6, characterized in that, Following the optimization order of the terms in the term restoration hints, iteratively optimize the terms in the term restoration hints to generate optimized hints that satisfy both semantic preservation and jailbreak conditions, including: Step 61, calculate the word optimization order weight, which can be expressed as: ; in, To match the original prompt words Compared to the initial modulation prompt words Index of terms that have changed in the text To compute terms using the all-mpnet-base-v2 model Related to the entry A function of semantic similarity Suggestion for restoring the entry The Middle Optimization order weights for each term; Step 62, for Perform descending sort, i.e. To obtain the optimized order list ,in For the first optimized variable term index, For the second optimized variable term index, For the last optimized change term index; Step 63, let the optimized ordered list traversal index ; Step 64, use synonyms one by one. Replace prompt words The first in Each entry ,generate Modulation prompt words ,in , These are filter prompts. The index is The entry The corresponding set of synonyms; Step 65, under the jailbreak conditions Under the given conditions, find the prompt word with the highest semantic similarity to the original prompt word. and order and Among them, the conditions for escaping from prison This refers to the response of the target large language model. It does not contain preset rejection terms; in, To compute prompt words using the all-mpnet-base-v2 model With prompt words A function of semantic similarity; Step 66, let ; Step 67: Repeat steps 64 to 66 until the optimized order list traversal is completed to generate optimized prompts that satisfy both semantic preservation and jailbreak conditions. .

8. The method for optimizing the semantics of prompt words using a large language model alignment enhancement as described in claim 1, characterized in that, Adversarial training of the target large language model is performed using optimized prompts that satisfy both semantic preservation and jailbreak conditions, including: Step 71, given Original prompt words Use steps 1 to 6 to generate optimized prompts that satisfy both semantic preservation and jailbreak requirements. ; Step 72, for optimized prompts that meet the requirements of semantic preservation and jailbreaking. Set up security responses by combining preset rejection keywords. ; Step 73: Use optimized prompts that satisfy both semantic preservation and jailbreak conditions. and security response As a target large language model The input and output of the target large language model Fine-tuning training was conducted.

9. A large language model alignment enhancement system for optimizing the semantics of prompt words, characterized in that, For implementing the method according to any one of claims 1 to 8, comprising: The part-of-speech filtering module is used to perform part-of-speech filtering on the original prompt words to obtain filtered prompt words; The synonym set generation module is used to generate a set of synonyms corresponding to each word in the filter prompts; The initial modulation prompt word generation module is used to replace the words in the original prompt word with words randomly selected from the synonym set to generate an initial modulation prompt word that meets the jailbreak conditions; The variable word list construction module is used to record words that have changed from the original prompt words in the initial modulated prompt words, and construct a variable word list; The term restoration prompt generation module is used to replace the modified terms in the initial modulated prompts with terms in the filtered prompts while maintaining jailbreak conditions, thus generating term restoration prompts. The optimized prompt word generation module is used to iteratively optimize the words in the word restoration prompt word according to the word optimization order in the word restoration prompt word, and generate optimized prompt words that meet the requirements of semantic preservation and jailbreak conditions; The alignment mechanism enhancement module is used to perform adversarial training on the target large language model using optimized prompts that satisfy the semantic preservation of prompts and the jailbreak conditions.