Adversarial sample construction method and system based on voice generation type large model

Discrete speech representations are generated through self-supervised training, cross-modal instruction data sets are constructed and a three-stage training strategy is implemented to generate adversarial samples, which solves the problems of knowledge transfer difficulties and semantic information loss between modals of the speech generation model, improves the adversarial robustness and security of the model, and is suitable for the practical application of multimodal large language models.

CN120260549APending Publication Date: 2025-07-04ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510386526.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing speech generation model faces problems such as inter-modal knowledge transfer difficulties, emotional and tone information loss when processing continuous signals, and cannot understand semantic information, which limits the ability to cross-modal perception and generation. At the same time, existing adversarial samples are difficult to maintain naturalness and are easily recognized by defense mechanisms.

Method used

Self-supervised training is used to generate discrete speech representations, build a cross-modal instruction data set and implement a three-stage training strategy to generate adversarial samples through gradient search to ensure their naturalness and robustness.

Benefits of technology

It improves the adversarial robustness and security of the speech generation model, enhances the reliability and stability of the model in practical applications, and is suitable for the wide application of multimodal large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260549A_ABST
    Figure CN120260549A_ABST
Patent Text Reader

Abstract

The invention discloses an adversarial sample construction method and system based on a voice generation type large model, and relates to the technical field of natural language processing and voice processing. Discrete voice representation is generated through self-supervised training, a cross-modal instruction data set is constructed, and a three-stage training strategy is implemented; after the voice generation type large model is trained, an adversarial sample is generated through gradient search; according to the method, the anti-robustness and safety of the model can be improved, and powerful support is provided for practical application of the multi-modal large language model. The generation and defense technology of the confrontation sample can be further explored through future research, and wide application of the multi-mode large language model in more fields is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of natural language processing and speech processing, and more specifically, to an adversarial sample construction method and system based on a speech generative large model. Background Art

[0002] Currently, in recent years, multi-modal large language models (such as SpeechGPT) have unified speech and text modalities by converting speech signals into discrete units through discrete speech representation technology, thereby enhancing the model's perception and generation capabilities for multi-modal content. However, when dealing with continuous signals (such as images and speech), these models still face problems such as difficult knowledge transfer between modalities and loss of emotional and intonation information. In addition, although existing speech generation models can synthesize speech, they cannot understand its semantic information, which limits the true cross-modal perception and generation capabilities.

[0003] Therefore, how to generate effective adversarial samples and maintain the high naturalness of adversarial samples so that they are difficult to be recognized by defense mechanisms is an urgent problem for those skilled in the art to solve. Summary of the Invention

[0004] In view of this, the present invention provides an adversarial sample construction method and system for a speech generative large model, which can not only generate effective adversarial samples, but also maintain the high naturalness of these samples so that they are difficult to be recognized by defense mechanisms. This helps to improve the security and robustness of the model and ensure its reliability and stability in practical applications.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] An adversarial sample construction method based on a speech generative large model includes:

[0007] Encoding the original speech signal using a self-supervised trained speech model to generate discrete speech representations;

[0008] Constructing a cross-modal instruction dataset including speech-text pairs and chained modal instructions;

[0009] Designing and implementing a three-stage training strategy including modal adaptation pre-training, cross-modal instruction fine-tuning, and chained modal instruction fine-tuning;

[0010] Training a speech generative large model based on the three-stage training strategy, discrete speech representations, and cross-modal instruction dataset;

[0011] Generating the best adversarial trigger through a gradient search algorithm based on the trained speech generative large model to obtain adversarial samples.

[0012] Optionally, the generation of discrete speech representations specifically includes:

[0013] Self-supervised training: Learn speech feature representations from unlabeled speech data through self-supervised learning; during the training process, the speech model first segments the speech signal into fixed-length segments, and then learns speech features through self-supervised tasks.

[0014] Vocabulary expansion: Expand the discrete speech representations into the vocabulary of the speech generative large model, enabling the speech generative large model to directly process speech signals.

[0015] Optionally, the construction of the cross-modal instruction dataset includes:

[0016] Data collection: Extract speech-text pairs from existing ASR datasets to generate diverse cross-modal instructions; the ASR dataset contains speech signals and corresponding text transcripts.

[0017] Data annotation: Manually annotate the generated cross-modal instructions to ensure the accuracy and diversity of the cross-modal instructions; during the annotation process, it is necessary to ensure that the format and content of the instructions meet the requirements of the actual application scenario.

[0018] Optionally, the design and implementation of the three-stage training strategy specifically include:

[0019] Modal adaptation pre-training: Pre-train the speech generative large model using discrete speech representations and text data to enable the speech generative large model to adapt to speech and text modalities.

[0020] Cross-modal instruction fine-tuning: Fine-tune the speech generative large model using the constructed cross-modal instruction dataset to enhance the model's understanding and execution ability of multi-modal instructions.

[0021] Chain modal instruction fine-tuning: Further fine-tune the model through the chain modal instruction dataset to improve the model's performance in complex tasks; the chain modal instruction dataset contains instruction sequences of multiple modalities, and the speech generative large model needs to execute tasks step by step according to the instruction sequences.

[0022] Optionally, the generation of the best adversarial trigger through the gradient search algorithm specifically includes:

[0023] Gradient optimization: Search for the best adversarial trigger by calculating the gradient of the model with respect to the input, ensuring that the generated trigger can mislead the target model without affecting the human reading experience.

[0024] Naturalness evaluation: Combine the internal state of the speech generative large model and the external input, and ensure the naturalness and fluency of the generated trigger by dynamically adjusting the trigger generation process.

[0025] Optionally, the speech generation large model further includes an adaptive defense strategy, specifically:

[0026] Semantic naturalness evaluation: By calculating the perplexity and semantic similarity of sentences, identify potential adversarial triggers to reduce false positives and false negatives;

[0027] Defense mechanism optimization: Combine the internal characteristics of the speech generation large model and the external environment, and continuously optimize the defense mechanism to improve the security and robustness of the model.

[0028] Optionally, the obtained adversarial samples are not limited to classification tasks, but also include generation tasks, applicable to a wider range of model architectures and application scenarios. The method for generating adversarial samples can adapt to models of different scales, and can be effectively applied from small models to large models, ensuring the security and robustness of the speech generation large model in different tasks.

[0029] An adversarial sample construction system for a speech generation large model includes:

[0030] Discrete encoding module, which encodes the original speech signal using a self-supervised trained speech model to generate discrete speech representations;

[0031] Dataset construction module, which constructs a cross-modal instruction dataset, including speech-text pairs and chained modal instructions;

[0032] Training strategy design module, which designs and implements a three-stage training strategy, including modal adaptation pre-training, cross-modal instruction fine-tuning, and chained modal instruction fine-tuning;

[0033] Model training module, which trains the speech generation large model based on the three-stage training strategy, discrete speech representations, and cross-modal instruction dataset;

[0034] Adversarial sample generation module, which generates the best adversarial trigger through a gradient search algorithm based on the trained speech generation large model to obtain adversarial samples.

[0035] From the above technical solutions, compared with the prior art, the present invention discloses an adversarial sample construction method and system for a speech generation large model. By generating discrete speech representations through self-supervised training, constructing a cross-modal instruction dataset and implementing a three-stage training strategy, after training the speech generation large model, adversarial samples are generated through gradient search. The present invention can not only improve the adversarial robustness and security of the model, but also provide strong support for the practical application of multi-modal large language models. Future research will further explore the generation and defense technologies of adversarial samples to promote the wide application of multi-modal large language models in more fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0037] Figure 1 It is a schematic flowchart of the method provided by the present invention. Specific embodiments

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0039] The embodiments of the present invention disclose an adversarial sample construction method based on a speech generation large model, as Figure 1 shown, including:

[0040] Encoding the original speech signal using a self-supervised trained speech model to generate discrete speech representations;

[0041] Constructing a cross-modal instruction dataset, including speech-text pairs and chained modal instructions;

[0042] Designing and implementing a three-stage training strategy, including modal adaptation pre-training, cross-modal instruction fine-tuning, and chained modal instruction fine-tuning;

[0043] Training the speech generation large model based on the three-stage training strategy, discrete speech representations, and cross-modal instruction dataset;

[0044] Generating the best adversarial trigger through a gradient search algorithm based on the trained speech generation large model to obtain adversarial samples.

[0045] In a specific embodiment, generating discrete speech representations is specifically:

[0046] Self-supervised training: Encoding the original speech signal using a self-supervised trained speech model (such as HuBERT) to generate discrete speech representations. Specifically, the HuBERT model learns speech feature representations from a large amount of unlabeled speech data through self-supervised learning. During the training process, the model first divides the speech signal into fixed-length segments, and then learns speech features through self-supervised tasks (such as predicting the category of the next speech segment).

[0047] Vocabulary Expansion: Expand the discrete speech representations into the vocabulary of the speech generation large model, enabling the speech generation large model to directly process speech signals. Specifically, map the discrete speech representations into the vocabulary of the speech generation large model to ensure that the speech generation large model can directly receive and generate discrete speech representations. For example, if the vocabulary of the speech generation large model contains 100,000 words, the discrete speech representations can be mapped to a part of these 100,000 words to ensure that speech signals and text signals are represented in the same vocabulary space.

[0048] In a specific embodiment, constructing a cross-modal instruction dataset includes:

[0049] Data Collection: Extract speech-text pairs from existing ASR datasets. These datasets usually contain a large number of speech signals and corresponding text transcripts. Use models such as GPT-4 to generate diverse cross-modal instructions. These instructions can cover various tasks, such as classification, translation, conversation, etc.

[0050] Data Annotation: Manually annotate the generated instructions to ensure the accuracy and diversity of the instructions. During the annotation process, it is necessary to ensure that the format and content of the instructions meet the requirements of the actual application scenario. For example, for a classification task, instructions like "Please judge whether the emotion of this speech is positive or negative" can be generated; for a translation task, instructions like "Please translate this speech into English" can be generated.

[0051] In a specific embodiment, designing and implementing a three-stage training strategy specifically includes:

[0052] Modal Adaptation Pre-training: Pre-train the speech generation large model using discrete speech representations and text data to make the speech generation large model adapt to the speech and text modalities. The pre-training tasks can include prediction of discrete speech representations, text generation, etc. For example, a sequence prediction task of discrete speech representations can be used to train the model to predict the category of the next speech segment; a text generation task can also be used to train the speech generation large model to generate corresponding text according to the given speech segment.

[0053] Cross-modal Instruction Fine-tuning: Fine-tune the speech generation large model using the constructed cross-modal instruction dataset to enhance the model's understanding and execution ability of multi-modal instructions. The fine-tuning tasks can include instruction parsing, task execution, etc. For example, an instruction parsing task can be used to train the speech generation large model to parse out the task type and task parameters according to the given instruction; a task execution task can also be used to train the speech generation large model to execute the corresponding task according to the parsed task type and parameters.

[0054] Chain - based Modal Instruction Fine - Tuning: Further fine - tune the speech - generating large - model through a chain - based modal instruction dataset to improve the performance of the speech - generating large - model in complex tasks. The chain - based modal instruction dataset contains instruction sequences of multiple modalities, and the speech - generating large - model needs to execute tasks step by step according to the instruction sequences. For example, an instruction sequence like "Please first translate this piece of speech into English, and then determine whether the sentiment of the translated text is positive or negative" can be generated to train the speech - generating large - model to execute tasks step by step.

[0055] In a specific embodiment, the generating of the optimal adversarial trigger through the gradient search algorithm specifically includes:

[0056] Gradient Optimization: Use the gradient search algorithm to generate Universal Adversarial Triggers (UATs) with high naturalness. The gradient search algorithm is specifically LinkPrompt. This algorithm searches for the optimal adversarial trigger through gradient optimization techniques, which can not only effectively attack target pre - trained language models (PLMs) and prompt - based fine - tuned models (PFMs), but also maintain the naturalness and fluency of the trigger; When the LinkPrompt algorithm generates adversarial triggers, it combines the internal state of the speech - generating large - model and external inputs, and ensures that the generated triggers can adapt to different model architectures and task requirements by dynamically adjusting the trigger generation process. For example, the gradient ascent method can be used to gradually adjust the words or speech segments in the input to make the output of the speech - generating large - model change as expected.

[0057] Naturalness Evaluation: Combine the internal state of the speech - generating large - model and external inputs, and ensure that the generated triggers have high naturalness and fluency by dynamically adjusting the trigger generation process. Naturalness evaluation can be achieved by calculating metrics such as the perplexity and semantic similarity of sentences.

[0058] In a specific embodiment, the speech - generating large - model also includes an adaptive defense strategy, specifically:

[0059] Semantic Naturalness Evaluation: Identify potential adversarial triggers by calculating metrics such as the perplexity and semantic similarity of sentences to reduce false positives and false negatives. Specifically, a pre - trained language model can be used to calculate the perplexity of the input sentence. According to the baseline perplexity distribution, a suitable threshold is selected. If the perplexity of the sentence exceeds this threshold, it may contain an adversarial trigger. Semantic similarity can be calculated through a pre - trained word vector model or sentence encoding model. A common method is to calculate the cosine similarity between the embedding vectors of two sentences: where v1 and v2 represent the embedding vectors of sentences S1 and S2 respectively.

[0060] Defense mechanism optimization: Combine the internal features of the model and the external environment to continuously optimize the defense mechanism and improve the security and robustness of the model. Specifically, reinforcement learning techniques can be used to learn from human feedback, continuously adjust the defense strategy, and improve the defense effect of the speech generation large model.

[0061] The adaptive defense strategy detects and resists adversarial sample attacks by evaluating the semantic naturalness of sentences, improving the robustness of the speech generation large model; the defense strategy combines the internal features of the speech generation large model and the external environment, and identifies potential adversarial triggers by calculating metrics such as sentence perplexity and semantic similarity, thereby reducing false positives and false negatives and improving the security of the model.

[0062] The models used include but are not limited to RoBERTa-large, BERT-large-cased, Llama2-7B, and the language model GPT-3.5-turbo accessed through the API, and can be widely applied to the security evaluation and defense of various multimodal large language models; the present invention supports continuous optimization of the speech generation large model, learns from human feedback through reinforcement learning, further improves the concealment and attack effect of adversarial samples, and ensures that the model still has good defense capabilities in the face of new attacks.

[0063] In a specific embodiment, the obtained adversarial samples are not limited to classification tasks, but also include generation tasks such as translation and dialogue, and are applicable to a wider range of model architectures and application scenarios, such as models based on Transformer and models based on attention mechanisms; the method for generating adversarial samples can adapt to models of different scales, and can be effectively applied to small models and large models (such as RoBERTa-large, BERT-large-cased, Llama2-7B, etc.), ensuring the security and robustness of the speech generation large model in different tasks.

[0064] An adversarial sample construction system for a speech generation large model, comprising:

[0065] A discrete encoding module that encodes the original speech signal using a self-supervised trained speech model to generate a discrete speech representation;

[0066] A dataset construction module that constructs a cross-modal instruction dataset, including speech-text pairs and chained modal instructions;

[0067] A training strategy design module that designs and implements a three-stage training strategy, including modal adaptation pre-training, cross-modal instruction fine-tuning, and chained modal instruction fine-tuning;

[0068] A model training module trains a speech generative large model based on a three-stage training strategy, discrete speech representations, and a cross-modal instruction dataset;

[0069] An adversarial sample generation module generates the best adversarial trigger through a gradient search algorithm based on the trained model to obtain adversarial samples.

[0070] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, please refer to the description in the method section.

[0071] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for constructing adversarial samples for speech generation large models, characterized in that, Including: Encoding the original speech signal using a self-supervised trained speech model to generate discrete speech representations; Constructing a cross-modal instruction dataset, including speech-text pairs and chained-modal instructions; Designing and implementing a three-stage training strategy, including modal adaptation pre-training, cross-modal instruction fine-tuning, and chained-modal instruction fine-tuning; Training a speech generative large model based on the three-stage training strategy, discrete speech representations, and cross-modal instruction dataset; Based on the trained speech generative large model, generating the best adversarial trigger through a gradient search algorithm to obtain adversarial samples.

2. The method for constructing adversarial samples based on a speech generation large model according to claim 1, wherein The specific generation of discrete speech representations is as follows: Self-supervised training: Learning speech feature representations from unlabeled speech data through self-supervised learning; during the training process, the speech model first segments the speech signal into fixed-length segments and then learns speech features through self-supervised tasks; Vocabulary expansion: Expanding the discrete speech representations into the vocabulary of the speech generative large model, enabling the speech generative large model to directly process speech signals.

3. The method for constructing adversarial samples based on a speech generation large model according to claim 1, wherein, The construction of the cross-modal instruction dataset includes: Data collection: Extracting speech-text pairs from existing ASR datasets to generate diverse cross-modal instructions; the ASR dataset contains speech signals and corresponding text transcripts; Data annotation: Manually annotating the generated cross-modal instructions to ensure the accuracy and diversity of the cross-modal instructions; during the annotation process, it is necessary to ensure that the format and content of the instructions meet the requirements of the actual application scenario.

4. A method for constructing adversarial samples based on a speech generation large model according to claim 1, characterized in that The specific design and implementation of the three-stage training strategy include: Modal adaptation pre-training: Pre-training the speech generative large model using discrete speech representations and text data to enable the speech generative large model to adapt to speech and text modalities; Cross-modal instruction fine-tuning: Fine-tuning the speech generative large model using the constructed cross-modal instruction dataset to enhance the model's understanding and execution ability of multi-modal instructions; Chained-modal instruction fine-tuning: Further fine-tuning the model through the chained-modal instruction dataset to improve the model's performance in complex tasks; the chained-modal instruction dataset contains instruction sequences of multiple modalities, and the speech generative large model needs to execute tasks step by step according to the instruction sequences.

5. A method for constructing adversarial samples based on a speech generation large model according to claim 1, characterized in that, The specific generation of the best adversarial trigger through the gradient search algorithm includes: Gradient optimization: Searching for the best adversarial trigger by calculating the gradient of the model with respect to the input, ensuring that the generated trigger can mislead the target model without affecting the human reading experience; Naturalness evaluation: Combining the internal state and external input of the speech generative large model, and ensuring the naturalness and fluency of the generated trigger by dynamically adjusting the trigger generation process.

6. The method for constructing adversarial samples based on a speech generation large model according to claim 1, wherein, The speech generative large model also includes an adaptive defense strategy, specifically: Semantic naturalness evaluation: Identifying potential adversarial triggers by calculating the perplexity and semantic similarity of sentences to reduce false positives and false negatives; Defense mechanism optimization: Continuously optimizing the defense mechanism by combining the internal features and external environment of the speech generative large model to improve the security and robustness of the model.

7. A method for constructing adversarial examples based on a speech generation large model according to claim 1, characterized in that, The obtained adversarial examples are not limited to classification tasks, but also include generation tasks, and are applicable to a wider range of model architectures and application scenarios. The method for generating adversarial examples can adapt to models of different scales and can be effectively applied to both small and large models, ensuring the security and robustness of large speech generation models in different tasks.

8. An adversarial sample construction system based on a speech generation large model, characterized in that, Applying the method for constructing adversarial examples for large speech generation models according to any one of claims 1-7, comprising: A discrete encoding module that encodes the original speech signal using a self-supervised trained speech model to generate a discrete speech representation; A dataset construction module that constructs a cross-modal instruction dataset, including speech-text pairs and chained modal instructions; A training strategy design module that designs and implements a three-stage training strategy, including modal adaptation pre-training, cross-modal instruction fine-tuning, and chained modal instruction fine-tuning; A model training module that trains a large speech generation model based on the three-stage training strategy, the discrete speech representation, and the cross-modal instruction dataset; An adversarial example generation module that generates the best adversarial trigger through a gradient search algorithm based on the trained large speech generation model to obtain adversarial examples.