A voiceprint comparison method and device based on a large language model and a readable medium

CN120496537BActive Publication Date: 2026-09-29XIAMEN KUAISHANGTONG TECH CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510608005.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2026-09-29
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

但是目前在声纹对比的场景中并没有充分利用大语言模型技术,因此,亟需开发一种基于大语言模型的声纹比对方法

Benefits of technology

[0024](1)本发明提出的基于大语言模型的声纹比对方法将两个不同的语音进行拼接,得到合成语音,再结合提示词输入到基于大语言模型的声纹比对模型中,利用预训练的大语言模型的本体结构的上下文学习能力,识别出合并语音中的两个语音是否是同一人的语音。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496537B_ABST
    Figure CN120496537B_ABST
Patent Text Reader

Abstract

The application discloses a voiceprint comparison method and device based on a large language model and readable medium, comprising: obtaining first speech and second speech to be compared which are respectively collected and spliced into merged speech; inputting the merged speech and a prompt word into a trained voiceprint comparison model, wherein the merged speech is first subjected to an audio encoder to obtain speech coding features; the prompt word is subjected to a text encoder to obtain text coding features; inputting the speech coding features into an adapter to convert the dimensions of the speech coding features to obtain dimensionally converted speech coding features; inputting the text coding features and the dimensionally converted speech coding features into an ontology structure of an improved large language model after splicing; adding output features of the ontology structure of the pre-trained large language model and output features of a LoRA module to obtain an output token sequence; and subjecting the output token sequence to a text decoder to obtain corresponding output text, thereby improving the robustness and accuracy of model algorithm in voiceprint discrimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voiceprint comparison, specifically to a voiceprint comparison method, apparatus, and readable medium based on a large language model. Background Technology

[0002] Voiceprint matching technology is used to compare whether two audio recordings belong to the same person. Traditional voiceprint matching methods rely on signal processing techniques to extract acoustic features, then combine them with methods such as Gaussian mixture models for identity modeling, and use likelihood ratios to complete the voiceprint comparison. Due to its weak feature representation ability, poor model flexibility, and insufficient anti-interference ability, traditional methods have been gradually replaced by deep learning methods. Deep learning methods use neural network models such as x-vector or ECAPA-TDNN for voiceprint matching. However, this method requires collecting a large amount of labeled speech training data to train the neural network model to achieve good recognition results. Otherwise, it is prone to problems with robustness and poor comparison accuracy. In addition, the computational complexity and cost are relatively high.

[0003] With the development of large language model technology, large model techniques based on audio or video have emerged for direct processing of audio or video data, not just text data. However, large language model technology is not fully utilized in current voiceprint comparison scenarios. Therefore, there is an urgent need to develop a voiceprint comparison method based on large language models. Summary of the Invention

[0004] The purpose of this application is to propose a voiceprint comparison method, device, and readable medium based on a large language model to address the aforementioned technical problems.

[0005] In a first aspect, the present invention provides a voiceprint matching method based on a large language model, comprising the following steps:

[0006] The first and second voice samples collected separately are acquired and concatenated into a merged voice sample.

[0007] A voiceprint matching model based on a large language model is constructed and trained to obtain a trained voiceprint matching model. The voiceprint matching model includes an audio encoder, an adapter, a text encoder, an improved ontology structure of the large language model, and a text decoder. The improved ontology structure of the large language model includes a pre-trained ontology structure of the large language model set in parallel and a LoRA module.

[0008] The system acquires prompt words to determine whether the first and second voice samples to be compared belong to the same person by merging the voice samples. The merged voice samples and prompt words are input into a trained voiceprint comparison model. The merged voice samples are first processed by an audio encoder to obtain voice encoding features; the prompt words are processed by a text encoder to obtain text encoding features. The voice encoding features are then input into an adapter to perform dimensionality transformation to obtain dimensionality-transformed voice encoding features. The text encoding features and the dimensionality-transformed voice encoding features are concatenated and input into the ontology structure of the improved large language model. The output features of the pre-trained large language model ontology structure are added to the output features of the LoRA module to obtain an output token sequence. The output token sequence is then processed by a text decoder to obtain the corresponding output text.

[0009] Preferably, the audio encoder includes a pre-trained speech encoder based on a transformer structure or a pre-trained wav2vec model. The pre-trained speech encoder based on a transformer structure includes a feature extraction module and an encoding module of a wishper model.

[0010] Preferably, the adapter includes a first convolutional layer, a first ReLU activation function layer, a second convolutional layer, and a second ReLU activation function layer connected in sequence.

[0011] Preferably, the first and second voice samples to be compared in the merged speech are distinguished by duration; multiple prompt words are set, and one of them is randomly selected each time; the output text includes yes or no, if the output text is yes, it is determined that the first and second voice samples to be compared are the voices of the same person, if the output text is no, it is determined that the first and second voice samples to be compared are the voices of different people.

[0012] As a preferred option, the voiceprint matching model is trained using two supervised training sessions, specifically:

[0013] During the first training process, the parameters of the ontology structure of the pre-trained large language model are frozen, and the parameters of the LoRA module, audio encoder and adapter are adjusted to obtain the voiceprint comparison model after one training.

[0014] In the second training process, the voiceprint comparison model after the first training is fine-tuned. The ontology structure of the pre-trained large language model, the parameters of the LoRA module after the first training, the audio encoder and the adapter are adjusted to obtain the trained voiceprint comparison model.

[0015] As a preferred option, pre-trained large language models include the Thousand Questions Large Model.

[0016] Secondly, the present invention provides a voiceprint comparison device based on a large language model, comprising:

[0017] The voice acquisition module is configured to acquire the first voice and the second voice to be compared, which are collected separately, and then concatenate them into a merged voice.

[0018] The model building module is configured to build and train a voiceprint matching model based on a large language model to obtain a trained voiceprint matching model. The voiceprint matching model includes an audio encoder, an adapter, a text encoder, an improved ontology structure of the large language model, and a text decoder. The improved ontology structure of the large language model includes a pre-trained ontology structure of the large language model set in parallel and a LoRA module.

[0019] The comparison module is configured to acquire prompt words used to determine whether the first and second voice samples to be compared belong to the same person by merging the voice samples. The merged voice samples and prompt words are input into a trained voiceprint comparison model. The merged voice samples are first processed by an audio encoder to obtain voice encoding features; the prompt words are processed by a text encoder to obtain text encoding features; the voice encoding features are input into an adapter to perform dimensionality transformation to obtain dimensionality-transformed voice encoding features; the text encoding features and the dimensionality-transformed voice encoding features are concatenated and input into the ontology structure of the improved large language model; the output features of the pre-trained large language model ontology structure are added to the output features of the LoRA module to obtain an output token sequence; the output token sequence is processed by a text decoder to obtain the corresponding output text.

[0020] Thirdly, the present invention provides an electronic device including one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.

[0021] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the implementations of the first aspect.

[0022] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any of the implementations in the first aspect.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] (1) The speaker matching method based on the large language model proposed in this invention splices two different voices to obtain synthesized voice, and then combines the prompt words into the speaker matching model based on the large language model. By utilizing the context learning ability of the ontology structure of the pre-trained large language model, it can identify whether the two voices in the merged voice are the voices of the same person.

[0025] (2) The voiceprint matching method based on the large language model proposed in this invention utilizes the technical advantages of the large language model to realize the voiceprint matching technology, improve the robustness and accuracy of the model algorithm for voiceprint discrimination, and improve the voiceprint matching effect. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart illustrating the voiceprint comparison method based on a large language model, as an embodiment of this application.

[0028] Figure 2 This is a schematic diagram of the voiceprint matching model of the voiceprint matching method based on a large language model, which is an embodiment of this application.

[0029] Figure 3 This is a schematic diagram of a voiceprint comparison device based on a large language model, which is an embodiment of this application.

[0030] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0032] Figure 1 The present application illustrates an embodiment of a voiceprint matching method based on a large language model, comprising the following steps:

[0033] S1, acquire the first speech and the second speech collected separately and concatenate them into a merged speech.

[0034] Specifically, one of the first and second voice samples to be compared can be from an unknown person, while the other can be from a known person. Therefore, by comparing the voiceprints of the first and second voice samples, it can be determined whether they belong to the same person. If they do, the identity of the unknown person is the same as that of the known person. If both the first and second voice samples belong to an unknown person, voiceprint comparison can still be used to determine if they belong to the same person, and voice samples from the same person can be categorized accordingly.

[0035] S2, Construct and train a voiceprint matching model based on a large language model to obtain a trained voiceprint matching model. The voiceprint matching model includes an audio encoder, an adapter, a text encoder, an improved ontology structure of the large language model, and a text decoder. The improved ontology structure of the large language model includes a pre-trained ontology structure of the large language model set in parallel and a LoRA module.

[0036] In a specific embodiment, the audio encoder includes a pre-trained speech encoder based on a transformer structure or a pre-trained wav2vec model. The pre-trained speech encoder based on a transformer structure includes a feature extraction module and an encoding module of a wishper model.

[0037] In a specific embodiment, the adapter includes a first convolutional layer, a first ReLU activation function layer, a second convolutional layer, and a second ReLU activation function layer connected in sequence.

[0038] In a specific implementation, the pre-trained large language model includes the Thousand Questions Large Model.

[0039] For details, please refer to Figure 2The voiceprint comparison model proposed in this application consists of an audio encoder, an adapter, a text encoder, an improved large language model ontology structure, and a text decoder. The audio encoder encodes the merged speech, converting it into vector data. This audio encoder can be a pre-trained transformer-based speech encoder or a pre-trained WAV2VEC model. Specifically, the pre-trained transformer-based speech encoder can use the feature extraction module and encoding module of the Wishper model, without requiring a Wishper model decoding module. The merged speech sequentially passes through the pre-trained Wishper model's feature extraction module and encoding module, or through the pre-trained WAV2VEC model, to obtain the corresponding vector representation. This vector representation is the speech encoding feature. This speech encoding feature is then input into the adapter to convert it into a format acceptable to the pre-trained large language model ontology structure. This mainly involves dimensionality transformation of the speech encoding feature, resulting in a dimensionality-transformed speech encoding feature that is identical to the dimensionality of the text encoding feature obtained after the prompt word passes through the text encoder. For example, if the dimension of the speech coding features is 96*N, while the text coding features and the pre-trained large language model's ontology structure require 112*N, then the dimension 96 needs to be converted to 112, where N represents the data length. Therefore, the dimension-converted speech coding features and text coding features can be directly concatenated, and the concatenated features are simultaneously input into the LoRA module of the improved large language model's ontology structure and the pre-trained large language model's ontology structure. The output features of the pre-trained large language model's ontology structure and the output features of the LoRA module are added together to obtain the output features of the improved large language model's ontology structure, i.e., the output token sequence. The output token sequence is then decoded by the text decoder to obtain the corresponding output text.

[0040] In one embodiment, the improved large language model's ontology structure is built upon the traditional large language model's ontology structure by adding a LoRA module in parallel. That is, a low-rank matrix is ​​added to the side path of the pre-trained large language model's ontology structure. During the training process, the LoRA module is used as an intermediate component. The parameters of the large language model's ontology structure are frozen during training, and only the parameters of the low-rank matrix of the LoRA module are adjusted, thus effectively reducing the training load. The LoRA module is an existing structure, and its details will not be elaborated here. The text encoder and text decoder use a rule-based traditional word segmenter; neither belongs to the neural network and therefore does not participate in training. In one example, the pre-trained large language model can be a Qwen2 model. The ontology structure of the large language model includes an embedding layer, a transformer structure, and an output layer. The ontology structure itself has context learning capabilities, thus possessing the potential to determine the differences between preceding and following parts of a synthesized speech.

[0041] Before training, cue words need to be constructed to specify the task type. In the embodiments of this application, the cue words are used to determine whether the first speech to be compared and the second speech to be compared are from the same person by merging the speech. As an example, 3 to 5 cue words can be constructed, and one is randomly selected each time it is used. By randomly selecting cue words, the form of the cue words can be enriched and the model's dependence on the expression of the cue words can be reduced.

[0042] In a specific embodiment, the training process of the voiceprint comparison model employs two supervised training sessions, specifically:

[0043] During the first training process, the parameters of the ontology structure of the pre-trained large language model are frozen, and the parameters of the LoRA module, audio encoder and adapter are adjusted to obtain the voiceprint comparison model after one training.

[0044] In the second training process, the voiceprint comparison model after the first training is fine-tuned. The ontology structure of the pre-trained large language model, the parameters of the LoRA module after the first training, the audio encoder and the adapter are adjusted to obtain the trained voiceprint comparison model.

[0045] Specifically, the voiceprint comparison model in the embodiments of this application can use the same dataset for both supervised training sessions. This dataset includes merged speech and its corresponding labels, which are either yes or no, corresponding to whether the two speech segments in the merged speech are from the same person or different people, respectively. The loss function used during training is the cross-entropy loss function, used to compare the differences between the labels and the output text. The training process is as follows:

[0046] Step 1: Freeze the parameters of the ontology structure of the pre-trained large language model and train the LoRA module, adapter, and audio encoder;

[0047] Step 2: Fine-tuning training of the overall model, including training the LoRA module, adapter, audio encoder, and ontology structure of the pre-trained large language model;

[0048] Step two is to fully train the model and further improve its recognition performance.

[0049] S3: Obtain prompt words for determining whether the first and second voice samples to be compared belong to the same person by merging the voice samples. Input the merged voice samples and prompt words into the trained voiceprint comparison model. The merged voice samples are first processed by an audio encoder to obtain voice encoding features; the prompt words are processed by a text encoder to obtain text encoding features. The voice encoding features are then input into an adapter to perform dimensionality transformation to obtain dimensionality-transformed voice encoding features. The text encoding features and the dimensionality-transformed voice encoding features are concatenated and input into the ontology structure of the improved large language model. The output features of the pre-trained large language model ontology structure are added to the output features of the LoRA module to obtain the output token sequence. The output token sequence is then processed by a text decoder to obtain the corresponding output text.

[0050] In a specific embodiment, the first and second voice samples to be compared in the merged speech are distinguished by duration; multiple prompt words are set, and one of them is randomly selected each time; the output text includes yes or no. If the output text is yes, it is determined that the first and second voice samples to be compared are from the same person. If the output text is no, it is determined that the first and second voice samples to be compared are from different people.

[0051] Specifically, the trained voiceprint comparison model is deployed and applied. For the first speech A and the second speech B to be compared, they are concatenated into a single speech segment. A segment of the prompt words used during training is randomly selected as the input prompt words. At the same time, the combined speech and the prompt words are input into the trained voiceprint comparison model. The model judges whether the output is "yes" or "no". If the output is "yes", the two speech segments are considered to be from the same person. If the output is "no", they are judged to be from different people.

[0052] It should be noted that the first and second voices in the embodiments of this application are distinguished by duration. In one example, the first and second voices can each be 4 seconds.

[0053] Further reference Figure 3As an implementation of the methods shown in the above figures, this application provides an embodiment of a voiceprint comparison device based on a large language model. This device embodiment is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0054] This application provides a speaker recognition device based on a large language model, comprising:

[0055] The voice acquisition module 1 is configured to acquire the first voice to be compared and the second voice to be compared respectively, and then concatenate them into a merged voice.

[0056] Model building module 2 is configured to build and train a voiceprint matching model based on a large language model to obtain a trained voiceprint matching model. The voiceprint matching model includes an audio encoder, an adapter, a text encoder, an improved ontology structure of the large language model, and a text decoder. The improved ontology structure of the large language model includes a pre-trained ontology structure of the large language model set in parallel and a LoRA module.

[0057] The comparison module 3 is configured to acquire prompt words for determining whether the first and second voice samples to be compared belong to the same person by merging the voice samples. The merged voice samples and prompt words are input into the trained voiceprint comparison model. The merged voice samples are first processed by an audio encoder to obtain voice encoding features; the prompt words are processed by a text encoder to obtain text encoding features; the voice encoding features are input into an adapter to perform dimensionality transformation to obtain dimensionality-transformed voice encoding features; the text encoding features and the dimensionality-transformed voice encoding features are concatenated and input into the ontology structure of the improved large language model; the output features of the pre-trained large language model ontology structure are added to the output features of the LoRA module to obtain the output token sequence; the output token sequence is processed by a text decoder to obtain the corresponding output text.

[0058] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. For example... Figure 4 As shown, the electronic device in this embodiment includes a processor 401 and a memory 402; wherein the memory 402 is used to store computer execution instructions; and the processor 401 is used to execute the computer execution instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.

[0059] Alternatively, the memory 402 can be either standalone or integrated with the processor 401.

[0060] When the memory 402 is set up independently, the electronic device also includes a bus 403 for connecting the memory 402 and the processor 401.

[0061] This invention also provides a computer storage medium storing computer execution instructions, which, when executed by processor 401, implement the above method.

[0062] This invention also provides a computer program product, including a computer program that, when executed by a processor 401, implements the above-described method.

[0063] In the embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0064] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.

[0065] Furthermore, the functional modules in the various embodiments of this invention can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit formed by the above modules can be implemented in hardware or in the form of hardware plus software functional units.

[0066] The integrated modules implemented as software functional modules described above can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor 401 to execute some steps of the methods of the various embodiments of this application.

[0067] It should be understood that the processor 401 described above can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor, or the processor 401 can be any conventional processor 401. The steps of the method disclosed in this invention can be directly manifested as the hardware processor 401 executing the steps, or as a combination of hardware and software modules within the processor 401 executing the steps.

[0068] The memory 402 may include high-speed RAM memory, and may also include non-volatile memory NVM, such as at least one disk storage device, and may also be a USB flash drive, portable hard drive, read-only memory, disk or optical disc, etc.

[0069] Bus 403 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 403 can be divided into address bus, data bus, control bus, etc. For ease of illustration, the bus 403 in the accompanying drawings of this application is not limited to only one bus 403 or one type of bus 403.

[0070] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.

[0071] An exemplary storage medium is coupled to processor 401, enabling processor 401 to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of processor 401. Processor 401 and storage medium can reside in application-specific integrated circuits (ASICs). Alternatively, processor 401 and storage medium can exist as discrete components in an electronic device or host device.

[0072] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voiceprint matching method based on a large language model, characterized in that, Includes the following steps: The first and second voice samples collected separately are acquired and concatenated into a merged voice sample. A voiceprint matching model based on a large language model is constructed and trained to obtain a trained voiceprint matching model. The voiceprint matching model includes an audio encoder, an adapter, a text encoder, an improved ontology structure of the large language model, and a text decoder. The improved ontology structure of the large language model includes a pre-trained ontology structure of the large language model and a LoRA module configured in parallel. The training process of the voiceprint matching model employs two supervised training iterations, specifically: During the first training process, the parameters of the ontology structure of the pre-trained large language model are frozen, and the parameters of the LoRA module, audio encoder and adapter are adjusted to obtain the voiceprint comparison model after one training. During the second training process, the voiceprint comparison model after the first training is fine-tuned. The ontology structure of the pre-trained large language model, the parameters of the LoRA module after the first training, the audio encoder and the adapter are adjusted to obtain the trained voiceprint comparison model. A prompt word is obtained to determine whether the first and second voice samples to be compared belong to the same person. The merged voice and the prompt word are input into the trained voiceprint comparison model. The merged voice is first processed by the audio encoder to obtain voice encoding features. The prompt word is processed by the text encoder to obtain text encoding features. The voice encoding features are input into the adapter to perform dimensionality transformation to obtain dimensionality-transformed voice encoding features. The text encoding features and the dimensionality-transformed voice encoding features are concatenated and input into the ontology structure of the improved large language model. The output features of the pre-trained large language model ontology structure and the output features of the LoRA module are added to obtain an output token sequence. The output token sequence is processed by the text decoder to obtain the corresponding output text.

2. The voiceprint comparison method based on a large language model according to claim 1, characterized in that, The audio encoder includes a pre-trained speech encoder based on a transformer structure or a pre-trained wav2vec model. The pre-trained speech encoder based on a transformer structure includes a feature extraction module and an encoding module of a Whisper model.

3. The voiceprint comparison method based on a large language model according to claim 1, characterized in that, The adapter includes a first convolutional layer, a first ReLU activation function layer, a second convolutional layer, and a second ReLU activation function layer connected in sequence.

4. The voiceprint comparison method based on a large language model according to claim 1, characterized in that, The first and second voice samples to be compared in the merged speech are distinguished by duration; multiple prompt words are set, and one of them is randomly selected each time; The output text includes yes or no. If the output text is yes, it is determined that the first voice and the second voice to be compared are the voices of the same person. If the output text is no, it is determined that the first voice and the second voice to be compared are the voices of different people.

5. The voiceprint comparison method based on a large language model according to claim 1, characterized in that, The pre-trained large language model includes the Thousand Questions Large Model.

6. A voiceprint comparison device based on a large language model, characterized in that, include: The voice acquisition module is configured to acquire the first voice and the second voice to be compared, which are collected separately, and then concatenate them into a merged voice. The model building module is configured to build and train a speaker recognition model based on a large language model, resulting in a trained speaker recognition model. This model includes an audio encoder, an adapter, a text encoder, an improved ontology structure of the large language model, and a text decoder. The improved ontology structure of the large language model includes a pre-trained ontology structure of the large language model and a LoRA module, all configured in parallel. The training process of the speaker recognition model employs two supervised training iterations, specifically: During the first training process, the parameters of the ontology structure of the pre-trained large language model are frozen, and the parameters of the LoRA module, audio encoder and adapter are adjusted to obtain the voiceprint comparison model after one training. During the second training process, the voiceprint comparison model after the first training is fine-tuned. The ontology structure of the pre-trained large language model, the parameters of the LoRA module after the first training, the audio encoder and the adapter are adjusted to obtain the trained voiceprint comparison model. The comparison module is configured to acquire prompt words for determining whether the first and second voice samples to be compared belong to the same person based on the merged speech. The merged speech and prompt words are input into the trained voiceprint comparison model. The merged speech first passes through the audio encoder to obtain speech encoding features; the prompt words pass through the text encoder to obtain text encoding features; the speech encoding features are input into the adapter, where they undergo dimensionality transformation to obtain dimensionally transformed speech encoding features; the text encoding features and the dimensionally transformed speech encoding features are concatenated and input into the ontology structure of the improved large language model; the output features of the pre-trained large language model's ontology structure are added to the output features of the LoRA module to obtain an output token sequence; the output token sequence passes through the text decoder to obtain the corresponding output text.

7. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Counterfeit voice detection method, system and device fused with large language model, and medium

    CN117577119A

  • Speech recognition method, device and equipment based on end-to-end cross-language large model

    CN119252228A