Performance Detection Method, Device, Equipment and Storage Medium for Voiceprint Recognition System
By constructing and sending various voice spoofing attacks to voice recognition systems, the method evaluates their anti-counterfeiting capabilities, enhancing security by identifying and rating their performance against such attacks.
Patent Information
- Application Number
- CN202111222370.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-10-20
AI Technical Summary
There is a lack of effective methods in the prior art to detect the anti-counterfeiting performance of voiceprint recognition systems, resulting in malicious voiceprint attacks that seriously affect the system security.
By obtaining character audio from the audio database and performing splicing processing, the first attack voiceprint and the second attack voiceprint are constructed, the target user's voice is simulated, and the voiceprint recognition system is sent to the voiceprint recognition system and its recognition response is analyzed to obtain performance detection results.
It can evaluate the anti-counterfeiting performance level of the voiceprint recognition system, identify the ability to forge voiceprints, and improve the security of the system.
Smart Images

Figure CN114023331B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of biometrics and information security, and particularly to a method, device, equipment, and storage medium for detecting the performance of a voiceprint recognition system. Background Art
[0002] Currently, voiceprint recognition systems have been widely applied in multiple business scenarios such as login and payment in Internet finance. A voiceprint recognition system uses an authentication technology based on voiceprint recognition to verify the identity of a user and ensure transaction security. However, at the same time, malicious voiceprint attacks against voiceprint recognition systems are gradually increasing. Attackers imitate, collect, and generate the voiceprints of the attacked to impersonate the identity of the attacked, seriously affecting the security of the voiceprint recognition system. Therefore, it is necessary to detect the anti-counterfeiting performance of the voiceprint recognition system to provide a reference for the security of the voiceprint recognition system.
[0003] However, currently, there is no detection method for detecting the anti-counterfeiting performance of a voiceprint recognition system. Therefore, detecting the anti-counterfeiting performance of a voiceprint recognition system has become an urgent problem to be solved. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, device, equipment, and storage medium for detecting the performance of a voiceprint recognition system that can detect the anti-counterfeiting performance of the voiceprint recognition system.
[0005] In a first aspect, a method for detecting the performance of a voiceprint recognition system is provided. The method includes:
[0006] Obtain multiple character audios from an audio database, and perform splicing processing on the multiple character audios to obtain a first attack voiceprint, where the character audio is an audio segment corresponding to a single character; obtain the voiceprint feature of a target user, and generate a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature; send the first attack voiceprint and the second attack voiceprint to the voiceprint recognition system to obtain an identification response of the voiceprint recognition system; and obtain a performance detection result of the voiceprint recognition system according to the identification response.
[0007] In one embodiment, the first attack voiceprint includes a replay voiceprint. The step of obtaining multiple character audios from the audio database and performing splicing processing on the multiple character audios to obtain the first attack voiceprint includes: randomly obtaining the multiple character audios from the audio database, and performing splicing processing on the randomly obtained multiple character audios to obtain the replay voiceprint.
[0008] In one embodiment, the first attack voiceprint includes a constructed voiceprint. Multiple character audios are obtained from an audio database, and the multiple character audios are spliced to obtain the first attack voiceprint, including: obtaining attack text content, where the attack text content includes multiple text characters; obtaining the multiple character audios respectively corresponding to the text characters from the audio database; splicing the obtained multiple character audios according to the arrangement order of the multiple text characters in the attack text content to obtain the constructed voiceprint.
[0009] In one embodiment, the method further includes: collecting an original voiceprint; performing cleaning processing on the original voiceprint to remove noise in the original voiceprint to obtain a candidate voiceprint; performing segmentation processing on the candidate voiceprint to obtain multiple character audios; and constructing the audio database based on the multiple character audios.
[0010] In one embodiment, obtaining the voiceprint feature of a target user and generating a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature includes: inputting the voiceprint of the target user into a feature extraction neural network to obtain a voiceprint feature vector of the target user; fusing the voiceprint feature vector with voiceprint text to obtain a Mel spectrogram; and performing transformation processing on the Mel spectrogram to obtain the second attack voiceprint.
[0011] In one embodiment, inputting the voiceprint of the target user into the feature extraction neural network to obtain the voiceprint feature vector of the target user includes: performing segmentation processing on the voiceprint of the target user to obtain multiple voiceprint segments, inputting the multiple voiceprint segments into the feature extraction neural network respectively to obtain voiceprint feature vectors corresponding to the voiceprint segments, and taking the average of the voiceprint feature vectors corresponding to the voiceprint segments to obtain the voiceprint feature vector of the target user.
[0012] In one embodiment, before inputting the voiceprint of the target user into the feature extraction neural network, the method further includes: obtaining a training sample set, where the training sample set includes sample voiceprints and voiceprint labels corresponding to the sample voiceprints, and the voiceprint labels are used to indicate whether the sample voiceprints are normal voiceprints or malicious voiceprints; training a classification neural network based on the training sample set, where the classification neural network includes a feature extraction layer; and using the feature extraction layer included in the classification neural network as the feature extraction neural network.
[0013] In one embodiment, obtaining the performance detection result of the voiceprint recognition system according to the recognition response includes: if the recognition response corresponding to the replayed voiceprint is successful, determining that the performance of the voiceprint recognition system is at the first level; if the recognition response corresponding to the replayed voiceprint fails, and the recognition response corresponding to the constructed voiceprint is successful, determining that the performance of the voiceprint recognition system is at the second level; if the recognition responses corresponding to the replayed voiceprint and the constructed voiceprint are both successful, and the recognition response corresponding to the second attack voiceprint fails, determining that the performance of the voiceprint recognition system is at the third level; if the recognition responses corresponding to the replayed voiceprint, the constructed voiceprint, and the second attack voiceprint all fail, determining that the performance of the voiceprint recognition system is at the fourth level; the performance of the voiceprint recognition system represented by the first level, the second level, the third level, and the fourth level increases in sequence.
[0014] In a second aspect, a performance detection device for a voiceprint recognition system is provided. The device includes:
[0015] A first acquisition module, configured to acquire a plurality of character audios from an audio database, and perform splicing processing on the plurality of character audios to obtain a first attack voiceprint, where the character audio is an audio segment corresponding to a single character; a second acquisition module, configured to acquire the voiceprint feature of a target user, and generate a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature; a sending module, configured to send the first attack voiceprint and the second attack voiceprint to the voiceprint recognition system to obtain the recognition response of the voiceprint recognition system; a third acquisition module, configured to obtain the performance detection result of the voiceprint recognition system according to the recognition response.
[0016] In one embodiment, the first attack voiceprint includes a replayed voiceprint. The first acquisition module is specifically configured to: randomly acquire the plurality of character audios from the audio database, and perform splicing processing on the randomly acquired plurality of character audios to obtain the replayed voiceprint.
[0017] In one embodiment, the first attack voiceprint includes a constructed voiceprint. The first acquisition module is specifically configured to: acquire attack text content, where the attack text content includes a plurality of text characters; acquire the plurality of character audios corresponding to the text characters from the audio database respectively; perform splicing processing on the acquired plurality of character audios in the arrangement order of the plurality of text characters in the attack text content to obtain the constructed voiceprint.
[0018] In one embodiment, the device further includes:
[0019] The acquisition module is used to acquire the original voiceprint; the cleaning module is used to clean the original voiceprint to remove the noise in the original voiceprint and obtain the candidate voiceprint; the segmentation module is used to segment the candidate voiceprint to obtain multiple character audios; the construction module is used to construct the audio database based on the multiple character audios.
[0020] In one embodiment, the second acquisition module is specifically configured to: input the voiceprint of the target user into the feature extraction neural network to obtain the voiceprint feature vector of the target user; fuse the voiceprint feature vector with the voiceprint text to obtain the Mel spectrogram; perform a conversion process on the Mel spectrogram to obtain the second attack voiceprint.
[0021] In one embodiment, the second acquisition module is specifically configured to: segment the voiceprint of the target user to obtain multiple voiceprint segments, input the multiple voiceprint segments into the feature extraction neural network respectively to obtain the voiceprint feature vectors corresponding to the voiceprint segments, and take the average value of the voiceprint feature vectors corresponding to the voiceprint segments to obtain the voiceprint feature vector of the target user.
[0022] In one embodiment, the device further includes:
[0023] The fourth acquisition module is used to acquire a training sample set, where the training sample set includes sample voiceprints and the voiceprint labels corresponding to the sample voiceprints, and the voiceprint labels are used to indicate whether the sample voiceprints are normal voiceprints or malicious voiceprints; the training module is used to train a classification neural network based on the training sample set, and the classification neural network includes a feature extraction layer; use the feature extraction layer included in the classification neural network as the feature extraction neural network.
[0024] In one embodiment, the third acquisition module is specifically configured to: if the recognition response corresponding to the replayed voiceprint is successful, determine that the performance of the voiceprint recognition system is at the first level; if the recognition response corresponding to the replayed voiceprint fails, and the recognition response corresponding to the constructed voiceprint is successful, determine that the performance of the voiceprint recognition system is at the second level; if the recognition responses corresponding to the replayed voiceprint and the constructed voiceprint are both successful, and the recognition response corresponding to the second attack voiceprint fails, determine that the performance of the voiceprint recognition system is at the third level; if the recognition responses corresponding to the replayed voiceprint, the constructed voiceprint, and the second attack voiceprint all fail, determine that the performance of the voiceprint recognition system is at the fourth level; the first level, the second level, the third level, and the fourth level represent the performance of the voiceprint recognition system increasing in sequence.
[0025] In a third aspect, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above first aspects are implemented.
[0026] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above first aspects are implemented.
[0027] For the above method, device, equipment, and storage medium for detecting the performance of the voiceprint recognition system, by obtaining multiple character audios from an audio database and performing splicing processing on the multiple character audios to obtain a first attack voiceprint, that is, constructing a forged voiceprint; by obtaining the voiceprint feature of the target user and generating a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature, that is, constructing another forged voiceprint; since the character audio for constructing the first attack voiceprint is an audio segment corresponding to a single character, and the construction of the second attack voiceprint is based on the voiceprint feature of the target user, therefore, the forging complexities of the constructed first attack voiceprint and the second attack voiceprint are different; by sending the first attack voiceprint and the second attack voiceprint with different forging complexities to the voiceprint recognition system to obtain the recognition response of the voiceprint recognition system, the performance detection result of the voiceprint recognition system can be obtained according to the recognition response, and the anti-counterfeiting performance and anti-counterfeiting level of the voiceprint recognition system can be detected. Description of the Drawings
[0028] Figure 1 It is an application environment diagram of a method for detecting the performance of a voiceprint recognition system provided by an embodiment of the present application;
[0029] Figure 2 It is a flowchart of a method for detecting the performance of a voiceprint recognition system provided by an embodiment of the present application;
[0030] Figure 3 It is a flowchart of constructing an audio database provided by an embodiment of the present application;
[0031] Figure 4 It is a flowchart of constructing a replay voiceprint provided by an embodiment of the present application;
[0032] Figure 5 It is a flowchart of constructing a constructed voiceprint provided by an embodiment of the present application;
[0033] Figure 6 It is a schematic diagram of obtaining a feature extraction neural network provided by an embodiment of the present application;
[0034] Figure 7 It is a schematic diagram of constructing a second attack voiceprint provided by an embodiment of the present application;
[0035] Figure 8 Schematic diagram of obtaining a voiceprint feature vector provided by an embodiment of the present application;
[0036] Figure 9 Schematic diagram of a method for detecting the anti-counterfeiting performance of a voiceprint recognition system provided by an embodiment of the present application;
[0037] Figure 10 Block diagram of a performance detection device for a voiceprint recognition system provided by an embodiment of the present application;
[0038] Figure 11 Block diagram of a performance detection device for a second voiceprint recognition system provided by an embodiment of the present application;
[0039] Figure 12 Block diagram of a performance detection device for a third voiceprint recognition system provided by an embodiment of the present application;
[0040] Figure 13 Block diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0041] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0042] With the rapid development of the Internet and intelligent devices, voiceprint recognition systems based on machine learning and deep learning have been widely applied in multiple business scenarios such as the login and payment of Internet financial services. The voiceprint recognition system uses an identity verification technology based on voiceprint recognition to verify the identity of users, greatly improving the usability and security of the services.
[0043] At the same time, malicious voiceprint attacks against these application scenarios have gradually increased. Attackers imitate, collect, arrange, and generate the voices and voiceprint characteristics of the attacked person to bypass the voiceprint recognition system for key business transactions, seriously affecting the security of the voiceprint recognition system and causing a greater impact on the security and ecological health of voiceprint recognition. Therefore, it is necessary to detect the anti-counterfeiting performance of the voiceprint recognition system to provide a reference for the security of the voiceprint recognition system.
[0044] However, each website and platform has gradually incorporated the security detection of the voiceprint recognition system into the security management work, but the protection methods adopted by each party are different, and the evaluation and protection levels are uneven. At present, there is no detection method for detecting the anti-counterfeiting performance of the voiceprint recognition system. Therefore, detecting the anti-counterfeiting performance of the voiceprint recognition system has become an urgent problem to be solved.
[0045] The performance detection method for the voiceprint recognition system provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown in the figure. The attack voiceprint construction system 101 is communicatively connected to the voiceprint recognition system 102. The attack voiceprint construction system sends the constructed first attack voiceprint and second attack voiceprint to the voiceprint recognition system. The voiceprint recognition system receives the first attack voiceprint and the second attack voiceprint and performs recognition on them to obtain corresponding recognition responses. In subsequent steps, according to the recognition responses of the voiceprint recognition system to the first attack voiceprint and the second attack voiceprint, the performance detection result of the voiceprint recognition system is obtained. Among them, the attack voiceprint construction system 101 can be, but is not limited to, a server, a personal computer, a laptop computer, etc., and the voiceprint recognition system 102 can be, but is not limited to, a server or a server cluster, various computer devices, a laptop computer, a smart phone, a tablet computer, etc.
[0046] In the embodiments of the present application, as Figure 2 shown in the figure, it shows a flowchart of a performance detection method for a voiceprint recognition system provided by the embodiments of the present application. Taking the application of this method to the Figure 1 attack voiceprint construction system 101 in the figure as an example, the method includes the following steps:
[0047] Step 201: Obtain a plurality of character audio from the audio database, and perform splicing processing on the plurality of character audio to obtain a first attack voiceprint, where the character audio is an audio segment corresponding to a single character.
[0048] Among them, various audible continuous audio is composed of each character audio. The content of each character audio is, for example, text or numbers, etc. Each character is output in audio form so as to be heard or recognized by the device. That is, each valid character corresponds to an audio segment, which is the character audio. The audio database is composed of a plurality of character audio; the voiceprint is a continuous audio that can be played and is composed of a plurality of character audio. The voiceprint recognition system is used to recognize the voiceprints of each user. After the recognition is passed, the user can perform other operations in the next step; in the scenario where the voiceprint recognition system performs voiceprint recognition on a certain user, the user is used as the target user, and the voiceprint naturally generated by the user is the voiceprint of the target user. To detect whether the voiceprint recognition system can accurately recognize the voiceprint of the target user, that is, to detect the anti-counterfeiting ability of the voiceprint recognition system, a forged voiceprint can be constructed and input into the voiceprint recognition system for recognition, and based on the recognition result of the forged voiceprint by the voiceprint recognition system, the anti-counterfeiting ability of the voiceprint recognition system can be judged; among them, the first attack voiceprint is the forged voiceprint, and it can be formed by obtaining a plurality of independent character audio from the audio database and performing splicing processing on the plurality of independent character audio to form a continuous audio that can be played as the first attack voiceprint. The first attack voiceprint is not the voiceprint naturally generated by the target user, but the voiceprint constructed by splicing processing.
[0049] Step 202: Obtain the voiceprint feature of the target user, and generate a second attack voiceprint for simulating the voice of the target user based on this voiceprint feature.
[0050] Among them, the voiceprint of the target user is continuous audio, which is composed of multiple character audios. The voiceprint feature of the target user can be the feature information extracted from the voiceprint of the target user. For example, this feature information can be a feature vector. Based on the voiceprint feature of the target user, simulate the voiceprint of the target user to obtain a simulated voiceprint that simulates the voice of the target user, and use this simulated voiceprint as the second attack voiceprint. The second attack voiceprint is composed of each character audio similar to each character audio included in the voiceprint of the target user.
[0051] Step 203: Send the first attack voiceprint and the second attack voiceprint to the voiceprint recognition system to obtain the recognition response of the voiceprint recognition system.
[0052] Among them, obviously, the above-obtained first attack voiceprint is spliced, and the second attack voiceprint is simulated. The forgery degree of the second attack voiceprint is higher than that of the first attack voiceprint. By sending the first attack voiceprint and the second attack voiceprint to the voiceprint recognition system respectively, the voiceprint recognition system recognizes the first attack voiceprint and the second attack voiceprint, and obtains the recognition results corresponding to the first attack voiceprint and the second attack voiceprint respectively. This recognition result is the recognition response of the voiceprint recognition system.
[0053] Step 204: Obtain the performance detection result of the voiceprint recognition system according to this recognition response.
[0054] Among them, this recognition response indicates whether the voiceprint recognition system successfully recognizes the first attack voiceprint and the second attack voiceprint. If successful, it means that the voiceprint recognition system does not recognize that each attack voiceprint is a forged voiceprint, indicating that the anti-counterfeiting performance of the voiceprint recognition system needs to be improved. Among them, according to the recognition responses of the voiceprint recognition system to the first attack voiceprint and the second attack voiceprint with different forgery degrees, the anti-counterfeiting performance level of the voiceprint recognition system can be judged.
[0055] The performance detection method of the above voiceprint recognition system obtains multiple character audios from an audio database, splices and processes the multiple character audios to obtain a first attack voiceprint, that is, constructs a forged voiceprint; obtains the voiceprint feature of the target user, and generates a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature, that is, constructs another forged voiceprint; since the character audio for constructing the first attack voiceprint is an audio segment corresponding to a single character, and the second attack voiceprint is constructed based on the voiceprint feature of the target user, therefore, the forging complexity of the constructed first attack voiceprint and the second attack voiceprint is different; by sending the first attack voiceprint and the second attack voiceprint with different forging complexities to the voiceprint recognition system, obtaining the recognition response of the voiceprint recognition system, so that the performance detection result of the voiceprint recognition system can be obtained according to the recognition response, and the anti-counterfeiting performance and anti-counterfeiting level of the voiceprint recognition system can be detected.
[0056] In the embodiment of the present application, as Figure 3 shown, it shows a flowchart of constructing an audio database provided by the embodiment of the present application. The method further includes the following steps:
[0057] Step 301, collect the original voiceprint.
[0058] Among them, to construct an audio database, it is necessary to first collect voiceprint data from different scenarios. The various voiceprint data obtained by the collection are the original voiceprints; the original voiceprints can be obtained by methods such as social engineering, on-site conversation recording, telephone recording, offline and online meeting recording, on-site speech recording, self-media platform download, phishing page induction, etc.; among them, the original voiceprint collection device can be a mobile phone microphone, a professional recording device, a sound card set and other audio collection devices.
[0059] Step 302, perform a cleaning process on the original voiceprint to remove the noise in the original voiceprint and obtain candidate voiceprints.
[0060] Among them, most of the acquisition scenarios of the original voiceprints are relatively noisy, and there are many invalid noises in each of the obtained original voiceprints, such as background sounds, etc.; therefore, in order to obtain clean original voiceprints, it is necessary to perform a cleaning process on each of the obtained original voiceprints. The cleaning process is to remove the invalid noises in each original voiceprint. The cleaning process can be to remove the invalid noises in each original voiceprint by using methods such as inverse filtering denoising, so as to obtain the pure and effective audio corresponding to each original voiceprint, and use each pure and effective audio as a candidate voiceprint.
[0061] Step 303, perform a segmentation process on the candidate voiceprint to obtain multiple character audios.
[0062] After obtaining each candidate voiceprint, use an automated tool to perform segmentation processing on each candidate voiceprint, segment each candidate voiceprint into single-character audio to obtain multiple character audios. Among them, the segmentation processing of each candidate voiceprint can be achieved by using an automated tool. The segmentation process of this automated tool can be as follows: For example, segment a candidate voiceprint containing numbers from 0 to 9. Denote the waveform function corresponding to this candidate voiceprint as f(t). Obtain the invalid noise filtered out during the cleaning process of the original voiceprint corresponding to this candidate voiceprint, and denote the waveform function corresponding to this invalid noise as S(t). Denote the peak value of the function S(t) as S1. Select a time window T containing the target number, and the target number can be from 0 to 9. When f(t) is greater than the peak value S1 for the first time and remains for a certain period, it is considered that the position t 0S is the starting point of the audio of the number 0. When f(t) is less than the peak value S1 for the first time and remains for a certain period, it is considered that the position t 0e is the end point of the number 0, and segment the audio from t 0S to t 0e to obtain the audio segment of the number 0. By analogy, the audio segments of the numbers from 0 to 9 can be segmented.
[0063] Step 304, construct the audio database based on the multiple character audios.
[0064] Segment each candidate voiceprint to obtain multiple character audios, store all the character audios in the attack voiceprint construction system as the audio database, and some character audios in the audio database can be directly called when constructing the voiceprint.
[0065] By collecting the original voiceprints from various different scenarios, the randomness of the original voiceprints is ensured, and the data volume of the original voiceprints is ensured to be sufficient. By performing cleaning and segmentation processing on the original voiceprints, not only can multiple character audios be obtained, but also each obtained character audio is a clean audio, so that it will not be interfered by invalid noise during subsequent use.
[0066] In the embodiment of the present application, as Figure 4 shown, it shows a flowchart of constructing a reproduced voiceprint provided by the embodiment of the present application. The first attack voiceprint includes the reproduced voiceprint. Obtain multiple character audios from the audio database and perform splicing processing on the multiple character audios to obtain the first attack voiceprint, including:
[0067] Step 401, randomly obtain the multiple character audios from the audio database.
[0068] Step 402, perform splicing processing on the randomly obtained multiple character audios to obtain the reproduced voiceprint.
[0069] Among them, the first attack voiceprint may include a reproduced voiceprint. In the process of constructing the reproduced voiceprint, first randomly select a preset number of character audios from the audio database, directly splice the selected character audios to form a continuous audio, and use this continuous audio as the reproduced voiceprint; where the preset number can be set to different values according to the actual situation.
[0070] Since this reproduced voiceprint is composed of a plurality of randomly obtained character audios, in the scenario where the target user performs voiceprint recognition in the voiceprint recognition system, the character audios included in this reproduced voiceprint are meaningless. Therefore, the similarity between this reproduced voiceprint and the target user's voiceprint is relatively low. Send this reproduced voiceprint to the voiceprint recognition system. If the voiceprint recognition system fails to recognize this reproduced voiceprint, it means that the voiceprint recognition system recognizes that this reproduced voiceprint is different from the target voiceprint and is a forged voiceprint. Thus, it can be shown that the anti-counterfeiting performance of the voiceprint recognition system meets a lower level; therefore, this reproduced voiceprint can be used to detect whether the voiceprint recognition system has anti-counterfeiting ability at a lower level.
[0071] In the embodiments of the present application, as Figure 5 shown, it shows a flowchart of constructing a constructed voiceprint provided by the embodiments of the present application. The first attack voiceprint includes a constructed voiceprint. Obtain a plurality of character audios from the audio database and perform splicing processing on the plurality of character audios to obtain the first attack voiceprint, including:
[0072] Step 501, obtain attack text content, where the attack text content includes a plurality of text characters.
[0073] Among them, in the scenario where the target user performs voiceprint recognition in the voiceprint recognition system, the target user outputs a voiceprint containing the text content given by the voiceprint recognition system. The text content includes a plurality of text characters, and the characters can be text, numbers, etc.; use this text content as the attack text content. Correspondingly, the attack text content includes a plurality of text characters.
[0074] Step 502, obtain the plurality of character audios respectively corresponding to each text character from the audio database.
[0075] Among them, according to each text character in the obtained attack text content, select the character audios respectively corresponding to each text character from the audio database to obtain a plurality of character audios.
[0076] Step 503, perform splicing processing on the obtained plurality of character audios according to the arrangement order of the plurality of text characters in the attack text content to obtain the constructed voiceprint.
[0077] Among them, the order of each character audio included in the voiceprint of the target user is determined according to the text content given by the voiceprint recognition system, that is, the order of each text character included in the attack text content is the same as the order of each character audio included in the voiceprint of the target user. Through the order of each character audio included in the voiceprint of the target user, the arrangement order of each text character included in the attack text content is determined, and the multiple character audios corresponding to each text character are spliced according to this arrangement order to form a continuous audio, and this continuous audio is used as the constructed voiceprint.
[0078] Since the order of each character audio included in this constructed voiceprint is determined according to the text content given by the voiceprint recognition system, therefore, in the scenario where the target user performs voiceprint recognition in the voiceprint recognition system, compared with the replayed voiceprint, this constructed voiceprint is meaningful, so this constructed voiceprint has a certain similarity with the voiceprint of the target user; in the scenario where the target user performs voiceprint recognition in the voiceprint recognition system, this constructed voiceprint is sent to the voiceprint recognition system. If the voiceprint recognition system fails to recognize this constructed voiceprint, it means that the voiceprint recognition system recognizes that this constructed voiceprint is different from the target voiceprint and is a forged voiceprint, thus it can be shown that the anti-counterfeiting performance of the voiceprint recognition system meets the general level; therefore, this constructed voiceprint can be used to detect whether the voiceprint recognition system has the anti-counterfeiting ability of the general level.
[0079] In the embodiment of the present application, as described above, in the process of obtaining the second attack voiceprint, first, the voiceprint feature of the target user needs to be obtained. Optionally, a feature extraction neural network can be used to obtain the voiceprint feature of the target user; please refer to Figure 6 , which shows a flowchart of obtaining a feature extraction neural network provided by the embodiment of the present application. Before inputting the voiceprint of the target user into the feature extraction neural network, the method further includes:
[0080] Step 601, obtain a training sample set, which includes sample voiceprints and voiceprint labels corresponding to the sample voiceprints, and the voiceprint labels are used to indicate whether the sample voiceprints are normal voiceprints or malicious voiceprints.
[0081] Among them, each obtained original voiceprint is used as a sample voiceprint, and a simple classifier is used to classify each sample voiceprint into a normal voiceprint or a malicious voiceprint. The normal voiceprint refers to the voiceprint directly output by the user in the scenario where each original voiceprint is obtained. The malicious voiceprint refers to the voiceprint that has been forged and used in these scenarios in the scenario where each original voiceprint is obtained; after each sample voiceprint is classified into a normal voiceprint or a malicious voiceprint, each normal voiceprint is bound to a normal voiceprint label, and the normal voiceprint label is used to indicate that the bound sample voiceprint is a normal voiceprint; each malicious voiceprint is bound to a malicious voiceprint label, and the malicious voiceprint label is used to indicate that the bound sample voiceprint is a malicious voiceprint; each normal voiceprint and the corresponding normal voiceprint label and each malicious voiceprint and the corresponding malicious voiceprint label are used as a training sample set.
[0082] Step 602: Train a classification neural network based on the training sample set. The classification neural network includes a feature extraction layer.
[0083] Among them, the obtained sample training set is used to train the classification neural network. The trained classification neural network can identify an input target voiceprint to determine whether the target voiceprint is a normal voiceprint or a malicious voiceprint; among them, in the classification neural network, there is a feature extraction layer for extracting a feature vector of the input target voiceprint during the recognition process.
[0084] Step 603: Use the feature extraction layer included in the classification neural network as the feature extraction neural network.
[0085] Use the feature extraction layer included in the above classification neural network as the feature extraction neural network, and the feature extraction neural network is used to output the feature vector corresponding to the input target voiceprint.
[0086] In the embodiments of the present application, as Figure 7 shown, it shows a flowchart of constructing a second attack voiceprint provided by the embodiments of the present application. Obtain the voiceprint feature of the target user, and generate a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature, including:
[0087] Step 701: Input the voiceprint of the target user into the feature extraction neural network to obtain the voiceprint feature vector of the target user.
[0088] Step 702: Perform a fusion process on the voiceprint feature vector and voiceprint text to obtain a Mel spectrum.
[0089] Among them, the voiceprint of the target user for recognition in the voiceprint recognition system is obtained, and the voiceprint of the target user is input into the feature extraction neural network, so that the voiceprint feature vector corresponding to the voiceprint of the target user can be obtained; the characters corresponding to each character audio included in the voiceprint of the target user are obtained, which are the voiceprint characters of the target user. The voiceprint feature vector corresponding to the voiceprint of the target user is fused with the voiceprint characters of the target user, and then the corresponding Mel spectrogram can be obtained.
[0090] Step 703: Perform a conversion process on the Mel spectrogram to obtain the second attack voiceprint.
[0091] Among them, the Mel spectrogram obtained after the above feature fusion is the voiceprint information in the frequency domain. Therefore, it is necessary to perform a conversion process on the Mel spectrogram to convert it into the voiceprint information in the time domain. The voiceprint information in the time domain includes multiple audio characters and can be played normally. The voiceprint information in the time domain is used as the second attack voiceprint.
[0092] Since the second attack voiceprint is obtained by fusing the voiceprint feature vector corresponding to the voiceprint of the target user with the voiceprint characters of the target user, in the scenario where the target user performs voiceprint recognition in the voiceprint recognition system, the second attack voiceprint has a high similarity to the voiceprint of the target user; in the scenario where the target user performs voiceprint recognition in the voiceprint recognition system, the second attack voiceprint is sent to the voiceprint recognition system. If the voiceprint recognition system fails to recognize the second attack voiceprint, it means that the voiceprint recognition system recognizes that the second attack voiceprint is different from the target voiceprint and is a forged voiceprint, which can indicate that the anti-counterfeiting performance of the voiceprint recognition system is high; therefore, the second attack voiceprint can be used to detect whether the voiceprint recognition system has a high-level anti-counterfeiting ability.
[0093] In the embodiments of the present application, as Figure 8 shown, it shows a flowchart of obtaining a voiceprint feature vector provided by the embodiments of the present application. Inputting the voiceprint of the target user into the feature extraction neural network to obtain the voiceprint feature vector of the target user includes:
[0094] Step 801: Perform a segmentation process on the voiceprint of the target user to obtain multiple voiceprint segments.
[0095] Step 802: Input the multiple voiceprint segments into the feature extraction neural network respectively to obtain the voiceprint feature vectors corresponding to the respective voiceprint segments.
[0096] Among them, the voiceprint of the target user can be segmented by seconds to obtain multiple voiceprint segments, and each voiceprint segment contains a section of audio; in the order of segmentation, the respective voiceprint segments are input into the feature extraction neural network respectively, and each voiceprint segment obtains the corresponding voiceprint feature vector, so as to obtain the voiceprint feature vectors corresponding to the respective voiceprint segments.
[0097] Step 803: Take the average value of the voiceprint feature vectors corresponding to each voiceprint segment to obtain the voiceprint feature vector of the target user.
[0098] Among them, by performing an averaging process on the voiceprint feature vectors corresponding to each voiceprint segment, a feature vector can be obtained, and this feature vector is used as the voiceprint feature vector of the target user for subsequent processing.
[0099] In the embodiment of the present application, according to the recognition response, the performance detection result of the voiceprint recognition system is obtained, including: if the recognition response corresponding to the replayed voiceprint is successful, it is determined that the performance of the voiceprint recognition system is at the first level; if the recognition response corresponding to the replayed voiceprint fails and the recognition response corresponding to the constructed voiceprint is successful, it is determined that the performance of the voiceprint recognition system is at the second level; if the recognition responses corresponding to the replayed voiceprint and the constructed voiceprint are both successful and the recognition response corresponding to the second attack voiceprint fails, it is determined that the performance of the voiceprint recognition system is at the third level; if the recognition responses corresponding to the replayed voiceprint, the constructed voiceprint, and the second attack voiceprint all fail, it is determined that the performance of the voiceprint recognition system is at the fourth level; the performance of the voiceprint recognition system represented by the first level, the second level, the third level, and the fourth level increases in sequence.
[0100] Among them, in the scenario where the target user performs voiceprint recognition in the voiceprint recognition system, the attack voiceprint construction system respectively sends the constructed replayed voiceprint, constructed voiceprint, and second attack voiceprint to the voiceprint recognition system in batches; first, the replayed voiceprint is sent to the recognition server in the voiceprint recognition system, and the recognition server recognizes the replayed voiceprint and outputs the corresponding recognition response to the attack voiceprint construction system; when the attack voiceprint construction system receives that the recognition response corresponding to the replayed voiceprint sent by the recognition server is successful, it sends the constructed voiceprint to the recognition server, and the recognition server recognizes the constructed voiceprint and outputs the corresponding recognition response to the attack voiceprint construction system; when the attack voiceprint construction system receives that the recognition response corresponding to the constructed voiceprint sent by the recognition server is successful, it sends the second attack voiceprint to the recognition server, and the recognition server recognizes the second attack voiceprint and outputs the corresponding recognition response to the attack voiceprint construction system; among them, when the recognition response output by the recognition server is successful, it indicates that the voiceprint is not recognized by the recognition server as a forged voiceprint.
[0101] For each recognition response result obtained by the system constructed based on the attack voiceprint, if the recognition response corresponding to the reproduced voiceprint is successful, the performance of the voiceprint recognition system is determined to be the first level; if the recognition response corresponding to the reproduced voiceprint fails, and the recognition response corresponding to the constructed voiceprint is successful, the performance of the voiceprint recognition system is determined to be the second level; if the recognition responses corresponding to both the reproduced voiceprint and the constructed voiceprint are successful, and the recognition response corresponding to the second attack voiceprint fails, the performance of the voiceprint recognition system is determined to be the third level; if the recognition responses corresponding to the reproduced voiceprint, the constructed voiceprint, and the second attack voiceprint all fail, the performance of the voiceprint recognition system is determined to be the fourth level; the performance of the voiceprint recognition system represented by the first level, the second level, the third level, and the fourth level increases in sequence.
[0102] By sending forged voiceprints with different similarities to the target user's voiceprint to the recognition server, the obtained recognition response results also reflect the ability of the recognition server to recognize forged voiceprints with different similarities, so that the recognition ability of the recognition server, that is, the anti-counterfeiting performance of the voiceprint recognition system, can be determined according to each response result.
[0103] In the embodiments of the present application, as Figure 9 shown, it shows a flowchart of a method for detecting the anti-counterfeiting performance of a voiceprint recognition system provided by the embodiments of the present application, including:
[0104] Step 901, collect the original voiceprint.
[0105] In the attack voiceprint construction system, various audio data are obtained through methods such as social engineering, on-site conversation recording, telephone recording, offline and online meeting recording, on-site speech recording, self-media platform downloading, phishing page induction, etc., and each audio data is used as the original voiceprint; among them, the original voiceprint acquisition device can be an audio acquisition device such as a mobile phone microphone, a professional recording device, and a sound card set.
[0106] Step 902, perform cleaning processing on the original voiceprint to remove the noise in the original voiceprint and obtain candidate voiceprints.
[0107] To obtain clean original voiceprints, it is necessary to perform cleaning processing on the obtained original voiceprints. This cleaning processing is to remove the invalid noise in each original voiceprint, such as background noise, etc.; this cleaning processing can be to remove the invalid noise in each original voiceprint by using common filtering methods such as inverse filtering denoising, so as to obtain the pure and effective audio corresponding to each original voiceprint, and use each pure and effective audio as the candidate voiceprint.
[0108] Step 903, perform segmentation processing on the candidate voiceprints to obtain multiple character audios, and construct an audio database based on the multiple character audios.
[0109] After obtaining each candidate voiceprint, use an automated tool to segment each candidate voiceprint, segment each candidate voiceprint into single-character audio to obtain multiple character audios; use the set of character audios composed of the multiple character audios as the audio database.
[0110] Step 904, obtain a training sample set, which includes sample voiceprints and voiceprint labels corresponding to the sample voiceprints, and the voiceprint labels are used to indicate whether the sample voiceprints are normal voiceprints or malicious voiceprints.
[0111] Use the obtained original voiceprints as sample voiceprints, classify each sample voiceprint using a simple classifier, classify each sample voiceprint as a normal voiceprint or a malicious voiceprint. The normal voiceprint refers to the voiceprint directly output by the user in the scenario where each original voiceprint is obtained. The malicious voiceprint refers to the voiceprint that has been forged and used in these scenarios in the scenario where each original voiceprint is obtained; after each sample voiceprint is classified as a normal voiceprint or a malicious voiceprint, bind each normal voiceprint to a normal voiceprint label, and the normal voiceprint label is used to indicate that the bound sample voiceprint is a normal voiceprint; bind each malicious voiceprint to a malicious voiceprint label, and the malicious voiceprint label is used to indicate that the bound sample voiceprint is a malicious voiceprint; use each normal voiceprint and the corresponding normal voiceprint label and malicious voiceprint and the corresponding malicious voiceprint label as the training sample set.
[0112] Step 905, train a classification neural network based on the training sample set. The classification neural network includes a feature extraction layer, and use the feature extraction layer included in the classification neural network as the feature extraction neural network.
[0113] Use the above-obtained sample training set to train the classification neural network. The trained classification neural network can identify an input target voiceprint to determine whether the target voiceprint is a normal voiceprint or a malicious voiceprint; among them, in the classification neural network, there is a feature extraction layer for extracting feature vectors of the input target voiceprint during the recognition process, and use the feature extraction layer included in the classification neural network as the feature extraction neural network.
[0114] Step 906, randomly obtain multiple character audios from the audio database, and splice the randomly obtained multiple character audios to obtain a reproduced voiceprint.
[0115] Randomly select a preset number of multiple character audios from the audio database, directly splice the selected multiple character audios to form a continuous audio, and use the continuous audio as the reproduced voiceprint; where the preset number can be set to different values according to the actual situation.
[0116] Step 907: Obtain the attack text content, which includes multiple text characters. Retrieve multiple character audios corresponding to each text character from the audio database respectively, and splice the obtained multiple character audios according to the arrangement order of the multiple text characters in the attack text content to obtain a constructed voiceprint.
[0117] In the scenario where the target user performs voiceprint recognition in the voiceprint recognition system, the target user outputs a voiceprint containing the text content given by the voiceprint recognition system. This text content includes multiple text characters and their corresponding arrangement order. The character can be a text, a number, etc.; take this text content as the attack text content; according to each text character in the obtained attack text content, select the character audio corresponding to each text character from the audio database respectively to obtain multiple character audios; according to the corresponding arrangement order of the text content given by the voiceprint recognition system, splice the multiple character audios corresponding to each text character selected from the audio character library according to this arrangement order to form a continuous audio, and take this continuous audio as the constructed voiceprint.
[0118] Step 908: Perform segmentation processing on the voiceprint of the target user to obtain multiple voiceprint segments. Input the multiple voiceprint segments into the feature extraction neural network respectively to obtain the voiceprint feature vectors corresponding to each voiceprint segment, and take the average value of the voiceprint feature vectors corresponding to each voiceprint segment to obtain the voiceprint feature vector of the target user.
[0119] Among them, the voiceprint of the target user can be segmented by seconds to obtain multiple voiceprint segments, and each voiceprint segment contains a section of audio; according to the segmentation order, input each voiceprint segment into the feature extraction neural network respectively, and each voiceprint segment obtains the corresponding voiceprint feature vector, so as to obtain the voiceprint feature vectors corresponding to each voiceprint segment; take the average value of the voiceprint feature vectors corresponding to each voiceprint segment to obtain the voiceprint feature vector of the target user.
[0120] Step 909: Perform fusion processing on the voiceprint feature vector and the voiceprint text to obtain a Mel spectrum, and perform transformation processing on the Mel spectrum to obtain the second attack voiceprint.
[0121] Obtain the characters corresponding to each character audio included in the voiceprint of the target user, which is the voiceprint text of the target user. Perform feature fusion on the voiceprint feature vector corresponding to the voiceprint of the target user and the voiceprint text of the target user to obtain the corresponding Mel spectrum; among them, the Mel spectrum obtained after feature fusion is the voiceprint information in the frequency domain. Therefore, perform transformation processing on this Mel spectrum to transform it into the voiceprint information in the time domain, and take this voiceprint information in the time domain as the second attack voiceprint.
[0122] Step 910: Send the replay voiceprint, the constructed voiceprint and the second attack voiceprint to the voiceprint recognition system to obtain the recognition response of the voiceprint recognition system.
[0123] In the scenario where a target user performs voiceprint recognition in a voiceprint recognition system, the reproduced voiceprint, the constructed voiceprint, and the second attack voiceprint constructed in the attack voiceprint construction system are respectively sent to the voiceprint recognition system for recognition.
[0124] First, send the reproduced voiceprint to the recognition server in the voiceprint recognition system. The recognition server recognizes the reproduced voiceprint and outputs the corresponding recognition response to the attack voiceprint construction system. When the attack voiceprint construction system receives that the recognition response corresponding to the reproduced voiceprint sent by the recognition server is successful, it sends the constructed voiceprint to the recognition server. The recognition server recognizes the constructed voiceprint and outputs the corresponding recognition response to the attack voiceprint construction system. When the attack voiceprint construction system receives that the recognition response corresponding to the constructed voiceprint sent by the recognition server is successful, it sends the second attack voiceprint to the recognition server. The recognition server recognizes the second attack voiceprint and outputs the corresponding recognition response to the attack voiceprint construction system. Among them, when the recognition response output by the recognition server is successful, it indicates that the voiceprint is not recognized by the recognition server as a forged voiceprint.
[0125] Step 911, obtain the anti-counterfeiting performance detection result of the voiceprint recognition system according to the recognition response.
[0126] According to the respective response recognition results obtained by the attack voiceprint construction system, if the recognition response corresponding to the reproduced voiceprint is successful, it is determined that the performance of the voiceprint recognition system is at the first level; if the recognition response corresponding to the reproduced voiceprint fails and the recognition response corresponding to the constructed voiceprint is successful, it is determined that the performance of the voiceprint recognition system is at the second level; if the recognition responses corresponding to both the reproduced voiceprint and the constructed voiceprint are successful and the recognition response corresponding to the second attack voiceprint fails, it is determined that the performance of the voiceprint recognition system is at the third level; if the recognition responses corresponding to the reproduced voiceprint, the constructed voiceprint, and the second attack voiceprint all fail, it is determined that the performance of the voiceprint recognition system is at the fourth level. The performance of the voiceprint recognition system represented by the first level, the second level, the third level, and the fourth level increases in sequence. Thus, the anti-counterfeiting performance level of the voiceprint recognition system can be determined according to this level.
[0127] It should be understood that although Figure 2-9 the steps in the flowchart of Figure 2-9At least some of the steps may include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.
[0128] In the embodiments of the present application, as Figure 10 shown, which shows a block diagram of a performance detection device for a voiceprint recognition system provided by an embodiment of the present application. The performance detection device 1000 of the voiceprint recognition system includes: a first acquisition module 1001, a second acquisition module 1002, a sending module 1003, and a third acquisition module 1004, where:
[0129] The first acquisition module 1001 is configured to acquire multiple character audios from an audio database, and perform splicing processing on the multiple character audios to obtain a first attack voiceprint, where the character audio is an audio segment corresponding to a single character;
[0130] The second acquisition module 1002 is configured to acquire the voiceprint feature of a target user, and generate a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature;
[0131] The sending module 1003 is configured to send the first attack voiceprint and the second attack voiceprint to the voiceprint recognition system to obtain an identification response of the voiceprint recognition system;
[0132] The third acquisition module 1004 is configured to obtain a performance detection result of the voiceprint recognition system according to the identification response.
[0133] In the embodiments of the present application, the first attack voiceprint includes a replay voiceprint. The first acquisition module is specifically configured to: randomly acquire the multiple character audios from the audio database, and perform splicing processing on the randomly acquired multiple character audios to obtain the replay voiceprint.
[0134] In the embodiments of the present application, the first attack voiceprint includes a constructed voiceprint. The first acquisition module is specifically configured to: acquire attack text content, where the attack text content includes multiple text characters; acquire the multiple character audios corresponding to the text characters from the audio database; and perform splicing processing on the acquired multiple character audios according to the arrangement order of the multiple text characters in the attack text content to obtain the constructed voiceprint.
[0135] In the embodiments of the present application, the second acquisition module is specifically configured to: input the voiceprint of the target user into a feature extraction neural network to obtain a voiceprint feature vector of the target user; perform fusion processing on the voiceprint feature vector and voiceprint text to obtain a Mel spectrum; and perform transformation processing on the Mel spectrum to obtain the second attack voiceprint.
[0136] In the embodiment of the present application, the second acquisition module is specifically configured to: segment the voiceprint of the target user to obtain a plurality of voiceprint segments, input the plurality of voiceprint segments into the feature extraction neural network respectively to obtain the voiceprint feature vectors corresponding to the voiceprint segments, and take the average value of the voiceprint feature vectors corresponding to the voiceprint segments to obtain the voiceprint feature vector of the target user.
[0137] In the embodiment of the present application, as Figure 11 shown, it shows a block diagram of a performance detection device for a second voiceprint recognition system provided by the embodiment of the present application. The performance detection device 1100 of the voiceprint recognition system further includes: an acquisition module 1005, a cleaning module 1006, a segmentation module 1007, and a construction module 1008, where:
[0138] The acquisition module 1005 is configured to acquire the original voiceprint;
[0139] The cleaning module 1006 is configured to perform a cleaning process on the original voiceprint to remove the noise in the original voiceprint and obtain a candidate voiceprint;
[0140] The segmentation module 1007 is configured to perform a segmentation process on the candidate voiceprint to obtain a plurality of character audios;
[0141] The construction module 1008 is configured to construct the audio database based on the plurality of character audios.
[0142] In the embodiment of the present application, as Figure 12 shown, it shows a block diagram of a performance detection device for a third voiceprint recognition system provided by the embodiment of the present application. The performance detection device 1200 of the voiceprint recognition system further includes: a fourth acquisition module 1009 and a training module 1010, where:
[0143] The fourth acquisition module 1009 is configured to acquire a training sample set, where the training sample set includes sample voiceprints and voiceprint labels corresponding to the sample voiceprints, and the voiceprint labels are used to indicate whether the sample voiceprints are normal voiceprints or malicious voiceprints;
[0144] The training module 1010 is configured to train a classification neural network based on the training sample set. The classification neural network includes a feature extraction layer; and use the feature extraction layer included in the classification neural network as the feature extraction neural network.
[0145] In the embodiments of the present application, the third acquisition module is specifically configured to: if the recognition response corresponding to the reproduced voiceprint is successful, determine that the performance of the voiceprint recognition system is at the first level; if the recognition response corresponding to the reproduced voiceprint fails, and the recognition response corresponding to the constructed voiceprint is successful, determine that the performance of the voiceprint recognition system is at the second level; if the recognition responses corresponding to the reproduced voiceprint and the constructed voiceprint are both successful, and the recognition response corresponding to the second attack voiceprint fails, determine that the performance of the voiceprint recognition system is at the third level; if the recognition responses corresponding to the reproduced voiceprint, the constructed voiceprint, and the second attack voiceprint all fail, determine that the performance of the voiceprint recognition system is at the fourth level; the performance of the voiceprint recognition system represented by the first level, the second level, the third level, and the fourth level increases in sequence.
[0146] For the specific limitations of the performance detection device of the voiceprint recognition system, reference can be made to the limitations of the performance detection method of the voiceprint recognition system in the foregoing text, which will not be elaborated here. Each module in the above-mentioned performance detection device of the voiceprint recognition system can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0147] In the embodiments of the present application, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 13 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the performance detection data of the voiceprint recognition system. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a performance detection method of a voiceprint recognition system.
[0148] Those skilled in the art can understand that Figure 13 the structure shown in
[0149] In one embodiment of the present application, a computer device is provided. The computer device can be a server. The computer device includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0150] Obtain a plurality of character audios from an audio database, and perform splicing processing on the plurality of character audios to obtain a first attack voiceprint. The character audio is an audio segment corresponding to a single character; obtain the voiceprint feature of the target user, and generate a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature; send the first attack voiceprint and the second attack voiceprint to a voiceprint recognition system to obtain an identification response of the voiceprint recognition system; obtain a performance detection result of the voiceprint recognition system according to the identification response.
[0151] In one embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0152] Randomly obtain the plurality of character audios from the audio database, and perform splicing processing on the randomly obtained plurality of character audios to obtain the replay voiceprint.
[0153] In one embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0154] Obtain attack text content, which includes a plurality of text characters; obtain the plurality of character audios respectively corresponding to the text characters from the audio database; perform splicing processing on the obtained plurality of character audios according to the arrangement order of the plurality of text characters in the attack text content to obtain the constructed voiceprint.
[0155] In one embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0156] Collect an original voiceprint; perform cleaning processing on the original voiceprint to remove noise in the original voiceprint to obtain a candidate voiceprint; perform segmentation processing on the candidate voiceprint to obtain a plurality of the character audios; construct the audio database based on the plurality of character audios.
[0157] In one embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0158] Input the voiceprint of the target user into a feature extraction neural network to obtain a voiceprint feature vector of the target user; perform fusion processing on the voiceprint feature vector and voiceprint text to obtain a Mel spectrogram; perform conversion processing on the Mel spectrogram to obtain the second attack voiceprint.
[0159] In one embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0160] Segment the voiceprint of the target user to obtain multiple voiceprint segments, input the multiple voiceprint segments into a feature extraction neural network respectively to obtain voiceprint feature vectors corresponding to the respective voiceprint segments, and take the average value of the voiceprint feature vectors corresponding to the respective voiceprint segments to obtain the voiceprint feature vector of the target user.
[0161] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0162] Obtain a training sample set, where the training sample set includes sample voiceprints and voiceprint labels corresponding to the sample voiceprints, and the voiceprint labels are used to indicate whether the sample voiceprints are normal voiceprints or malicious voiceprints; train a classification neural network based on the training sample set, where the classification neural network includes a feature extraction layer; use the feature extraction layer included in the classification neural network as the feature extraction neural network.
[0163] In an embodiment of the present application, when the processor executes the computer program, the following steps are further implemented:
[0164] If the recognition response corresponding to the replay voiceprint is successful, determine that the performance of the voiceprint recognition system is at the first level; if the recognition response corresponding to the replay voiceprint fails and the recognition response corresponding to the constructed voiceprint is successful, determine that the performance of the voiceprint recognition system is at the second level; if the recognition responses corresponding to the replay voiceprint and the constructed voiceprint are both successful and the recognition response corresponding to the second attack voiceprint fails, determine that the performance of the voiceprint recognition system is at the third level; if the recognition responses corresponding to the replay voiceprint, the constructed voiceprint, and the second attack voiceprint all fail, determine that the performance of the voiceprint recognition system is at the fourth level; the performance of the voiceprint recognition system represented by the first level, the second level, the third level, and the fourth level increases in sequence.
[0165] The computer device provided in the embodiment of the present application has the same implementation principle and technical effects as those in the above method embodiment, and will not be elaborated here.
[0166] In an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0167] Obtain multiple character audios from the audio database, splice the multiple character audios to obtain a first attack voiceprint, where the character audio is an audio segment corresponding to a single character; obtain the voiceprint feature of the target user, and generate a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature; send the first attack voiceprint and the second attack voiceprint to the voiceprint recognition system to obtain the recognition response of the voiceprint recognition system; obtain the performance detection result of the voiceprint recognition system according to the recognition response.
[0168] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are implemented:
[0169] Randomly obtain the multiple character audios from the audio database, and splice the randomly obtained multiple character audios to obtain the replay voiceprint.
[0170] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are implemented:
[0171] Obtain attack text content, where the attack text content includes multiple text characters; obtain the multiple character audios respectively corresponding to the text characters from the audio database; splice the obtained multiple character audios in the arrangement order of the multiple text characters in the attack text content to obtain the constructed voiceprint.
[0172] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are implemented:
[0173] Collect the original voiceprint; perform cleaning processing on the original voiceprint to remove the noise in the original voiceprint to obtain a candidate voiceprint; perform segmentation processing on the candidate voiceprint to obtain multiple character audios; construct the audio database based on the multiple character audios.
[0174] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are implemented:
[0175] Input the voiceprint of the target user into the feature extraction neural network to obtain the voiceprint feature vector of the target user; fuse the voiceprint feature vector with voiceprint text to obtain a Mel spectrogram; perform transformation processing on the Mel spectrogram to obtain the second attack voiceprint.
[0176] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are implemented:
[0177] Perform segmentation processing on the voiceprint of the target user to obtain multiple voiceprint segments, input the multiple voiceprint segments into the feature extraction neural network respectively to obtain the voiceprint feature vectors corresponding to the voiceprint segments, and take the average value of the voiceprint feature vectors corresponding to the voiceprint segments to obtain the voiceprint feature vector of the target user.
[0178] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are implemented:
[0179] Obtain a training sample set, where the training sample set includes sample voiceprints and voiceprint labels corresponding to the sample voiceprints, and the voiceprint labels are used to indicate whether the sample voiceprints are normal voiceprints or malicious voiceprints; train a classification neural network based on the training sample set, where the classification neural network includes a feature extraction layer; use the feature extraction layer included in the classification neural network as the feature extraction neural network.
[0180] In an embodiment of the present application, when the computer program is executed by a processor, the following steps are implemented:
[0181] If the recognition response corresponding to the replayed voiceprint is successful, determine that the performance of the voiceprint recognition system is at the first level; if the recognition response corresponding to the replayed voiceprint fails, and the recognition response corresponding to the constructed voiceprint is successful, determine that the performance of the voiceprint recognition system is at the second level; if the recognition responses corresponding to the replayed voiceprint and the constructed voiceprint are both successful, and the recognition response corresponding to the second attack voiceprint fails, determine that the performance of the voiceprint recognition system is at the third level; if the recognition responses corresponding to the replayed voiceprint, the constructed voiceprint, and the second attack voiceprint all fail, determine that the performance of the voiceprint recognition system is at the fourth level; the performance of the voiceprint recognition system characterized by the first level, the second level, the third level, and the fourth level increases in sequence.
[0182] The computer-readable storage medium provided in this embodiment has the same implementation principle and technical effects as the above method embodiment, and will not be elaborated here.
[0183] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in M forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Symchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0184] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0185] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for detecting the performance of a voiceprint recognition system, characterized in that, The method includes: Obtaining a plurality of character audios from an audio database, and performing splicing processing on the plurality of character audios to obtain a first attack voiceprint, where the character audio is an audio segment corresponding to a single character; wherein, the first attack voiceprint includes a replay voiceprint and a constructed voiceprint; Obtaining the voiceprint feature of a target user, and generating a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature; Sending the first attack voiceprint and the second attack voiceprint to a voiceprint recognition system to obtain an identification response of the voiceprint recognition system; Obtaining a performance detection result of the voiceprint recognition system according to the identification response; wherein, the performance detection result of the voiceprint recognition system is a performance level obtained by judging the identification response of the voiceprint recognition system to the first attack voiceprint and the second attack voiceprint; the performance levels include: a first level, a second level, a third level, and a fourth level; Wherein, the obtaining a plurality of character audios from an audio database, and performing splicing processing on the plurality of character audios to obtain a first attack voiceprint includes: Obtaining attack text content, where the attack text content includes a plurality of text characters; Obtaining the plurality of character audios respectively corresponding to the respective text characters from the audio database; Performing splicing processing on the obtained plurality of character audios according to the arrangement order of the plurality of text characters in the attack text content to obtain the constructed voiceprint.
2. The method according to claim 1, wherein The obtaining a plurality of character audios from an audio database, and performing splicing processing on the plurality of character audios to obtain a first attack voiceprint includes: Randomly obtaining the plurality of character audios from the audio database, and performing splicing processing on the randomly obtained plurality of character audios to obtain the replay voiceprint.
3. The method according to any one of claims 1 to 2, characterized in that, The method further includes: Collecting an original voiceprint; Performing cleaning processing on the original voiceprint to remove noise in the original voiceprint to obtain a candidate voiceprint; Performing segmentation processing on the candidate voiceprint to obtain a plurality of the character audios; Constructing the audio database based on the plurality of character audios.
4. The method according to claim 1, wherein The obtaining the voiceprint feature of a target user, and generating a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature includes: Inputting the voiceprint of the target user into a feature extraction neural network to obtain a voiceprint feature vector of the target user; Fusing the voiceprint feature vector of the target user with voiceprint text to obtain a Mel spectrogram; Performing transformation processing on the Mel spectrogram to obtain the second attack voiceprint.
5. The method according to claim 4, wherein The inputting the voiceprint of the target user into a feature extraction neural network to obtain a voiceprint feature vector of the target user includes: Performing segmentation processing on the voiceprint of the target user to obtain a plurality of voiceprint segments; Inputting the plurality of voiceprint segments into the feature extraction neural network respectively to obtain voiceprint feature vectors corresponding to the respective voiceprint segments; Taking the average value of the voiceprint feature vectors corresponding to the respective voiceprint segments to obtain the voiceprint feature vector of the target user.
6. The method according to claim 4 or 5, characterized in that, Before the inputting the voiceprint of the target user into a feature extraction neural network, the method further includes: Obtain a training sample set, where the training sample set includes sample voiceprints and voiceprint labels corresponding to the sample voiceprints, and the voiceprint labels are used to indicate whether the sample voiceprints are normal voiceprints or malicious voiceprints; Train a classification neural network based on the training sample set, where the classification neural network includes a feature extraction layer; Use the feature extraction layer included in the classification neural network as the feature extraction neural network.
7. The method according to claim 1, wherein The obtaining the performance detection result of the voiceprint recognition system according to the recognition response includes: If the recognition response corresponding to the replayed voiceprint is successful, determine that the performance of the voiceprint recognition system is at the first level; If the recognition response corresponding to the replayed voiceprint fails, and the recognition response corresponding to the constructed voiceprint is successful, determine that the performance of the voiceprint recognition system is at the second level; If the recognition responses corresponding to both the replayed voiceprint and the constructed voiceprint are successful, and the recognition response corresponding to the second attack voiceprint fails, determine that the performance of the voiceprint recognition system is at the third level; If the recognition responses corresponding to the replayed voiceprint, the constructed voiceprint, and the second attack voiceprint all fail, determine that the performance of the voiceprint recognition system is at the fourth level; The performance of the voiceprint recognition system represented by the first level, the second level, the third level, and the fourth level increases in sequence.
8. A performance detection device for a voiceprint recognition system, characterized in that The device includes: A first acquisition module, configured to acquire multiple character audios from an audio database, and perform splicing processing on the multiple character audios to obtain a first attack voiceprint, where the character audio is an audio segment corresponding to a single character; wherein, the first attack voiceprint includes a replayed voiceprint and a constructed voiceprint; A second acquisition module, configured to acquire the voiceprint feature of a target user, and generate a second attack voiceprint for simulating the voice of the target user based on the voiceprint feature; A sending module, configured to send the first attack voiceprint and the second attack voiceprint to the voiceprint recognition system to obtain the recognition response of the voiceprint recognition system; A third acquisition module, configured to obtain the performance detection result of the voiceprint recognition system according to the recognition response; wherein, the performance detection result of the voiceprint recognition system is a performance level obtained by judging the recognition response of the voiceprint recognition system to the first attack voiceprint and the second attack voiceprint; the performance level includes: the first level, the second level, the third level, and the fourth level; Among them, the first acquisition module is specifically configured to: acquire attack text content, where the attack text content includes multiple text characters; acquire the multiple character audios corresponding to the respective text characters from the audio database; perform splicing processing on the acquired multiple character audios in the arrangement order of the multiple text characters in the attack text content to obtain the constructed voiceprint.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.