Data processing method and device, computing equipment and storage medium

By dividing the voiceprint samples into sub-samples and calculating the similarity, the problems of large amount and low efficiency of anti-sample detection in the prior art are solved, efficient and accurate anti-sample detection are achieved, and the security of social and economic activities and commercial transactions are ensured.

CN120067845APending Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311626763.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing adversarial sample detection technology has too much calculation volume and low operating efficiency, making it difficult to accurately identify adversarial samples, which affects the security of social and economic activities and commercial transactions.

Method used

By dividing the voiceprint sample into multiple subsamples and calculating the similarity between the subsamples and the standard sample, it is determined whether the voiceprint sample is an adversarial sample based on the similarity.

Benefits of technology

This method reduces calculation and operation steps, reduces calculation amount, improves operation efficiency, and can accurately determine whether the soundprint sample is an adversarial sample, ensuring the accuracy and safety of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067845A_ABST
    Figure CN120067845A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method, a data processing device, computing equipment and a computer readable storage medium. The method is applied to a computing device and comprises the steps of obtaining a voiceprint sample; segmenting the voiceprint sample into a plurality of sub-samples; calculating the similarity between the plurality of sub-samples and a standard sample stored in the computing device; and judging whether the voiceprint sample is an adversarial sample based on the similarity. According to the method and the device, whether the voiceprint sample is the confrontation sample or not can be accurately judged through relatively simple calculation, the judgment cost is reduced, and the method and the device have relatively high applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biometric technologies, and particularly to a data processing method, a data processing device, a computing device, and a computer-readable storage medium. Background Art

[0002] In this field, there are detection technologies for detecting whether a voiceprint sample is an adversarial sample. An adversarial sample refers to a voiceprint sample that is not the voice of the speaker himself / herself, and is used to deceive a voiceprint recognition model, causing the model to misjudge a voice that is not spoken by the speaker as the true voice of the speaker, thereby illegally passing relevant verifications or obtaining corresponding permissions. Existing adversarial sample detection technologies all have the problems of excessive computational complexity and low operating efficiency. Therefore, there is an urgent need in this field for an adversarial sample detection technology with a relatively low computational complexity, a relatively high operating efficiency, and the ability to ensure the detection accuracy, which can accurately identify adversarial samples at a relatively low cost, thereby ensuring the security of social and economic activities and commercial transactions. Summary of the Invention

[0003] To this end, this application is committed to providing a data processing method, a data processing device, a computing device, and a computer-readable storage medium, which can accurately determine whether a voiceprint sample is an adversarial sample with relatively simple calculations, reduce the judgment cost, and have high applicability.

[0004] In one aspect, this application provides a data processing method, which is applied to a computing device. The method includes: obtaining a voiceprint sample; splitting the voiceprint sample into multiple sub-samples; calculating the similarity between the multiple sub-samples and a standard sample stored in the computing device; and determining whether the voiceprint sample is an adversarial sample based on the similarity.

[0005] According to this aspect, for an adversarial sample, the similarity situation between the split sub-samples and the standard sample is different from that of a voiceprint sample from the speaker's true voice. Utilizing this point, it is possible to effectively determine whether a voiceprint sample is an adversarial sample. By splitting the voiceprint sample into multiple sub-samples and determining the similarity between the sub-samples and the standard sample to determine whether the voiceprint sample is an adversarial sample, the required calculations and operation steps are fewer. Therefore, the computational complexity is less, and the operating efficiency is higher.

[0006] In a specific embodiment of this application, determining whether the voiceprint sample is an adversarial sample based on the similarity includes: determining that the voiceprint sample is an adversarial sample when the number of sub-samples with a similarity greater than the similarity threshold is greater than or equal to the preset number threshold; and determining that the voiceprint sample is not an adversarial sample when the number of sub-samples with a similarity greater than the similarity threshold is less than the preset number threshold.

[0007] According to this embodiment, by determining whether the number of subsamples with a similarity greater than a threshold in the subsamples is higher than a preset number threshold, it is possible to determine whether it is an adversarial sample, and the voiceprint sample can be accurately detected with a relatively simple calculation process. The inventors of this application found that the vast majority of subsamples of adversarial samples have a high similarity to the standard sample, which is different from real samples. Based on this discovery, the above detection method is proposed in this application, which can detect adversarial samples only by judging the similarity and through simple mathematical statistics, improving the operation efficiency and ensuring the detection accuracy.

[0008] In a particular embodiment of this application, the number of multiple subsamples is determined according to the security requirement level. In the case of a higher security requirement level, the number of subsamples is larger.

[0009] According to this embodiment, in a scenario with a higher security requirement level, the more subsamples are segmented, the more accurately it can be determined whether the number of subsamples with a similarity greater than the threshold exceeds the normal level, and thus it is possible to more accurately determine the adversarial sample.

[0010] In a particular embodiment of this application, the overlap degree between every two adjacent subsamples among multiple subsamples is determined according to the security requirement level. In the case of a higher security requirement level, the overlap degree between two adjacent subsamples is larger.

[0011] According to this embodiment, in a scenario with a higher security requirement level, the larger the overlap degree between two adjacent subsamples, the more accurate the judgment of the number of subsamples with a similarity greater than the threshold. The same segment is detected more times, making the detection result more accurate.

[0012] In a particular embodiment of this application, the length of each subsample among multiple subsamples is between 1 millisecond and half of the length of the voiceprint sample.

[0013] According to this embodiment, the length of the subsample has a range. The inventors of this application found through research and experiments that for subsamples less than 1 millisecond, the detection efficiency is low, and for subsamples exceeding half of the length of the voiceprint sample, the detection results do not differ much. Therefore, setting the length of the subsample within this range can ensure the detection efficiency and detection accuracy.

[0014] In a particular embodiment of this application, multiple subsamples do not completely cover the voiceprint sample.

[0015] According to this embodiment, all subsamples do not completely cover the voiceprint sample, which can speed up the detection speed. There is no need to detect the entire length of the voiceprint sample, and only a segmented subsample needs to be detected, which can reduce the calculation amount and improve the detection efficiency.

[0016] In a particular embodiment of the present application, the voiceprint sample is sliced into multiple sub-samples, including: cutting at the amplitude low points where the amplitude of the voiceprint sample is less than the low point threshold; or, cutting the voiceprint sample such that the obtained sub-samples cover the amplitude high points where the amplitude of the voiceprint sample is greater than the high point threshold.

[0017] According to this embodiment, cutting at the amplitude low points or making the cut sub-samples cover the high points is conducive to making the sub-samples cover the main part of the speaker's voice. The inventors of the present application found that the amplitude low points often occur when the speaker pauses. Cutting at this place can cut out a complete piece of language, which is conducive to detecting the speaker's voice completely and avoiding too much part without the speaker's voice in the sub-samples, resulting in low detection efficiency.

[0018] In a particular embodiment of the present application, before determining that the voiceprint sample is an adversarial sample when the number of sub-samples with a similarity greater than the similarity threshold is greater than or equal to the preset number threshold, it further includes: for multiple true voiceprint samples, each true voiceprint sample is sliced into a first number of true sub-samples, and the second number of true sub-samples with a similarity greater than the similarity threshold between the true voiceprint sample and the standard sample is calculated, where the first number is variable for different true voiceprint samples; for multiple adversarial voiceprint samples, each adversarial voiceprint sample is sliced into a third number of adversarial sub-samples, and the fourth number of adversarial sub-samples with a similarity greater than the similarity threshold between the adversarial voiceprint sample and the standard sample is calculated, where the third number is variable for different adversarial voiceprint samples; according to the second number and the fourth number, the preset number threshold is determined.

[0019] According to this embodiment, by testing known true samples and adversarial samples, a more accurate number threshold is obtained, which can improve the accuracy of the number threshold and avoid blindly obtaining the preset number threshold, thus causing the judgment of adversarial samples to lose objectivity and accuracy.

[0020] In a particular embodiment of the present application, determining the preset number threshold according to the second number and the fourth number includes: for multiple true voiceprint samples, calculating the maximum value of the ratio of the second number to the first number; for multiple adversarial voiceprint samples, calculating the minimum value of the ratio of the fourth number to the third number; according to the smaller value of the maximum value and the minimum value, determining the ratio of the preset number threshold to the number of multiple sub-samples; according to the ratio, determining the preset number threshold.

[0021] According to this embodiment, by statistically analyzing the ratio between the sub-samples of known true samples and adversarial samples and the sub-samples with a similarity greater than the similarity threshold, a more accurate preset number threshold is obtained, which is conducive to obtaining a more objective and accurate number threshold, thereby making the judgment of adversarial samples more accurate and reliable.

[0022] In a particular embodiment of the present application, before calculating the similarity between multiple sub-samples and a standard sample stored in a computing device, it further includes: identifying the features of the sub-samples, where the features include one or more of amplitude, frequency, time length, language, and accent; and determining the standard sample according to the features.

[0023] According to this embodiment, there may be multiple standard samples. In order to select a more suitable standard sample for comparison, the features of the sub-samples can be judged first, and then a standard sample with features similar to those of the sub-samples can be selected from multiple standard samples to calculate the similarity, which can more accurately judge the similarity and make the detection effect more precise.

[0024] In a particular embodiment of the present application, the voiceprint sample includes a voiceprint sample for inputting into a voiceprint recognition model for recognition. Calculating the similarity between multiple sub-samples and a standard sample stored in a computing device includes: inputting each sub-sample into the voiceprint recognition model to obtain the similarity between the sub-sample and the standard sample.

[0025] According to this embodiment, since the voiceprint sample is used to input into the voiceprint recognition model for verification or detection, the detection of the sub-samples also uses this voiceprint recognition model, which is beneficial to improving the detection efficiency, avoiding using a new model to detect the sub-samples separately, saving the development cost, and improving the applicability.

[0026] In a particular embodiment of the present application, in the case where it is determined that the voiceprint sample is not an adversarial sample, voiceprint recognition is performed on the voiceprint sample to determine whether the voiceprint sample comes from the real voice of the speaker.

[0027] According to this embodiment, judging whether it is an adversarial sample before performing voiceprint sample recognition is beneficial to improving the accuracy of voiceprint sample recognition.

[0028] On the other hand, the present application provides a data processing device, which is applied to a computing device. The device includes: an acquisition module for acquiring a voiceprint sample; a segmentation module for segmenting the voiceprint sample into multiple sub-samples; a calculation module for calculating the similarity between the multiple sub-samples and a standard sample stored in the computing device; and a judgment module for judging whether the voiceprint sample is an adversarial sample based on the similarity.

[0029] In a particular embodiment of the present application, the judgment module is further configured to: determine that the voiceprint sample is an adversarial sample when the number of sub-samples with a similarity greater than the similarity threshold is greater than or equal to the preset number threshold; and determine that the voiceprint sample is not an adversarial sample when the number of sub-samples with a similarity greater than the similarity threshold is less than the preset number threshold.

[0030] In a particular embodiment of the present application, the number of multiple sub-samples is determined according to the security requirement level. The higher the security requirement level, the greater the number of sub-samples.

[0031] In a particular embodiment of the present application, the overlap degree between every two adjacent sub-samples among the multiple sub-samples is determined according to the security requirement level. The higher the security requirement level, the greater the overlap degree between two adjacent sub-samples.

[0032] In a particular embodiment of the present application, the length of each sub-sample among the multiple sub-samples is between 1 millisecond and half of the length of the voiceprint sample.

[0033] In a particular embodiment of the present application, the multiple sub-samples do not completely cover the voiceprint sample.

[0034] In a particular embodiment of the present application, the segmentation module is further configured to: perform cutting at the amplitude low point where the amplitude of the voiceprint sample is less than the low point threshold; or, perform cutting on the voiceprint sample such that the sub-samples obtained by cutting cover the amplitude high points where the amplitude of the voiceprint sample is greater than the high point threshold.

[0035] In a particular embodiment of the present application, the judgment module is further configured to: for multiple real voiceprint samples, cut each real voiceprint sample into a first number of real sub-samples, and calculate a second number of real sub-samples of the real voiceprint sample whose similarity to the standard sample is greater than the similarity threshold, where the first number is variable for different real voiceprint samples; for multiple adversarial voiceprint samples, cut each adversarial voiceprint sample into a third number of adversarial sub-samples, and calculate a fourth number of adversarial sub-samples of the adversarial voiceprint sample whose similarity to the standard sample is greater than the similarity threshold, where the third number is variable for different adversarial voiceprint samples; determine a preset number threshold according to the second number and the fourth number.

[0036] In a particular embodiment of the present application, the judgment module is further configured to: for multiple real voiceprint samples, calculate the maximum value of the ratio of the second number to the first number; for multiple adversarial voiceprint samples, calculate the minimum value of the ratio of the fourth number to the third number; determine the ratio of the preset number threshold to the number of multiple sub-samples according to the smaller value of the maximum value and the minimum value; determine the preset number threshold according to the ratio.

[0037] In a particular embodiment of the present application, the device is further configured to: identify the features of the sub-samples, where the features include one or more of amplitude, frequency, time length, language, accent; determine the standard sample according to the features.

[0038] In a particular embodiment of the present application, the voiceprint sample includes a voiceprint sample for inputting into a voiceprint recognition model for recognition, and the calculation module is further configured to: input each sub-sample into the voiceprint recognition model to obtain the similarity between the sub-sample and the standard sample.

[0039] On the other hand, the present application provides a computing device, which includes a processor and a memory. The processor is used to execute a computer program stored in the memory to implement the above data processing method.

[0040] On the other hand, the present application provides a computer-readable storage medium, which stores a computer program. The computer program is used to execute the above data processing method.

[0041] On the other hand, the present application provides a computer program product, which includes program code. When the computer runs the computer program product, the computer is enabled to implement the above data processing method.

[0042] Any of the above-provided data processing method, data processing device, computing device, computer-readable storage medium or computer program product is used to execute the data processing method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding solutions in the corresponding method provided above, and will not be elaborated here. Description of the Drawings

[0043] Hereinafter, the specific embodiments of the present application will be described in detail with reference to the drawings, where:

[0044] Figure 1 Shows a schematic architecture diagram of a data processing method according to an embodiment of the present application;

[0045] Figure 2 Shows a schematic flow diagram of a data processing method according to an embodiment of the present application;

[0046] Figure 3 Shows according to Figure 2 A schematic diagram of the segmentation steps of the data processing method of the embodiment;

[0047] Figure 4 Shows a schematic flow diagram of a data processing method according to an embodiment of the present application;

[0048] Figure 5 Shows a schematic structural diagram of a data processing device according to an embodiment of the present application;

[0049] Figure 6 Shows a schematic structural diagram of a computing device according to an embodiment of the present application. Detailed Description of the Embodiments

[0050] To enable those skilled in the art to more clearly understand the concepts and ideas of this application, the following describes this application in detail with specific embodiments. It should be understood that the embodiments given herein are only a part of all possible embodiments of this application. After reading the specification of this application, those skilled in the art are capable of making improvements, modifications, or substitutions to part or all of the following embodiments, and these improvements, modifications, or substitutions are also included within the scope protected by this application.

[0051] In this article, the terms "a", "an", and other similar words do not intend to mean that there is only one such thing, but rather that the relevant description is only directed to one of such things, and such things may have one or more. In this article, the terms "comprise", "include", and other similar words are intended to represent a logical relationship and should not be regarded as representing a spatial structural relationship. For example, "A includes B" is intended to mean that logically B belongs to A, rather than meaning that B is located inside A in terms of space. Additionally, the meanings of the terms "comprise", "include", and other similar words should be regarded as open-ended rather than closed. For example, "A includes B" is intended to mean that B belongs to A, but B does not necessarily constitute the whole of A, and A may also include other elements such as C, D, E, etc.

[0052] In this article, the terms "first", "second", and other similar words do not intend to imply any order, quantity, or importance, but are only used to distinguish different elements. In this article, the terms "embodiment", "the present embodiment", "one embodiment", "a single embodiment" do not mean that the relevant description only applies to a specific embodiment, but rather that these descriptions may also apply to one or more other embodiments. Those skilled in the art should understand that in this article, any description made for a particular embodiment can be substituted, combined, or otherwise combined with the relevant descriptions in one or more other embodiments, and the new embodiments thus produced are easily conceivable by those skilled in the art and fall within the scope of protection of this application.

[0053] In various embodiments of this application, a voiceprint may refer to the identity information contained in the sound emitted by a human. For example, a voiceprint may be the sound wave spectrum carrying speech information displayed by an electroacoustic instrument. A voiceprint not only has specificity but also has the characteristic of relative stability. After adulthood, a person's voice can remain relatively stable for a long time. Whether the speaker deliberately imitates the voices and tones of others or whispers softly, even if the imitation is extremely vivid, their voiceprint remains unchanged. Based on these two characteristics of the voiceprint, investigators can conduct inspection and comparison through voiceprint identification technology to quickly identify criminals. Voiceprints can also be used in security verification scenarios to identify whether the current operation is made by an authorized person.

[0054] In various embodiments of the present application, a voiceprint sample may refer to a voice sample used for voiceprint recognition. A real voiceprint sample is the voice emitted by a speaker, and a fake voiceprint sample may be the voice emitted by another speaker or a sample that is artificially synthesized and not a human voice. In various embodiments of the present application, adversarial sample detection may refer to the detection process and technology for detecting whether a voiceprint sample is a real sample or an adversarial sample, and is used to judge the authenticity of the voiceprint sample. There is a difference between adversarial sample detection and voiceprint detection that directly verifies the voiceprint identity. The former is used to judge whether the voiceprint sample is real, and the latter is used to judge whether the voiceprint is the voice of the target speaker, that is, to judge whether the identity information contained in the voiceprint points to a specific speaker. In other words, even if it is proved through adversarial sample detection that it is not an adversarial sample or an adversarial sample, it may not necessarily pass the voiceprint detection and thus obtain the corresponding permission or pass the corresponding verification. In some embodiments, adversarial sample detection may be a part of or a pre - procedure for voiceprint detection. Only by passing the adversarial sample detection can the next step of voiceprint detection or identity verification be carried out. Other combination methods of adversarial sample detection and voiceprint detection are also conceivable, and the present application is not limited thereto.

[0055] In recent years, with the rapid development of deep learning, deep learning has become one of the most common technologies in artificial intelligence, affecting and changing people's lives in all aspects. Typical applications include smart home, autonomous driving, speech recognition, voiceprint recognition and other fields. Voiceprint recognition models are vulnerable to adversarial attacks, that is, adding slight perturbations to the input samples can lead to abnormal behaviors of the models. Adversarial samples are formed by deliberately adding subtle interferences to the input data, and such samples cause the model to give a wrong output with high confidence. Simply put, adversarial samples cause deep learning models to produce classification errors by superimposing carefully constructed perturbations that are difficult for humans to detect on the elemental data. At present, the adversarial attack and defense in the direction of voiceprint authentication are relatively lacking, and there is an urgent need for a comprehensive, efficient and prior - knowledge - independent adversarial sample detection method.

[0056] Traditional voiceprint adversarial sample defense methods include adversarial sample detection, adversarial training, verifiable robustness, etc. Among them, the improvement effect of adversarial training on the robustness of the model depends on the diversity of the generated adversarial samples. Moreover, the attacker can still generate adversarial samples through stronger attacks, and the accuracy of the model itself may decrease; verifiable robustness consumes a large amount of computing resources and takes longer time in the process of calculating the robust space of the samples; while the detection - based method not only has a lower computational cost, but also can improve the robustness of the model while maintaining the original performance of the model.

[0057] Voiceprint authentication is a biometric authentication technology that is currently widely used in identity authentication scenarios such as account login, intelligent assistants, and financial payments. However, voiceprint authentication models based on deep neural networks are at risk of adversarial attacks, that is, adding tiny adversarial noise to non-personal audio can achieve the effect of allowing attackers to pass the identity authentication. The adversarial attack defense methods for voiceprint authentication systems in this field are difficult to defend against multiple attack methods simultaneously, especially black-box transfer attacks. And to ensure effectiveness, efficiency is sacrificed, and prior knowledge of the attack method is required at the same time.

[0058] In some technologies in this field, a detection method based on noise reduction is proposed. By calculating the difference in voiceprint similarity scores before and after denoising a segment of audio, adversarial samples are detected: this detection method uses a Vocoder (voice coder) model to resynthesize the audio from the spectrum of the audio to achieve noise reduction. After denoising a real sample, the audio loss is small, and the change in the voiceprint similarity score of the audio is not obvious; after denoising an adversarial sample, the audio loss is large, and the voiceprint similarity score of the audio changes significantly. By comparing the difference in voiceprint similarity scores before and after audio denoising, it is determined whether the audio is an adversarial sample. The problem is that this method cannot effectively detect black-box adversarial samples, and the detection success rate for black-box transfer attacks is not high. At the same time, this method needs to call the ASV (automatic speaker verification) model twice and the Vocoder model once to detect a segment of audio, with low algorithm operation efficiency and high operation energy consumption.

[0059] In some other technologies in this field, a detection method based on the decision boundary is proposed. By using the decision boundary attack method to attack a segment of audio until the labels of the voiceprint model and the detection model change, the proportion of adversarial perturbation is obtained. Because the clean sample audio itself has no noise, a relatively large proportion of adversarial perturbation is required; the adversarial sample audio itself contains adversarial noise, and a relatively small proportion of adversarial perturbation is required. By comparing the obtained proportion of adversarial noise with the threshold, it is determined whether the audio is an adversarial sample. The problem is that this method needs to introduce a detection model to detect a segment of audio, and calling the detection model requires additional storage and computing resources. At the same time, this method needs to add noise to the input sample until the label changes, and multiple calls to the target model and calculation of the backpropagation gradient are required during the addition process, with a large amount of algorithm calculation and low operation efficiency.

[0060] The inventors of the present application found that during the generation of voiceprint adversarial samples, most adversarial attack algorithms add adversarial noise to the audio and then attack the voiceprint authentication system. Under the action of the adversarial noise, even if the adversarial samples are divided into very small segments, the similarity scores between them and the registered person's speech are higher than those of normal speech segments. This is because in order to obtain a higher voiceprint similarity score, the adversarial attack algorithm adds global adversarial noise to the audio. The global adversarial noise not only improves the similarity between the speech part and the registered person's speech, but also improves the similarity between the non-speech part and the registered person's speech, resulting in that even very small segments of adversarial samples can achieve very high similarity scores, much higher than those of normal sample segments.

[0061] Based on this characteristic, some embodiments of the present application propose a slicing-based adversarial sample detection method for a voiceprint authentication system, which calculates the proportion of segments greater than the threshold of the voiceprint authentication model to determine whether the input audio is an adversarial sample.

[0062] Some embodiments of the present application propose a general, efficient and prior-knowledge-independent adversarial text detection method. This method can correctly distinguish adversarial samples and normal samples on the premise of ensuring the accuracy of the model itself, and improve the robustness of the model.

[0063] Some embodiments of the present application construct a slicing-based adversarial sample detection method for a voiceprint authentication model, which improves the model robustness without affecting the accuracy of the original voiceprint model.

[0064] Some embodiments of the present application use the characteristic that the similarity scores of adversarial sample segments are much higher than those of normal samples to detect adversarial samples. This method does not need to use the knowledge of prior adversarial samples, can detect any attack algorithm, and gets rid of the limitation of only being able to detect fixed attack algorithms. This method uses the protected voiceprint model for detection, getting rid of the dependence on an additional detection model. This method requires low power consumption and is easy to deploy.

[0065] Some embodiments of the present application use the characteristic that the similarity between the subsamples cut from the adversarial samples and the target speaker is higher than that of normal samples to cut the input audio into subsamples to detect whether it is an adversarial sample. This method does not require the type of attack, greatly improves the detection efficiency of adversarial samples, and at the same time ensures the detection accuracy. At the same time, this method requires low power consumption and is easy to deploy.

[0066] Figure 1 The schematic diagram of the architecture of the data processing method according to an embodiment of the present application is shown.

[0067] As Figure 1As shown, for the voiceprint sample to be detected, it is first segmented into multiple sub-samples. Sub-sample 1, Sub-sample 2, and Sub-sample 3 are shown in the figure. It should be understood that the number of sub-samples is not limited to this, and the voiceprint sample can be segmented into more sub-samples. The 3 sub-samples shown in the figure are only for illustration. For each sub-sample, its similarity to the standard sample is calculated to obtain the corresponding similarity value. Each sub-sample has a corresponding similarity. For example, in the figure, Sub-sample 1 corresponds to Similarity 1, Sub-sample 2 corresponds to Similarity 2, and Sub-sample 3 corresponds to Similarity 3. The similarity between each sub-sample and the standard sample can have different values. There can be only one standard sample, or there can be multiple standard samples. The standard sample used to judge the similarity of a certain sub-sample can be selected from multiple standard samples. According to the similarity between multiple sub-samples and the standard sample, it can be judged whether the voiceprint sample is an adversarial sample or a genuine sample.

[0068] In one embodiment, after obtaining the similarity between each sub-sample and the standard sample, it can continue to judge whether these similarities are greater than the similarity threshold, and count the number of sub-samples whose similarities are greater than the similarity threshold. After obtaining this number, it is further judged whether this number is greater than the preset number threshold. If it is greater than or equal to the preset number threshold, it can be determined that the voiceprint sample is an adversarial sample. If it is less than the number threshold, it can be determined that the voiceprint sample is a genuine sample.

[0069] In this embodiment, a sub-sample can refer to multiple segments cut from a complete voiceprint sample. A complete voiceprint sample can be all the sounds made by a speaker to pass a certain verification or obtain a certain permission. The cut sub-samples can be multiple small parts of the entire voiceprint sample, which do not have complete semantics or features and cannot be used to pass a certain identity verification.

[0070] In this embodiment, the standard sample can refer to a standard voice sample recorded by the speaker in advance for identity verification. The standard sample can be stored in advance in the computing device to which the technical solution of this application is applied. By comparing with the standard sample, it can be judged whether the currently received speaking voice is the voice of the speaker recorded before. There can be one standard sample, or there can be multiple standard samples. The standard sample can be recorded in advance, or can be obtained by other means such as computer synthesis. The standard sample can be pre-stored in the terminal for verifying the speaker's identity, or the speaker can be required to record it temporarily, or it can be pre-stored elsewhere such as in the cloud.

[0071] In this embodiment, the similarity may refer to the degree of similarity between a sub-sample and a standard sample. Various means can be used to determine the similarity. For example, elements such as the sound frequency and amplitude of the sub-sample can be analyzed to determine the similarity with the corresponding elements of the standard sample. The similarity can be determined by using a voiceprint recognition model. For example, the voiceprint recognition model originally used to recognize the voiceprint sample can be directly used to determine the similarity between the sub-sample and the standard sample. Since the voiceprint recognition model is originally used to determine the similarity between the voiceprint sample and the standard sample, it can definitely be used to determine the similarity between the sub-sample and the standard sample. If the recognition is passed, it means that the similarity between the sub-sample and the standard sample is greater than the threshold. If the recognition fails, it means that the similarity between the sub-sample and the standard sample is lower than the threshold.

[0072] In this embodiment, the similarity threshold may refer to the boundary between judging similarity and dissimilarity between a sub-sample and a standard sample. For example, being higher than the similarity threshold may mean that the sub-sample and the standard sample are voices emitted by the same speaker. Being lower than the similarity threshold, it can be judged that the sub-sample and the standard sample are voices emitted by different speakers, or the sub-sample is not a voice emitted by a human, etc.

[0073] In this embodiment, the preset quantity threshold may refer to the threshold for judging whether the number of sub-samples whose similarity between the sub-sample and the standard sample exceeds the similarity threshold exceeds a reasonable range. Generally speaking, if it is a real voiceprint sample, a considerable part of the sub-samples cut from it cannot pass the voiceprint model verification, or it is impossible to have a similarity sufficient to judge that it belongs to the speaker's voice, because the sub-sample does not have complete sound or semantics, and the similarity judgment result is relatively low. In contrast, for a false voiceprint sample used to deceive the voiceprint recognition model by means such as adding noise, each of its sub-samples will basically have a high similarity, so that the number of sub-samples with a similarity greater than the similarity threshold exceeds the normal value, and this normal value is the preset quantity threshold.

[0074] As an example, taking advantage of the characteristic that the local voiceprint similarity of the adversarial sample audio is higher than that of the normal audio, a segmentation-based adversarial sample detection method for the voiceprint authentication system can be provided, and adversarial samples and normal samples are distinguished by analyzing the similarity between the sub-samples of the input audio and the registered audio. This example is specifically divided into five steps: obtaining the input audio, determining the optimal detection threshold, segmenting the audio to be detected, calculating the similarity score, and detecting adversarial samples.

[0075] In the step of obtaining the input audio, obtain the voiceprint authentication model F as the target protection model, and obtain the threshold t of this voiceprint model. Obtain the input audio to be detected, and perform format conversion on the audio to ensure that the number of audio channels and the audio sampling rate meet the algorithm requirements.

[0076] In the step of determining the optimal detection threshold, by selecting a small number of samples and using a known adversarial sample generation algorithm to generate a batch of adversarial samples, the optimal threshold is determined as the optimal detection threshold τ.

[0077] In the step of splitting the audio to be detected, in order to calculate the local similarity score of the audio to be detected, the audio is first sliced. For the obtained audio to be detected, the input audio is sliced with a window of a fixed length to obtain n subsamples x = [x 1 , x 2 , …, x n , where x i is the i-th subsample obtained by splitting.

[0078] In the step of calculating the similarity score, for the subsamples x = [x 1 , x 2 , …, x n obtained by slicing, the target protection model F is used to calculate the voiceprint similarity score between each subsample and the target speaker (or standard sample). The voiceprint similarity score s i of the subsample x i is calculated by the formula

[0079] s i = F(x i )

[0080] The n subsamples are input into the voiceprint authentication model F to obtain the similarity score s = [s 1 , s 2 , …, s n with the registered audio (or standard sample), where s i is the similarity score between the i-th subsample and the target speaker (or standard sample).

[0081] In the step of detecting adversarial samples, first, according to the similarity score s = [s 1 , s 2 , …, s n obtained in the previous step, calculate the number H of subsamples with scores greater than the voiceprint authentication model threshold t. The calculation method is

[0082]

[0083] Then, based on the number H of subsamples greater than the voiceprint authentication model threshold, calculate the proportion of subsamples greater than the voiceprint authentication model threshold Judge whether the audio to be detected is an adversarial sample according to the predefined threshold (or the optimal detection threshold τ). If If it is greater than a predefined threshold (or the optimal detection threshold τ), the measured audio is an adversarial sample or an adversarial example; if If it is less than a predefined threshold (or the optimal detection threshold τ), the measured audio is a normal sample or a genuine sample.

[0084] Figure 2 The flowchart shows a data processing method according to an embodiment of the present application.

[0085] According to this embodiment, the data processing method includes steps S210 to S240, and each step is described in detail below.

[0086] S210. Obtain a voiceprint sample.

[0087] In this embodiment, a voiceprint sample can be obtained first. Specifically, an input audio to be detected can be obtained, and the format of the audio can be converted to ensure that the number of audio channels and the audio sampling rate meet the algorithm requirements.

[0088] S220. Split the voiceprint sample into multiple sub-samples.

[0089] In this embodiment, splitting the voiceprint sample may refer to intercepting the voiceprint sample in terms of length, and the obtained multiple small-segment audio samples are the sub-samples.

[0090] In this embodiment, in order to calculate the local similarity score of the audio to be measured, the audio is sliced first. For the obtained input audio to be detected, the input audio is sliced with a fixed-length window to obtain multiple sub-samples of the input audio.

[0091] In this embodiment, sub-samples are obtained by splitting the input audio, and the input is fed into the protected model to calculate the voiceprint similarity score with the target speaker to detect adversarial samples. According to this embodiment, there is no need to require the type of attack, and the detection efficiency of adversarial samples can be greatly improved while ensuring the detection accuracy. Since global noise addition is a common feature of adversarial samples, the similarity score between the sub-samples of the adversarial sample and the target speaker will be higher than that of the normal sample. By using this statistical feature to analyze the input audio, adversarial samples of a wide range of types can be detected.

[0092] As an example, the number of multiple sub-samples is determined according to the security requirement level. In the case of a higher security requirement level, the number of sub-samples is larger.

[0093] In this example, the security requirement level can be specific to a particular application scenario. The application scenario can refer to the actual scenario where the voiceprint sample is applied. In this scenario, by identifying whether the voiceprint sample is the voice of the speaker, it is determined whether the current person is the person recorded in the database or has the corresponding permissions. The security requirement level of the application scenario can refer to the degree of security requirements for the application scenario that needs to determine the authenticity of the voiceprint sample. For example, for the mobile phone unlocking scenario, it is necessary to determine whether the person requesting to unlock the phone is the owner of the phone. At this time, the security requirement level is relatively low because only the security of the personal information stored in the mobile phone is involved, and property security is not involved. For the bank transfer scenario, it is necessary to determine whether the person requesting the transfer is the owner of the account. At this time, the security requirement level is relatively high because it involves funds and economic interests. For other scenarios, the number of sub-samples can be appropriately set according to their security requirements. The security requirement level can be preset in the system or adjusted manually or automatically according to different application scenarios. Other methods for obtaining the security requirement level can also be conceived.

[0094] In this example, different segmentation window sizes are used for different scenarios. According to this example, a very small performance loss can be exchanged for a faster calculation speed, or the calculation speed can be exchanged for higher detection accuracy to meet the performance requirements of more application scenarios. Experiments have found that even with a very small number of audio segments, normal samples and adversarial samples can be distinguished. Therefore, when segmenting, the number of sub-samples (sub-audio) can be reduced to improve the detection efficiency and ensure the detection accuracy.

[0095] As an example, the overlap degree between every two adjacent sub-samples among multiple sub-samples is determined according to the security requirement level. The higher the security requirement level, the greater the overlap degree between two adjacent sub-samples.

[0096] In this example, the overlap degree between two sub-samples can refer to the length of the overlapping part of the two sub-samples in time. For example, for sub-sample 1 from 0 to 1 second and sub-sample 2 from 0.5 to 1.5 seconds, the length of their overlapping part is 0.5 second.

[0097] Figure 3 Shows Figure 2 A schematic diagram of the segmentation step in the data processing method in the embodiment. As Figure 3As shown, the voiceprint sample has multiple segments of equal length, and each segment can be one frame with a time length of 1 millisecond. The figure shows that the voiceprint sample has 20 frames. In scenarios with a relatively low security requirement level such as mobile phone unlocking, fewer sub-samples can be segmented, and the overlap degree between every two sub-samples can be relatively low. For example, in the segmentation method on the left side of the figure, the 20-frame voiceprint sample is segmented into 5 sub-samples, namely sub-sample x1, sub-sample x2, sub-sample x3, sub-sample x4, and sub-sample x5. The length of each sub-sample is 4 frames, and these sub-samples do not overlap with each other. In scenarios with a relatively high security requirement level such as bank transfer, more sub-samples can be segmented, and the overlap degree between every two sub-samples can be relatively high. For example, in the segmentation method shown on the right side of the figure, the 20-frame voiceprint sample is segmented into 10 sub-samples, namely sub-sample x1, sub-sample x2, sub-sample x3, sub-sample x4, sub-sample x5, sub-sample x6, sub-sample x7, sub-sample x8, sub-sample x9, and sub-sample x10. The length of each sub-sample is 3 frames or 2 frames. These sub-samples overlap with each other. For example, the overlap degree between sub-sample x1 and sub-sample x2 is 1 frame. In this way, in scenarios with a relatively high security requirement level, more sub-samples need to be detected, and some segments (overlap parts) need to be detected more than twice, so as to improve the detection accuracy and avoid missed detection due to too few detection times.

[0098] In this example, the technical solution of the present application can be deployed in common scenarios such as mobile phone unlocking, and can also be deployed in sensitive scenarios such as bank transfer. By adjusting the segmentation window, the efficiency and detection accuracy requirements of different sensitivity scenarios are met. When the technical solution of the present application is deployed in the mobile phone unlocking scenario, by increasing the segmentation window, reducing the number of overlaps between windows, reducing the computational amount required by the algorithm, and improving the detection efficiency; for more sensitive scenarios such as transfer, by reducing the segmentation window, increasing the overlap degree between windows, and increasing the number of segmented sub-samples, the detection accuracy is improved. For different deployment scenarios such as mobile phone unlocking and transfer payment, the technical solution of the present application can improve the detection efficiency or detection accuracy of the algorithm in this scenario by adjusting the size and overlap degree of the segmentation window.

[0099] As an example, the length of each sub-sample among multiple sub-samples is between 1 millisecond and half of the length of the voiceprint sample.

[0100] In this example, the length of the sub-samples is limited. Voiceprint samples have different lengths, but generally can maintain a degree that can be used for voiceprint verification and have obvious recognizable voiceprint features. For the sub-samples of voiceprint samples, since they are used to judge the authenticity of voiceprint samples, they do not need to have a length that can highlight the voiceprint features, and only a relatively short duration is required. The inventors of this application found through research and experiments that if the length of the sub-sample is less than 1 millisecond, it will cause the situation that the similarity with the standard sample cannot be successfully judged, resulting in low detection efficiency; if the length of the sub-sample is greater than half of the length of the voiceprint sample, then the result of detecting the sub-sample is basically no different from the result of detecting the entire voiceprint sample, making the segmentation of the sub-sample meaningless. Therefore, it is more appropriate to limit the length of the sub-sample between 1 millisecond and half of the length of the voiceprint sample.

[0101] As an example, multiple sub-samples do not completely cover the voiceprint sample.

[0102] In this example, the sum of the lengths of all sub-samples does not completely cover the segmented voiceprint sample. In other words, a part of the length of the voiceprint sample is not segmented into sub-samples for detection. This is because sometimes the voiceprint sample is extremely long. If the entire voiceprint sample is cut into sub-samples for detection, it will cause unnecessary detection and reduce the detection efficiency. Therefore, in some cases, only a part of the voiceprint sample can be selected for segmentation to obtain an appropriate number of sub-samples for subsequent operations.

[0103] As an example, the specific way to cut the voiceprint sample into multiple sub-samples can be: cut at the amplitude low point where the amplitude of the voiceprint sample is less than the low point threshold.

[0104] In this example, the amplitude low point is usually the part where the speaker's voice pauses. Cutting at the amplitude low point can usually cut out a complete sentence, which can be better used for detection. Although the cut sub-samples do not necessarily need to be able to completely and accurately reflect the speaker's voiceprint features like the voiceprint sample, they need to be able to effectively compare the similarity between the sub-sample and the standard sample. To successfully determine the similarity between the sub-sample and the standard sample, the sub-sample needs to have a certain length as described above. In addition to the length, the sub-sample can also have a certain degree of integrity. As described in this example, it can include a complete phrase or sentence between two pauses of the speaker, so as to be able to reflect the characteristics of the sub-sample as much as possible for better comparison of similarity.

[0105] As an example, the specific way to cut the voiceprint sample into multiple sub-samples can be: cut the voiceprint sample so that the cut sub-samples cover the amplitude high points where the amplitude of the voiceprint sample is greater than the high point threshold.

[0106] In this example, the amplitude high points are usually the places where the speaker emphasizes key points or is speaking. By cutting the sub-samples to cover the amplitude high points, the sub-samples can cover the parts that can reflect the speaker's voice characteristics and relatively complete semantics, thereby improving the detection accuracy.

[0107] S230. Calculate the similarity between multiple sub-samples and the standard samples stored in the computing device.

[0108] In this embodiment, there can be various methods for calculating the similarity between the sub-samples and the standard samples. For example, by signal processing, summarize the frequency characteristics of the audio and compare them with each other. It can also be calculated by means of an artificial intelligence model. The similarity calculation method of this application is not limited to this.

[0109] As an example, the voiceprint samples include voiceprint samples for inputting into a voiceprint recognition model for recognition. At this time, in order to calculate the similarity between each sub-sample and the standard sample, each sub-sample can be input into the voiceprint recognition model to obtain the similarity between the sub-sample and the standard sample.

[0110] In this example, the voiceprint recognition model can refer to an artificial intelligence model for recognizing the voiceprint characteristics of a piece of speech, such as the ASV model. The ASV model can refer to a method of verifying a speaker using computer technology. ASV determines whether the identity of the speaker is legal by analyzing the voice characteristics of the speaker and comparing them with the voice characteristics of a pre-registered target speaker. The principle of ASV is based on voiceprint recognition technology. A voiceprint refers to the unique voice characteristics of each person, which can be used to identify an individual just like a fingerprint. ASV extracts the voice characteristics of the speaker, establishes a voiceprint model and the voiceprint model of the target speaker for comparison, thereby determining the identity of the speaker. In ASV, feature extraction usually adopts a certain algorithm. By preprocessing the speech signal, performing Fourier transform and cepstrum analysis, the voice characteristics related to human ear auditory perception are extracted.

[0111] In this example, the voiceprint recognition model can refer to the voiceprint recognition model that the uncut voiceprint samples were originally used to be detected by. In order to improve the security of voiceprint recognition, the technical solution of this application proposes an adversarial sample detection method to determine whether the voiceprint sample is an adversarial sample or not, and then perform the detection of the voiceprint recognition model. This can improve the security of the voiceprint recognition model and avoid being deceived by adversarial samples, thus affecting the reliable stability of commercial transactions. By using the voiceprint recognition model originally used to identify the voiceprint samples to detect the sub-samples, the adversarial sample detection method proposed by the technical solution of this application no longer requires other models, reduces the computational amount, and improves the cost-effectiveness.

[0112] In this example, the target model is used to compare the sub-samples with the target speaker. In this way, the use of an additional detection model is eliminated, and the detection efficiency of adversarial samples is greatly improved. This example borrows the protected voiceprint model to calculate the similarity score, so no additional detection model is required, avoiding the computational cost brought by the detection model and improving the detection efficiency.

[0113] S240. Determine whether the voiceprint sample is an adversarial sample based on the similarity.

[0114] In this embodiment, an adversarial sample may refer to a sample used to deceive a voiceprint recognition model and not from the real voice of the speaker. The adversarial sample can be a voice sample emitted by a human or an audio sample synthesized by a machine. The adversarial sample can be formed by various means. In some embodiments, the adversarial sample is obtained by adding global noise to the voice of someone other than the speaker. Other means of obtaining adversarial samples can also be conceived, and this application is not limited thereto.

[0115] As an example, to determine whether the voiceprint sample is an adversarial sample based on the similarity, it can be determined that the voiceprint sample is an adversarial sample when the number of sub-samples with a similarity greater than the similarity threshold is greater than or equal to a preset quantity threshold, and it can be determined that the voiceprint sample is not an adversarial sample when the number of sub-samples with a similarity greater than the similarity threshold is less than the preset quantity threshold.

[0116] In this example, a sample with a similarity greater than the similarity threshold may refer to a sample determined by a machine to be from the real voice of the speaker. Since the sub-sample comes from the voiceprint sample, it only contains a fragment of the voiceprint sample, does not have complete voiceprint recognition features, and usually has noise or interference factors. Therefore, generally speaking, the probability that the similarity between a normal sub-sample and a standard sample exceeds the similarity threshold is not very high. However, in an adversarial sample, due to the addition of global noise, the similarity between any sub-sample and the standard sample is improved. Utilizing this point, it can be determined whether it is an adversarial sample by judging whether the number of sub-samples with a similarity greater than the similarity threshold is abnormal. The method for judging whether the number of these sub-samples is abnormal is to judge whether it is greater than a specific preset quantity threshold. The preset quantity threshold can be obtained through experimental methods, can also be obtained based on experience, or can be obtained through calculation or a model.

[0117] As an example, in the case where it is determined that the voiceprint sample is not an adversarial sample, voiceprint recognition is performed on the voiceprint sample to determine whether the voiceprint sample is from the real voice of the speaker.

[0118] In this embodiment, for Figure 2 the application scenario of the embodiment is described. In this scenario, Figure 2The data processing method described in the embodiment is used to determine whether a voiceprint sample is an adversarial sample before identifying the voiceprint sample, so as to prevent the adversarial sample from deceiving the voiceprint recognition model in the subsequent voiceprint sample recognition process and misjudging the adversarial sample that does not come from the real voice of the speaker as a real sample.

[0119] In this embodiment, adversarial sample detection is a pre - procedure for voiceprint sample detection. Those skilled in the art should understand that adversarial sample detection can also be an independent procedure and does not depend on the subsequent voiceprint sample detection procedure. In some embodiments, when it is necessary to determine the similarity between a sub - sample and a standard sample during the adversarial sample detection process, the detection method used in the voiceprint sample detection process can be adopted. That is, the voiceprint recognition model used in the voiceprint sample detection process can be used to determine the similarity between the sub - sample and the standard sample. Of course, other means can also be used to determine the similarity between the sub - sample and the standard sample, and this application is not limited thereto.

[0120] Figure 4 The flowchart showing the data processing method according to an embodiment of the present application is presented.

[0121] According to this embodiment, the data processing method includes steps S410 to S490, and each step is described in detail below.

[0122] S410. Obtain a voiceprint sample.

[0123] S420. Split the voiceprint sample into multiple sub - samples.

[0124] For the details of S410 and S420, refer to the detailed description of S210 and S220 in the above - mentioned Figure 2 embodiment, which will not be elaborated here.

[0125] S430. Identify the features of the sub - sample, where the features include one or more of amplitude, frequency, time length, language, and accent.

[0126] In this embodiment, the features of the sub - sample can refer to the recognizable traits or voice characteristics that can reflect the speaker's voice. Identifying the features of the sub - sample can be carried out through audio detection technology or artificial intelligence technology, etc. The amplitude feature can refer to the loudness of the speaker's voice, the frequency feature can refer to the sharpness or lowness of the speaker's voice, the time - length feature can refer to the time length when the sub - sample is played or recorded normally, the language feature can refer to the language spoken by the speaker such as Chinese, English, etc., and the accent feature can refer to the local or habitual accent in the speaker's language such as dialect. The sub - sample can also have other features.

[0127] S440. Determine the standard sample according to the features.

[0128] In this embodiment, by identifying the features of the sub-samples, the most suitable sample can be selected from multiple standard samples for similarity comparison. The multiple standard samples may refer to multiple standard samples that have been recorded by the speaker and stored in the database, which are used to select appropriate standard samples for identification and verification when verifying the speaker's identity. When there are multiple standard samples, in order to determine the similarity between the sub-samples and the standard samples, the standard sample closest to the features of the sub-samples can be selected, so as to determine as much as possible that the sub-samples and the standard samples are similar, and thus more strictly and accurately determine the adversarial samples.

[0129] S450. Calculate the similarity between multiple sub-samples and the standard samples stored in the computing device.

[0130] For the details of S450, refer to the detailed description of S230 in the above embodiment, which will not be elaborated here. Figure 2

[0131] S460. For multiple real voiceprint samples, each real voiceprint sample is segmented into a first number of real sub-samples, and calculate the second number of real sub-samples of the real voiceprint samples whose similarity with the standard samples is greater than the similarity threshold, where the first number is variable for different real voiceprint samples.

[0132] S470. For multiple adversarial voiceprint samples, each adversarial voiceprint sample is segmented into a third number of adversarial sub-samples, and calculate the fourth number of adversarial sub-samples of the adversarial voiceprint samples whose similarity with the standard samples is greater than the similarity threshold, where the third number is variable for different adversarial voiceprint samples.

[0133] S480. Determine a quantity threshold according to the second number and the fourth number.

[0134] As an example, in order to determine the quantity threshold according to the second number and the fourth number, for multiple real voiceprint samples, calculate the maximum value of the ratio of the second number to the first number; then, for multiple false voiceprint samples, calculate the minimum value of the ratio of the fourth number to the third number; then, according to the smaller value of the maximum value and the minimum value, determine the ratio of the quantity threshold to the number of multiple sub-samples; finally, determine the quantity threshold according to the ratio.

[0135] In this embodiment, experiments can be conducted on multiple known real voiceprint samples and fake voiceprint samples to obtain the optimal quantity threshold through testing. Specifically, multiple known real voiceprint samples can be segmented to obtain real sub-samples with different quantities, that is, the number of real sub-samples segmented from each real voiceprint sample is different. The similarity between these real sub-samples and the standard sample is judged to obtain the number of real sub-samples whose similarity exceeds the similarity threshold. Thus, for each real voiceprint sample, a ratio between the number of real sub-samples whose similarity exceeds the similarity threshold (i.e., the second quantity) and the total number of real sub-samples (i.e., the first quantity) can be obtained. Since it is known that these samples are all real samples and all come from the real voice of the speaker, the maximum value of this ratio among different real voiceprint samples should be the lower limit of the ratio used to judge adversarial samples (hereinafter referred to as lower limit 1). On the other hand, multiple known fake voiceprint samples can be segmented to obtain fake sub-samples with different quantities, that is, the number of fake sub-samples segmented from each fake voiceprint sample is different. The similarity between these adversarial samples and the standard sample is judged to obtain the number of fake sub-samples whose similarity exceeds the similarity threshold. Thus, for each fake sub-sample, a ratio between the number of fake sub-samples whose similarity exceeds the similarity threshold (i.e., the fourth quantity) and the total number of fake sub-samples (i.e., the third quantity) can be obtained. Since it is known that these samples are all adversarial samples and do not come from the real voice of the speaker, the minimum value of this ratio among different fake voiceprint samples should be the lower limit of the ratio for judging adversarial samples (hereinafter referred to as lower limit 2). By comparing the magnitudes of lower limit 1 and lower limit 2, it can be obtained what the lowest ratio should be for the number of sub-samples whose similarity is greater than the similarity threshold to account for the total number of sub-samples, that is, the smaller value between lower limit 1 and lower limit 2. After obtaining the lowest lower limit value, the quantity threshold can be determined according to the number of sub-samples during actual detection.

[0136] In this embodiment, by selecting a small number of samples and using existing adversarial sample generation algorithms to generate a batch of adversarial samples, the ratio of normal samples to adversarial samples can be obtained, and the optimal ratio threshold can be determined as the optimal detection threshold. If there are a part of known normal samples and a part of known adversarial samples, all samples can be cut to obtain corresponding sub-samples, and the threshold that can achieve the best detection effect on the known sub-samples is selected from a series of detection thresholds as the optimal detection threshold.

[0137] S490. When the number of sub-samples whose similarity is greater than the similarity threshold is greater than or equal to the preset quantity threshold, it is determined that the voiceprint sample is an adversarial sample.

[0138] S401. When the number of sub-samples whose similarity is greater than the similarity threshold is less than the preset quantity threshold, it is determined that the voiceprint sample is not an adversarial sample.

[0139] For details about S480 and S490, see the detailed description of S240 in the foregoing Figure 2 embodiment, which will not be elaborated here.

[0140] Based on the foregoing Figure 2 method embodiment, an embodiment of the present application further provides a data processing device, and its structural schematic diagram is as Figure 5 shown. The data processing device 500 is used to execute each of the foregoing Figure 2 steps.

[0141] According to this embodiment, the data processing device 500 is applied to a computing device. The device includes an acquisition module 510, a segmentation module 520, a calculation module 530, and a judgment module 540. The acquisition module 510 is used to acquire a voiceprint sample. The segmentation module 520 is used to segment the voiceprint sample into multiple sub-samples. The calculation module 530 is used to calculate the similarity between the multiple sub-samples and a standard sample stored in the computing device. The judgment module 540 is used to judge whether the voiceprint sample is an adversarial sample based on the similarity.

[0142] It should be noted that Figure 5 when the data processing device 500 provided in the shown embodiment executes the method, only the division of the above functional modules is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the data processing device 500 provided in the above embodiment and Figure 2 the data processing method embodiment shown respectively belong to the same concept. For the specific implementation process, see the method embodiment, which will not be elaborated here.

[0143] Figure 6 is a hardware structural schematic diagram of a computing device 600 provided in an embodiment of the present application.

[0144] See Figure 6 , the computing device 600 includes a processor 610, a memory 620, a communication interface 630, and a bus 640. The processor 610, the memory 620, and the communication interface 630 are connected to each other through the bus 640. The processor 610, the memory 620, and the communication interface 630 can also be connected in other connection ways except the bus 640.

[0145] Among them, the memory 620 can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical memory, hard disk, etc.

[0146] Among them, the processor 610 can be a general-purpose processor, and the general-purpose processor can be a processor that executes specific steps and / or operations by reading and executing the content stored in a memory (such as the memory 620). For example, the general-purpose processor can be a central processing unit (CPU). The processor 610 can include at least one circuit to execute Figure 2 all or part of the steps of the data processing method provided by the illustrated embodiment.

[0147] Among them, the communication interface 630 includes interfaces such as input / output (I / O) interfaces, physical interfaces, and logical interfaces for implementing interconnection of components inside the computing device 600, as well as interfaces for implementing interconnection between the computing device 600 and other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc. The communication interface 630 can be externally connected to an input device and an output device. For example, the input device can be a microphone or a microphone array for capturing voice input signals; it can be a communication network connector for receiving the collected input signals from the cloud or other devices; it can also include, for example, a keyboard, a mouse, etc. The output device can output various information to the outside, including the determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0148] Among them, the bus 640 can be of any type and is a communication bus for implementing interconnection of the processor 610, the memory 620, and the communication interface 630, such as a system bus.

[0149] The above-mentioned components can be respectively provided on independent chips, or at least partially or entirely provided on the same chip. Whether to independently provide each component on different chips or integrate them on one or more chips often depends on the needs of product design. The embodiments of the present application do not limit the specific implementation forms of the above-mentioned components.

[0150] Figure 6 The computing device 600 shown is merely exemplary. During implementation, the computing device 600 may further include other components, which will not be listed one by one herein.

[0151] An embodiment of the present application may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps in the data processing methods according to various embodiments of the present application described above in this specification.

[0152] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0153] The concepts, principles, and ideas of the present application have been described in detail above in conjunction with specific implementation manners (including embodiments and examples). Those skilled in the art should understand that the implementation manners of the present application are not limited to the several forms given above. After reading the present application document, those skilled in the art may make any possible improvements, substitutions, and equivalent forms to the steps, methods, devices, and components in the above implementation manners. These improvements, substitutions, and equivalent forms should be regarded as falling within the scope of the present application. The protection scope of the present application is subject to the claims only.

Claims

1. A data processing method, characterized in that, applied to a computing device, the method includes: obtaining a voiceprint sample; dividing the voiceprint sample into a plurality of sub-samples; calculating the similarity between the plurality of sub-samples and a standard sample stored in the computing device; judging whether the voiceprint sample is an adversarial sample based on the similarity.

2. The method according to claim 1, characterized in that, the judging whether the voiceprint sample is an adversarial sample based on the similarity includes: when the number of sub-samples with a similarity greater than a similarity threshold is greater than or equal to a preset number threshold, determining that the voiceprint sample is an adversarial sample; when the number of sub-samples with a similarity greater than a similarity threshold is less than a preset number threshold, determining that the voiceprint sample is not an adversarial sample.

3. The method according to claim 1, characterized in that, the number of the plurality of sub-samples is determined according to the security requirement level, and when the security requirement level is higher, the number of the sub-samples is more.

4. The method according to claim 1, characterized in that, the overlap degree between every two adjacent sub-samples among the plurality of sub-samples is determined according to the security requirement level, and when the security requirement level is higher, the overlap degree between the two adjacent sub-samples is greater.

5. The method according to claim 1, characterized in that, the length of each sub-sample among the plurality of sub-samples is between 1 millisecond and half of the length of the voiceprint sample.

6. The method according to claim 1, characterized in that, the plurality of sub-samples do not completely cover the voiceprint sample.

7. The method according to claim 1, characterized in that, the dividing the voiceprint sample into a plurality of sub-samples includes: cutting at an amplitude low point where the amplitude of the voiceprint sample is less than a low point threshold; or cutting the voiceprint sample so that the sub-samples obtained by cutting cover the amplitude high points where the amplitude of the voiceprint sample is greater than a high point threshold.

8. The method according to claim 2, characterized in that, before determining that the voiceprint sample is an adversarial sample when the number of sub-samples with a similarity greater than a similarity threshold is greater than or equal to a preset number threshold, it further includes: for a plurality of real voiceprint samples, dividing each real voiceprint sample into a first number of real sub-samples, and calculating a second number of real sub-samples of the real voiceprint sample with a similarity to the standard sample greater than the similarity threshold, wherein the first number is variable for different real voiceprint samples; for a plurality of adversarial voiceprint samples, dividing each adversarial voiceprint sample into a third number of adversarial sub-samples, and calculating a fourth number of adversarial sub-samples of the adversarial voiceprint sample with a similarity to the standard sample greater than the similarity threshold, wherein the third number is variable for different adversarial voiceprint samples; determining the preset number threshold according to the second number and the fourth number.

9. The method according to claim 8, characterized in that, the determining the preset number threshold according to the second number and the fourth number includes: For the multiple real voiceprint samples, calculate the maximum value of the ratio of the second quantity to the first quantity; For the multiple adversarial voiceprint samples, calculate the minimum value of the ratio of the fourth quantity to the third quantity; According to the smaller of the maximum value and the minimum value, determine the ratio of the preset quantity threshold to the quantity of the multiple sub-samples; According to the ratio, determine the preset quantity threshold.

10. The method according to claim 1, wherein, before calculating the similarity between the multiple sub-samples and the standard samples stored in the computing device, further comprising: identifying the features of the sub-samples, the features including one or more of amplitude, frequency, time length, language, accent; determining the standard samples according to the features.

11. The method according to claim 1, wherein, the voiceprint samples include voiceprint samples for inputting into a voiceprint recognition model for recognition, and calculating the similarity between the multiple sub-samples and the standard samples stored in the computing device includes: inputting each sub-sample into the voiceprint recognition model to obtain the similarity between the sub-sample and the standard sample.

12. The method according to any one of claims 1-11, wherein, in the case of determining that the voiceprint sample is not an adversarial sample, perform voiceprint recognition on the voiceprint sample to determine whether the voiceprint sample is from the real voice of the speaker.

13. A data processing device, wherein, applied to a computing device, the device includes: an acquisition module, configured to acquire voiceprint samples; a segmentation module, configured to segment the voiceprint samples into multiple sub-samples; a calculation module, configured to calculate the similarity between the multiple sub-samples and the standard samples stored in the computing device; a judgment module, configured to judge whether the voiceprint sample is an adversarial sample based on the similarity.

14. A computing device, wherein, the computing device includes a processor and a memory, and the processor is configured to execute a computer program stored in the memory to implement the data processing method according to any one of claims 1 to 12.

15. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and the computer program is configured to execute the data processing method according to any one of claims 1 to 12.