Method and device for identifying authenticity of sound in audio information

By encoding the audio information and comparing it with the prototype vector, using similarity to identify the authenticity of the sound in the audio information, the problem of difficulty in identifying the authenticity of the sound in the prior art is solved, and the accuracy of the recognition is improved.

CN120148554APending Publication Date: 2025-06-13ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510494504.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively identify the authenticity of sounds in audio information, which poses security risks, which may lead to the spread of false information, privacy leakage and damaging social trust.

Method used

By encoding the audio information to be identified, an encoded vector is obtained and compared with the prototype vector to determine the authenticity of the sound using the similarity. Prototype vectors are used to describe hidden categories within the hidden space of real and fake sounds.

Benefits of technology

Improve the accuracy of authenticity and false recognition of sounds in audio information, and can implicitly identify one or more data centers for a single classification category, enhancing the ability to identify fake sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148554A_ABST
    Figure CN120148554A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and a device for identifying the authenticity of sound in audio information, which can refine classification categories under the condition of identifying that the sound in the audio information is real sound or forged sound (such as synthetic sound), and specifically, the method and the device can be used for identifying the authenticity of the sound in the audio information. A real sound or counterfeit sound classification category is each refined into at least one hidden category within a hidden space, a single hidden category being characterized by a single prototype vector. According to the method, after the audio information to be recognized is coded to obtain the corresponding coding vector, the coding vector can be compared with each prototype vector to obtain each corresponding similarity, and then the sound authenticity of the audio information to be recognized is determined according to each similarity, namely, the audio information belongs to a real sound classification category or a forged sound classification category. Therefore, the accuracy of sound authenticity identification in the audio information can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a method and device for identifying the authenticity of voices in audio information. Background Art

[0002] Generative Artificial Intelligence (AIGC) is an artificial intelligence technology that can autonomously learn from data and generate new content including images, sounds, texts, etc. Taking synthetic speech as an example, text can be converted into audible voice audio. During the process of synthesizing speech, the computer will convert the text into natural speech through a speech synthesis engine according to the input text content. Potential risks that forged audio information may bring include, for example, the spread of false information, privacy leakage, and the destruction of social trust. Therefore, there are security risks in speech synthesis or conversion technology in the AIGC era. Taking corresponding measures to avoid these risks and prevent the abuse of this technology is beneficial to protecting personal privacy and social security. Summary of the Invention

[0003] One or more embodiments of this specification describe a method and device for identifying the authenticity of voices in audio information to solve one or more problems mentioned in the background art.

[0004] According to a first aspect, there is provided a method for identifying the authenticity of voices in audio information, the method comprising: encoding a first audio information to be recognized to obtain a first encoded vector; detecting respective similarities between the first encoded vector and respective prototype vectors, where a single prototype vector is used to describe a single hidden category within a hidden space of two classification categories of real voices or forged voices, and at least one hidden category corresponds to a single classification category; and determining the authenticity of the voice in the first audio information using the respective similarities.

[0005] In one embodiment, the encoding the first audio information to be recognized to obtain a first encoded vector comprises: extracting voice features from the first audio information, the voice features including at least one of the following: loudness, pitch, frequency, timbre, musical tone, duration, harmonic structure; and encoding or embedding the voice features to obtain the first encoded vector.

[0006] In a further embodiment, the encoding or embedding the voice features to obtain the first encoded vector comprises: before or after encoding or embedding the voice features, respectively normalizing the data on each feature dimension, where the normalization operation is one of the following operations: rescaling, mean normalization, standardization, unit length normalization.

[0007] In one embodiment, each prototype vector is adjusted in the training phase in the following manner: in a single parameter adjustment cycle: process the sample audio information of the current batch to obtain the predicted classification category; compare the predicted classification category with the corresponding classification label to obtain the current model loss; update each undetermined parameter according to the model loss to adjust each prototype vector; where the undetermined parameter includes one of the following: the values of each dimension in the prototype vector; the model parameters in the embedding network that embeds the one-hot representations of each hidden category to obtain the prototype vector.

[0008] In one embodiment, determining the authenticity of the sound in the first audio information by using each similarity includes: detecting the first classification category to which the prototype vector corresponding to the maximum similarity belongs, where the first classification category is a real sound classification category or a forged sound classification category; determining the authenticity of the sound in the first audio information according to the first classification category.

[0009] In one embodiment, determining the authenticity of the sound in the first audio information by using each similarity includes: using each similarity to determine the first fusion similarity and the second fusion similarity corresponding to the two classification categories of real sound and forged sound respectively, where the first fusion similarity or the second fusion similarity is obtained by performing a fusion operation on each similarity corresponding to each prototype vector corresponding to the corresponding classification category; determining the authenticity of the sound in the first audio information according to the larger value of the first fusion similarity and the second fusion similarity.

[0010] In a further embodiment, when the number of prototype vectors corresponding to the real sound classification category and the forged sound classification category is the same, the fusion operation includes at least one of: weighted summation, summation, mean value calculation, maximum value taking; when the number of prototype vectors corresponding to the real sound classification category and the forged sound classification category is different, the fusion operation includes at least one of: weighted summation, mean value calculation, maximum value taking, where the weight value in the weighted summation operation is negatively correlated with the number of prototype vectors corresponding to the corresponding classification category.

[0011] According to a second aspect, there is provided an apparatus for identifying the authenticity of a sound in audio information, the apparatus including:

[0012] An encoding unit configured to encode the first audio information to be identified to obtain a first encoded vector;

[0013] A detection unit configured to detect each similarity corresponding to the first encoded vector and each prototype vector, where a single prototype vector is used to describe a single hidden category in the hidden space corresponding to the two classification categories of real sound or forged sound, and at least one hidden category corresponds to a single classification category;

[0014] An identification unit configured to determine the authenticity of the voice in the first audio information by using respective similarity degrees.

[0015] According to a third aspect, there is provided a computer-readable storage medium having stored thereon a computer program which, when executed on a computer, causes the computer to execute the method of the first aspect.

[0016] According to a fourth aspect, there is provided a computing device including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method of the first aspect is implemented.

[0017] Through the apparatus and method provided in the embodiments of this specification, during the process of identifying whether the voice in the audio information is a real voice or a forged voice (such as a synthesized voice, etc.), the classification categories can be refined. Specifically, the classification categories of real voices or forged voices are each refined into hidden categories in the latent space, and a single hidden category is represented by a single prototype vector. After encoding the audio information to be identified to obtain a corresponding encoded vector, the encoded vector can be compared with each prototype vector respectively to obtain respective similarity degrees, and then, based on the respective similarity degrees, the authenticity of the voice in the audio information to be identified is determined, that is, whether it belongs to the real voice classification category or the forged voice classification category. In this way, one or more data centers can be implicitly determined for a single classification category, improving the accuracy of identifying the authenticity of voices in audio information. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0019] Figure 1 Shows a schematic diagram of a specific implementation architecture for identifying the authenticity of voices in audio information in the conventional technology;

[0020] Figure 2 Shows a schematic diagram of a specific implementation architecture for identifying the authenticity of voices in audio information according to the embodiments of this specification;

[0021] Figure 3 Shows a schematic diagram for comparing the effect of identifying the authenticity of voices in audio information in this specification with the conventional technology;

[0022] Figure 4 Shows a schematic flowchart for identifying the authenticity of voices in audio information according to an embodiment of this specification;

[0023] Figure 5The structural block diagram of a device for identifying the authenticity of voices in audio information according to an embodiment of this specification is shown. Detailed implementation manners

[0024] The solution provided in this specification will be described below with reference to the accompanying drawings.

[0025] Similar to text, audio, as a form of information expression, can convey relevant information through voices. Audio information can be understood as voice information. In this specification, in the audio information generated by intelligent devices such as computers, the voices are usually synthesized according to relevant voice characteristics, which can be understood as the voices in the audio information being forged voices, while the voices collected by means such as recording can be understood as real voices. Here, the real voices and forged voices are usually human voices, but the voices of other things (such as animal calls, etc.) are not excluded.

[0026] In conventional technologies, for the audio information to be detected, features are usually extracted, the features are encoded, and then through a binary classification model, it is mapped into two classifications: real voice and forged voice. As Figure 1 shown, an implementation solution of conventional technologies in the process of identifying the authenticity of voices in the audio information involved in this specification is shown. As Figure 1 shown, the identification of the authenticity of voices in audio information can be performed by a computing platform. This computing platform can be implemented by hardware (such as computer modules, etc.) or software (such as applications deployed on an operating system). Referring to Figure 1 shown, the computing platform can receive the audio information whose authenticity is to be identified, perform feature extraction on the audio information (described by a dotted box, indicating that it may be omitted in some embodiments), encode the extracted features, and then determine the authenticity of the voice according to the encoding result and output the determination result.

[0027] Considering the complexity of the real situation, traditional binary classification models may be limited by sample sources, sample quality, etc. In view of this, this specification can describe the hidden space of classification categories through prototype vectors, which is equivalent to expanding the number of classifications, so as to mine the differences in data distribution between forged voices and real voices from multiple dimensions and improve the generalization ability of identifying the authenticity of voices.

[0028] Figure 2 The schematic diagram of a specific implementation architecture of this specification is shown. Referring to Figure 2 shown, compared with the conventional technology example shown in Figure 1 in the technical concept of this specification, after performing feature extraction (optional) and encoding processing on the audio information to be identified, the obtained encoding vectors are mapped to each prototype vector, and each prototype vector can be regarded as a hidden category in the hidden space. Furthermore, the mapping results corresponding to the encoding vectors in the prototype vectors are fused, so as to determine the authenticity of the voice.

[0029] To make the technical concept of this specification clearer, refer to Figure 3 As shown, above the dotted line is equivalent to Figure 1 the technical solution shown, and below the dotted line corresponds to the technical concept of this specification. Among them, the audio information represented by the small circle ○, and the triangle ▲ symbol represents the classification space (or classification center). As Figure 3 shown, under the binary classification model, the audio information can be divided into two categories. The left side indicates that the sound is a real sound, and the right side indicates that the sound is a forged sound, corresponding to two triangle ▲ symbols respectively. At this time, the more marginal audio information in the circles 301 and 302 is more likely to be classified into the category of real sound, resulting in misclassification. Under the technical concept of this specification, the audio information is mapped to multiple prototype spaces through multiple prototype vectors, represented by multiple ▲ symbols. Considering that the characteristics of real sounds are more unified and the forged sounds in synthetic audio have their own characteristics, assuming that one ▲ on the left represents that the sound is true (real sound) and multiple ▲ on the right represent that the sound is false (forged sound), then the sound in the audio information close to any ▲ representing that the sound is false can be identified as a false sound. In this way, more accurate and effective discrimination results can be obtained. It should be noted that Figure 3 This is only an example. In practice, there can be multiple prototype vectors of real classification in the latent space, and the number of prototype vectors of forged sound classification in the latent space can also be other numbers, which are not limited here.

[0030] Figure 4 shows the flowchart of the method for identifying the authenticity of sound in audio information proposed in the embodiment of this specification. The execution subject of this process can be any computer, device, or server with certain computing capabilities. Refer to Figure 4 As shown, the process of identifying the authenticity of sound in audio information can include the following steps: Step 401, encode the first audio information to be recognized to obtain a first encoded vector; Step 402, detect the similarities between the first encoded vector and each prototype vector respectively. A single prototype vector is used to describe a single hidden category in the latent space corresponding to two classification categories of real sound or forged sound, and at least one hidden category corresponds to a single forged sound classification category; Step 403, determine the authenticity of the sound in the first audio information using each similarity.

[0031] First, in step 401, the first audio information to be recognized is encoded to obtain a first encoded vector.

[0032] Here, the first audio information can be any piece of audio information whose authenticity is to be recognized. The first encoded vector is a vector obtained by embedding or encoding the first audio information.

[0033] In one embodiment, encoding the first audio information may be directly encoding it using an encoding network, such as processing it using a multimodal large model, etc., to obtain a first encoded vector.

[0034] In another embodiment, the encoding of the first audio information may be extracting sound features from the first audio information and then encoding or embedding the extracted sound features to obtain a first encoded vector. The sound features may include, for example, but are not limited to, one of the following: loudness, pitch, frequency, timbre, musical tone, duration, harmonic structure, etc. Among them, loudness describes the size of the sound and is determined by the amplitude, pitch describes the high or low (high pitch, low pitch) of the sound and is determined by the "frequency", frequency is the number of sound waves passing through a given point per second, timbre is also known as the tone quality and is determined by the waveform, musical tone is a regular and pleasant sound, noise can be the sound outside the main sound source of the audio (such as the sound emitted by other sound sources) of a person, etc., and the harmonic structure can be the Fourier series decomposition of a periodic alternating quantity to obtain components with frequencies that are integer multiples greater than 1 of the fundamental frequency.

[0035] The sound features can be extracted through an audio feature extraction network. The audio feature extraction network is a network that can extract sound features from the waveform of the audio. The audio feature network can be various pre-trained sound feature extraction networks, which can use conventional sound feature extractors, such as the wav2vec[xlsr2_300m] model, the WavLM pre-trained model, or a feature extraction network determined by at least one neural network such as a fully connected layer, a convolutional neural network, etc., which is not limited here. Among them, both the WavLM pre-trained model and the wav2vec[xlsr2_300m] model are self-supervised models based on Transformer and can learn useful feature representations from the original audio waveform.

[0036] In some alternative embodiments, before extracting the sound features from the audio information, the audio information can also be preprocessed and the preprocessed audio information is used for sound feature extraction. The preprocessing can, for example, make the length of the audio meet a predetermined condition, such as a predetermined duration or the duration falls within a predetermined range. In the case where the audio information is too short, it can be copied to reach the length that meets the predetermined condition, and in the case where the audio information is too long, it can be truncated to meet the predetermined condition. In other examples, the preprocessing can also be noise filtering, etc., which will not be elaborated here.

[0037] It can be understood that encoding the voice features is a process of further processing the voice features, performing feature fusion or extracting higher-order features. Encoding the voice features can be carried out using an encoding network. The encoding network (or called the embedding network) can be composed of at least one of a fully connected network, an attention network, a convolutional neural network, etc., which will not be elaborated here. The encoding vector can be, for example, m-dimensional.

[0038] In an optional implementation, before or after encoding or embedding the voice features, normalization (feature scaling) processing can also be performed on each encoding vector. This is because different features may have their own measurement criteria. For example, the loudness of sound is usually described in decibels (dB), usually ranging from 0 to 130 dB, and audio is usually described in vibration hertz (Hz), such as 20 to 2000 Hz, and so on. The numerical ranges of different features may vary greatly. Then, in the process of feature data processing, features with larger numerical values may play a decisive role, while features with smaller numerical values are ignored. Therefore, before encoding the voice features or after encoding (such as embedding or fusion processing), the corresponding voice features or encoding vectors can be normalized.

[0039] The methods of normalization usually include: rescaling (min-max normalization, range scaling), mean normalization, standardization (Z-score normalization), scaling to unit length, etc. Among them, according to specific business requirements, various corresponding normalization methods can be used. Here, taking the rescaling normalization method as an example, a single-dimensional feature can be linearly mapped to the target range [a, b], then the scaled length is b - a, that is, the minimum value is mapped to a and the maximum value is mapped to b. In this way, according to the proportion position of the current feature value in the feature value range, its mapped position in the target range [a, b] can be determined, such as denoted as where min(x) is the minimum value of the current-dimensional feature x, and max(x) is the maximum value of the feature x. Assuming the target range is [0, 1], then a = 0, b = 1, so there is:

[0040] According to the actual business, the normalization operation on the encoding vector can be one of rescaling, mean normalization, standardization, and scaling to unit length, which will not be elaborated here. It is worth noting that the mean used in mean normalization can be determined by the mean of multiple samples or multiple audio data in the corresponding dimension.

[0041] Next, through step 402, the similarities between the first encoded vector and the respective prototype vectors are detected.

[0042] The prototype vectors can be the vectors of the respective dimensions that make up the latent space. Here, the latent space can be understood as a multi-dimensional hidden space. A single prototype vector is used to describe the center of a more refined hidden category within the latent space of one of the two classification categories, namely the real sound category or the forged sound category. The dimension of the prototype vector can generally be the same as that of the encoded vector.

[0043] The prototype vectors can be divided into prototype vectors corresponding to the real sound category and prototype vectors corresponding to the forged sound category. At least one prototype vector can correspond to a single classification category. The number of prototype vectors can be preset, and generally can include at least one prototype vector corresponding to the real sound and multiple prototype vectors corresponding to the forged sound.

[0044] In one embodiment, considering that the features of the real sound are relatively unified, the number of prototype vectors corresponding to the real sound is small, such as n1, and the number of prototype vectors corresponding to the forged sound is n2. Both n2 and n1 are natural numbers, and n2 > n1. Optionally, n1 can be set to 1.

[0045] In another embodiment, the number of prototype vectors corresponding to both the real sound and the forged sound is n. n is a natural number greater than 1. In other embodiments, the number of other prototype vectors can also be set according to experience, which is not limited here.

[0046] The similarity between a single prototype vector and the first encoded vector can be measured by one of the following: cosine similarity, Manhattan distance, Euclidean distance, Chebyshev distance, Pearson correlation coefficient, Jaccard coefficient, Hamming distance, variance, cross-entropy, vector angle, etc. Specifically, the vector similarity can be negatively correlated with at least one of the Manhattan distance, Euclidean distance, Chebyshev distance, Hamming distance, variance, vector angle, etc., or positively correlated with at least one of the cosine similarity, Pearson correlation coefficient, Jaccard coefficient, cross-entropy, etc. This is because the higher the similarity between a single prototype vector and the first encoded vector, the closer the first encoded vector approaches the single prototype vector. Considering the vector as the coordinates of a spatial point, the distance between the corresponding spatial points of the two is closer.

[0047] In an alternative implementation, before calculating the similarity between the encoded vector and each prototype vector, each prototype vector can be separately normalized (feature scaling). The normalization process for a single prototype vector is, for example, at least one of rescaling, standardization, mean normalization, unit length normalization, etc., which will not be elaborated here. It can be understood that the normalization of the prototype vector generally keeps the data distribution within the vector unchanged.

[0048] Further, in step 403, the authenticity of the sound in the first audio information is determined using each similarity.

[0049] It can be understood that regarding the prototype vector as the central coordinate of the hidden category, the prototype vector can be used to represent the center of a more refined hidden category (such as the triangle in Figure 3 ). In this way, a single similarity is equivalent to representing the distance between the encoded feature of the audio information to be recognized and the center, and the higher the similarity, the smaller the distance. Conversely, the lower the similarity, the larger the distance. Therefore, the authenticity classification of the sound in the audio information to be recognized can be determined based on each similarity, that is, classified into the classification category of real sound or the classification category of forged sound.

[0050] In some alternative implementations, the first classification category to which the prototype vector corresponding to the maximum similarity belongs can be detected, and the authenticity of the sound in the first audio information is determined based on the first classification category. Among them, the first classification category is one of the real sound classification category or the forged sound classification category. It can be understood that the prototype vector corresponding to the maximum similarity can be regarded as the center of the hidden category closest to the first audio information, and the classification category of real sound or forged sound corresponding to this hidden category center is the classification category to which the first audio information can be classified.

[0051] As a specific example, assume that the prototype vectors corresponding to real sounds are marked by numbers 1 to n1 (from 1 to n1), and the prototype vectors corresponding to forged sounds are marked by numbers n1 to n2 (from n1 to n2). If the number of the prototype vector corresponding to the maximum similarity falls within the range of 1 to n1, then the first audio information can be recognized as a real sound; if the number of the prototype vector corresponding to the maximum similarity falls within the range of n1 to n2, then the first audio information can be recognized as a forged sound.

[0052] In some other alternative implementations, the first fusion similarity and the second fusion similarity corresponding to the two classification categories of the real voice and the forged voice can be determined by using each similarity, and the authenticity of the voice in the first audio information can be determined according to the magnitudes of the first fusion similarity and the second fusion similarity. Among them, the first fusion similarity or the second fusion similarity can be obtained by performing a fusion operation on the similarities of the respective prototype vectors corresponding to the corresponding classification categories. In this way, it can be further clarified which classification category's prototype vector (or hidden classification) the first audio information tends to approach more in the latent space.

[0053] When the number of prototype vectors corresponding to the two classification categories of the real voice and the forged voice is equal, the fusion operation of each similarity can be, for example, one of summation, weighted summation, mean value calculation, maximum value extraction, etc. Among them, in the weighted summation process, the weights can be preset or determined by using the number of prototypes, etc., such as the reciprocal of the number of prototypes. Since the number of prototype vectors (i.e., the number of similarities) corresponding to the two classification categories is the same, summation is also equivalent to weighted summation with a weight of 1 for each similarity.

[0054] When the number of prototype vectors corresponding to the real voice classification category and the forged voice classification category is inconsistent, the fusion operation of each similarity can include at least one of weighted summation, mean value calculation, and maximum value extraction. Here, the weights in the weighted summation operation can be negatively correlated with the number of prototype vectors corresponding to the corresponding classification category. At this time, since the number of similarities corresponding to the two classification categories is different, directly summing the similarities may lead to unreasonable results. Therefore, the summation method is usually not used to determine the first fusion similarity and the second fusion similarity. In the weighted summation method, this imbalance in the number of prototype vectors can be balanced by the weights. For example, the weights can be negatively correlated with the number of prototype vectors corresponding to the corresponding classification category. As an example, assume that the number of prototype vectors corresponding to the real voice classification category is n1, and the number of prototype vectors corresponding to the forged voice classification category is n2. Then, the weights of the similarities of the respective prototype vectors used to determine the first fusion similarity can be positively correlated with 1 / n1, n2 / n1, n2 / (n1 + n2), etc. (these values are all negatively correlated with n1), and the weights of the similarities of the respective prototype vectors used to determine the second fusion similarity can be positively correlated with 1 / n2, n1 / n2, n1 / (n1 + n2), etc. (these values are all negatively correlated with n2).

[0055] It should be noted that in order to enable the prototype vectors to better describe the hidden categories in the latent space, the prototype vectors can be adjusted during the model training process, so as to effectively divide the real voice space and the forged voice space in the latent space.

[0056] Generally, the prototype vectors can be adjusted based on the sample audio information and the classification label indicating whether it is a real voice or a forged voice. For example, in a single parameter adjustment cycle, the Figure 4 process shown can be used to process the current batch of sample audio information to obtain the predicted classification category, and the predicted classification category is compared with the corresponding classification label to obtain the current model loss. Based on the model loss, the update gradients of each undetermined parameter can be determined, and then the undetermined parameters can be updated using a gradient update method such as gradient descent to adjust each prototype vector. In practice, the model loss can also be determined according to other reasonable methods to adjust each prototype vector, which is not limited here.

[0057] The adjustment process of the prototype vectors is distinguished according to the settings of the prototype vectors.

[0058] In one embodiment, the values of each dimension in the prototype vector can be used as undetermined parameters. At this time, each dimension in the prototype vector can be initialized to a random value, and the values of each dimension can be directly adjusted as undetermined parameters in subsequent parameter cycles.

[0059] In another embodiment, each prototype vector can be obtained by embedding an initial feature vector using an embedding network. At this time, the undetermined parameters can be the model parameters in the embedding network. The initial feature tensor can be set manually or determined by various encoding methods. For example, the initial feature vector is a one-hot representation. A single initial feature vector has one dimension of 1 and other dimensions of 0, and the one-hot representations corresponding to each prototype vector are different. At this time, since the representation of the initial feature vector is too single and not conducive to describing the data distribution, the corresponding prototype vector can be obtained by embedding the one-hot representation describing the hidden category using an embedding network. The model parameters in the embedding network can be used as undetermined parameters, so that the prototype vector can be updated by adjusting the model parameters in the embedding network.

[0060] In other embodiments, the prototype vectors can also be optimized in other ways, which will not be elaborated here. It can be understood that the prototype vectors optimized through multiple parameter adjustment cycles can be used as fixed reference vectors for the process of authenticating the authenticity of voices in audio information. In an alternative embodiment, the undetermined parameters adjusted by the model loss can also include the model parameters in the encoding network used to encode the audio information, which will not be elaborated here.

[0061] Reviewing the above process, the method for identifying the authenticity of voices in audio information provided under the technical concept of this specification can refine the classification categories. Specifically, the classification categories of real voices or forged voices are each refined into hidden categories within the latent space, and a single hidden category is represented by a single prototype vector. After encoding the audio information to be recognized to obtain the corresponding encoding vector, the encoding vector can be compared with each prototype vector respectively to obtain the corresponding similarities, and then, based on the similarities, the authenticity of the voice in the audio information to be recognized is determined, that is, whether it belongs to the real voice classification category or the forged voice classification category. In this way, one or more data centers can be implicitly determined for a single classification category, improving the accuracy of identifying the authenticity of voices in audio information.

[0062] According to an embodiment of another aspect, there is also provided an apparatus for identifying the authenticity of voices in audio information. The apparatus can be provided in a computer, a terminal, or a server with certain computing capabilities. More specifically, as Figure 1 shown in the computing platform. Figure 5 FIG. 5 shows an apparatus 500 for identifying the authenticity of voices in audio information according to an embodiment. As Figure 5 shown, the apparatus 500 may include:

[0063] An encoding unit 501 configured to encode the first audio information to be recognized to obtain a first encoding vector;

[0064] A detection unit 502 configured to detect the similarities respectively corresponding to the first encoding vector and each prototype vector, where a single prototype vector is used to describe a single hidden category within the latent space of two classification categories of real voices or forged voices, and at least one hidden category corresponds to a single classification category;

[0065] An identification unit 503 configured to determine the authenticity of the voice in the first audio information by using the similarities.

[0066] According to some optional implementation manners, the encoding unit 501 may further be configured to: extract voice features from the first audio information; encode or embed the voice features to obtain a first encoding vector. The voice features may include, for example, at least one of the following: loudness, pitch, frequency, timbre, musical tone, duration, and harmonic structure.

[0067] In one embodiment, the encoding unit 501 may further be configured to: before or after encoding or embedding the voice features, perform a normalization operation on the data in each feature dimension, and the normalization operation is one of the following operations: rescaling, mean normalization, standardization, and unit length normalization.

[0068] According to a possible design, the apparatus 500 further includes an adjustment unit (not shown), configured to adjust each prototype vector during the training phase in the following manner: in a single parameter adjustment cycle, process the sample audio information of the current batch to obtain the predicted classification category; compare the predicted classification category with the corresponding classification label to obtain the current model loss; update each undetermined parameter according to the model loss to adjust each prototype vector.

[0069] Wherein, the undetermined parameter may include one of the following: the values of each dimension in the prototype vector; the model parameters in the embedding network that embeds the one-hot representation of each hidden category to obtain the prototype vector.

[0070] In one embodiment, the recognition unit 503 may further be configured to: detect the first classification category to which the prototype vector corresponding to the maximum similarity belongs; determine the authenticity of the sound in the first audio information according to the first classification category. Wherein, the first classification category is a real sound classification category or a forged sound classification category.

[0071] In another embodiment, the recognition unit 503 may further be configured to: use each similarity to determine the first fusion similarity and the second fusion similarity corresponding to the two classification categories of real sound and forged sound respectively; determine the authenticity of the sound in the first audio information according to the magnitudes of the first fusion similarity and the second fusion similarity. Wherein, the first fusion similarity or the second fusion similarity is obtained by performing a fusion operation on each similarity corresponding to each prototype vector of the corresponding classification category.

[0072] It should be noted that, when the number of prototype vectors corresponding to the real sound classification category and the forged sound classification category is the same, the fusion operation includes at least one of: weighted summation, summation, mean value calculation, maximum value taking; when the number of prototype vectors corresponding to the real sound classification category and the forged sound classification category is different, the fusion operation includes at least one of: weighted summation, mean value calculation, maximum value taking, wherein the weight value in the weighted summation operation is negatively correlated with the number of prototype vectors corresponding to the corresponding classification category.

[0073] It should be noted that Figure 5 the illustrated apparatus 500 corresponds to Figure 4 the described method, Figure 4 and the corresponding descriptions in the illustrated method embodiments also apply to the apparatus 500 and will not be elaborated herein.

[0074] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the above computer program is executed in a computer, the computer is made to execute the method described in conjunction with Figure 4 etc.

[0075] According to an embodiment of still another aspect, a computing device is further provided, including a memory and a processor. Executable code is stored in the memory. When the processor executes the above executable code, the method described in combination with Figure 4 and so on is implemented. Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0076] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the technical concept of this specification. It should be understood that the above description is only the specific embodiments of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of this specification should be included in the protection scope of the technical concept of this specification.

Claims

1. A method for identifying the authenticity of sound in audio information, the method comprising: Encoding the first audio information to be identified to obtain a first encoding vector; Detecting the similarities between the first encoding vector and each prototype vector respectively, wherein a single prototype vector is used to describe a single hidden category in a latent space corresponding to two classification categories of real sound or fake sound, and a single classification category corresponds to at least one hidden category; The authenticity of the sound in the first audio information is determined by using the respective similarities.

2. The method of claim 1, wherein: The encoding of the first audio information to be identified to obtain a first encoding vector comprises: Extracting sound features from the first audio information, the sound features including at least one of the following: loudness, pitch, frequency, timbre, musical sound, duration, harmonic structure; The sound feature is encoded or embedded to obtain the first encoding vector.

3. The method of claim 2, wherein: The encoding or embedding of the sound feature to obtain the first encoding vector comprises: Before or after encoding or embedding the sound features, the data on each feature dimension are normalized respectively, and the normalization operation is one of the following operations: rescaling, mean normalization, standardization, and unit length normalization.

4. The method of claim 1, wherein: Each prototype vector is adjusted during the training phase in the following way: In a single parameter adjustment cycle: process the sample audio information of the current batch to obtain the predicted classification category; compare the predicted classification category with the corresponding classification label to obtain the current model loss; update each pending parameter according to the model loss to adjust each prototype vector; The undetermined parameters include one of the following: the numerical value of each dimension in the prototype vector; and the model parameters in the embedding network of the prototype vector obtained by embedding the one-hot representation of each hidden category.

5. The method of claim 1, wherein: Determining the authenticity of the sound in the first audio information by using each similarity comprises: Detecting a first classification category to which the prototype vector corresponding to the maximum similarity belongs, the first classification category being a real sound classification category or a forged sound classification category; The authenticity of the sound in the first audio information is determined according to the first classification category.

6. The method of claim 1, wherein: Determining the authenticity of the sound in the first audio information by using each similarity comprises: Using each similarity, determine a first fused similarity and a second fused similarity corresponding to two classification categories, namely, the real sound and the forged sound, respectively, wherein the first fused similarity or the second fused similarity is obtained by fusing each similarity corresponding to each prototype vector of the corresponding classification category; The authenticity of the sound in the first audio information is determined according to a larger value between the first fusion similarity and the second fusion similarity.

7. The method of claim 6, wherein: When the number of prototype vectors corresponding to the real sound classification category and the forged sound classification category is consistent, the fusion operation includes: at least one of weighted summation, addition, averaging, and maximum value; When the number of prototype vectors corresponding to the real sound classification category and the forged sound classification category are inconsistent, the fusion operation includes: at least one of weighted summation, averaging, and maximum value, wherein the weight in the weighted summation operation is negatively correlated with the number of prototype vectors corresponding to the corresponding classification category.

8. A device for identifying the authenticity of sound in audio information, the device comprising: An encoding unit, configured to encode the first audio information to be identified to obtain a first encoding vector; a detection unit configured to detect respective similarities respectively corresponding to the first encoding vector and respective prototype vectors, wherein a single prototype vector is used to describe a single hidden category in a latent space corresponding to two classification categories of real sound or forged sound, and a single classification category corresponds to at least one hidden category; The recognition unit is configured to determine the authenticity of the sound in the first audio information by using each similarity.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 7.

10. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Sound data processing method and device

    CN120636414A

  • Sound data processing method and apparatus

    CN120636414B

  • Block chain evidence storage method and system based on audio authenticity identification technology

    CN121191538A

  • A blockchain storage method and system based on audio authenticity identification technology

    CN121191538B