Optimization method and device for identifying sound authenticity hidden space in audio information

By expanding the classification category of audio information into multiple hidden categories in hidden space, optimizing the prototype vector, and using a multi-loss function optimization model, the problem of insufficient accuracy of authenticity and sound identification in the existing technology is solved, and higher identification accuracy and generalization capabilities are achieved.

CN120412641APending Publication Date: 2025-08-01ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510537258.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

When identifying authenticity and false sounds in audio information, the prior art has problems such as insufficient classification accuracy and poor generalization ability, especially when facing complex forged audio, it is difficult to effectively distinguish.

Method used

By extending the classification category into multiple hidden categories in hidden space, using prototype vectors to describe real sounds and forged sounds, the prototype vector is optimized to improve the identification accuracy, and a multi-loss function optimization model is adopted, including classification loss, intra-class loss and inter-class loss, and the similarity comparison between the encoded vector and the prototype vector is used for identification.

Benefits of technology

It improves the accuracy and generalization ability of authenticity of audio information, can more effectively identify complex forged audio, and enhances the classification effectiveness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412641A_ABST
    Figure CN120412641A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an optimization method and device for identifying sound authenticity hidden space in audio information, and on the basis that classification categories are refined into a plurality of hidden categories of the hidden space and prototype vectors are used for describing the hidden categories, the prototype vectors are optimized. In the optimization process, the sample audio information is coded, and after a corresponding coding vector is obtained, the coding vector and each prototype vector can be compared with each other to obtain each corresponding similarity. Therefore, intra-class loss is determined through similarity comparison between prototype vectors of the same class, inter-class loss is determined through similarity comparison between prototype vectors of different classes, and classification loss is determined through similarity comparison between coding vectors of audio information and prototype vectors of the classification class and prototype vectors of other classification classes. The optimized prototype vector can improve the classification effectiveness and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to an optimization method and device for an implicit space for identifying the authenticity of voices in audio information. Background Art

[0002] Generative Artificial Intelligence (AIGC) is an artificial intelligence technology that can autonomously learn from data and generate new content including images, sounds, texts, etc. Taking synthetic speech as an example, text can be converted into audible speech audio. During the process of synthesizing speech, the computer will convert the text into natural speech through a speech synthesis engine according to the input text content. Potential risks that forged audio information may bring include, for example, the spread of false information, privacy leakage, and the destruction of social trust. Therefore, there are security risks in speech synthesis or conversion technology in the AIGC era. Taking corresponding measures to avoid these risks and prevent the abuse of this technology is beneficial to protecting personal privacy and social security. Summary of the Invention

[0003] One or more embodiments of this specification describe an optimization method and device for an implicit space for identifying the authenticity of voices in audio information to solve one or more problems mentioned in the background art.

[0004] According to a first aspect, there is provided an optimization method for an implicit space for identifying the authenticity of voices in audio information. In the implicit space, each of the two classification categories of real voices and forged voices corresponds to at least one hidden category, and a single hidden category is described by a single prototype vector. The method includes multiple optimization cycles. In a single optimization cycle: obtaining first sample audio information and its corresponding first classification category, where the first classification category is one of the two classification categories of real voices and forged voices; encoding the first sample audio information to obtain a first encoded vector; determining a model loss by comparing the first encoded vector with each prototype vector, where the model loss includes: a first loss determined by comparing the similarity between the first encoded vector and each prototype vector according to the first classification category; a second loss determined based on the comparison between prototype vectors corresponding to the same classification category; and a third loss determined based on the comparison between prototype vectors corresponding to different classification categories; and optimizing each prototype vector according to the model loss.

[0005] In one embodiment, determining the model loss by comparing the first encoding vector with each prototype vector includes: comparing each first similarity corresponding to each prototype vector of the first classification category with the first encoding vector, and each second similarity corresponding to each prototype vector of other classification categories with the first encoding vector; determining the first loss according to each first similarity and each second similarity, where the first loss is negatively correlated with each first similarity and positively correlated with each second similarity.

[0006] In one embodiment, determining the model loss by comparing the first encoding vector with each prototype vector includes: comparing each third similarity between every two prototype vectors corresponding to the same classification category; determining the second loss according to each third similarity, and the second loss is negatively correlated with each third similarity.

[0007] In a further embodiment, the second loss includes: determining a first loss term and a second loss term, where the first loss term is obtained by adding a predetermined first balance value to each third similarity corresponding to the true voice classification category, and the second loss term is obtained by adding a predetermined second balance value to each third similarity corresponding to the forged voice classification category; performing a weighted sum on the first loss term and the second loss term to obtain the second loss, and before the weighted sum, normalizing the first loss term and the second loss term.

[0008] In one embodiment, determining the model loss by comparing the first encoding vector with each prototype vector includes: comparing the similarity between n1 prototype vectors corresponding to the true voice and n2 prototype vectors corresponding to the forged voice respectively, to obtain corresponding n1×n2 fifth similarities; determining the third loss according to the maximum value among each fifth similarity.

[0009] In a further embodiment, the maximum value is the smoothed maximum value of each fifth similarity.

[0010] In one embodiment, the model loss includes the weighted sum of the first loss, the second loss, and the third loss, and the weighted weights of each loss are preset in advance.

[0011] In one embodiment, optimizing each prototype vector according to the model loss includes: adjusting each undetermined parameter in the direction of reducing the model loss, where the undetermined parameter includes: the value of each dimension in the prototype vector, or the model parameter in the embedding network that embeds the one-hot representation of each hidden category to obtain the prototype vector; updating each prototype vector according to the adjusted undetermined parameter.

[0012] According to a second aspect, there is provided a method for identifying audio information using a sound authenticity latent space. In the latent space, each of the two classification categories of real sound and forged sound corresponds to at least one hidden category, and a single hidden category is described by a single prototype vector. Each prototype vector is optimized in the manner described in claim 1. The method includes: obtaining current audio information to be identified; encoding the current audio information to obtain a current encoded vector; and determining an identification result corresponding to the current audio information by comparing the similarity between the current encoded vector and each prototype vector respectively.

[0013] In one embodiment, determining an identification result corresponding to the current audio information by comparing the similarity between the current encoded vector and each prototype vector respectively includes: detecting each similarity between the current encoded vector and each prototype vector; determining the current classification category to which the prototype vector corresponding to the maximum value among the similarities belongs, where the current classification category is one of real sound and forged sound; and determining the sound authenticity identification result in the current audio information according to the current classification category.

[0014] In one embodiment, determining an identification result corresponding to the current audio information by comparing the similarity between the current encoded vector and each prototype vector respectively includes: detecting each similarity between the current encoded vector and each prototype vector; determining a first fusion similarity and a second fusion similarity respectively corresponding to the two classification categories of real sound and forged sound, where the first fusion similarity or the second fusion similarity is obtained by performing a fusion operation on each similarity corresponding to the corresponding classification category; and determining the sound authenticity identification result in the current audio information according to the larger value between the first fusion similarity and the second fusion similarity.

[0015] According to a third aspect, there is provided an optimization device for a latent space for identifying sound authenticity in audio information. In the latent space, each of the two classification categories of real sound and forged sound corresponds to at least one hidden category, and a single hidden category is described by a single prototype vector. The device includes:

[0016] An acquisition unit configured to acquire sample audio information and its corresponding classification category, where the classification category is one of the two classification categories of real sound and forged sound;

[0017] An encoding unit configured to encode the sample audio information to obtain a corresponding encoded vector;

[0018] A determination unit, configured to determine a model loss by comparing the encoded vector with each prototype vector, where the model loss includes: a first loss determined by comparing the corresponding encoded vector with each prototype vector according to the corresponding classification category, a second loss determined based on the comparison between the prototype vectors corresponding to the same classification category, and a third loss determined based on the comparison between the prototype vectors of different classification categories;

[0019] An adjustment unit, configured to optimize each prototype vector according to the model loss.

[0020] According to a fourth aspect, there is provided an apparatus for authenticating audio information using a voice authenticity latent space. In the latent space, each of the two classification categories of real voice and forged voice corresponds to at least one hidden category, and a single hidden category is described by a single prototype vector. Each prototype vector is optimized by the optimization apparatus according to the third aspect; the apparatus for authenticating audio information includes:

[0021] An acquisition unit, configured to acquire current audio information to be authenticated;

[0022] An encoding unit, configured to encode the current audio information to obtain a current encoded vector;

[0023] An authentication unit, configured to determine an authentication result corresponding to the current audio information by comparing the similarity between the current encoded vector and each prototype vector.

[0024] According to a fifth aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method according to the first aspect or the second aspect.

[0025] According to a sixth aspect, there is provided a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method according to the first aspect or the second aspect is implemented.

[0026] Through the apparatus and method provided by the embodiments of the present specification, on the basis of refining the classification category into multiple hidden categories in the latent space and describing the hidden categories with prototype vectors, the prototype vectors are optimized. During the optimization process, after encoding the sample audio information to obtain the corresponding encoded vector, the encoded vector and each prototype vector can be compared with each other to obtain the corresponding similarities. Thus, the within-class loss is determined by comparing the similarities between the prototype vectors of the same class, the between-class loss is determined by comparing the similarities between the prototype vectors of different classes, and the classification loss is determined by comparing the similarity between the encoded vector of the audio information and the prototype vectors of its own classification category and other classification categories. The prototype vectors optimized in this way can improve the classification effectiveness and generalization ability. Description of the Drawings

[0027] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0028] Figure 1 shows a schematic diagram of a specific implementation architecture for identifying the authenticity of voices in audio information in the conventional technology;

[0029] Figure 2 shows a schematic diagram of a specific implementation architecture for optimizing the hidden space for identifying the authenticity of voices in audio information according to an embodiment of the present specification;

[0030] Figure 3 shows a schematic diagram of the process for identifying audio information using the hidden space for voice authenticity;

[0031] Figure 4 shows a schematic diagram of the comparison of the principles or effects of identifying audio information using the hidden space for voice authenticity and the conventional technology for identifying audio information;

[0032] Figure 5 shows a schematic diagram of the optimization process for the hidden space for identifying the authenticity of voices in audio information according to an embodiment of the present specification;

[0033] Figure 6 shows a block diagram of the structure of an optimization device for the hidden space for identifying the authenticity of voices in audio information according to an embodiment of the present specification;

[0034] Figure 7 shows a block diagram of the structure of a device for identifying audio information using the hidden space for voice authenticity according to an embodiment of the present specification. Specific Embodiments

[0035] The following will describe the solutions provided in this specification in conjunction with the accompanying drawings.

[0036] Similar to text, audio, as a form of information expression, can convey relevant information through sound. Audio information can be understood as voice information. In this specification, in the audio information generated by intelligent devices such as computers, the sound is usually synthesized according to relevant sound characteristics, and can be understood as the sound in the audio information being forged, while the voice collected by means such as recording can be understood as real sound. Here, the real sound and forged sound are usually human voices, but other sounds made by other things (such as the calls of other animals, etc.) are not excluded.

[0037] In the conventional technology, for the audio information to be identified, usually the voice features are extracted, the voice features are encoded, and then through a binary classification model, it is mapped into two classifications: real voice and forged voice. As Figure 1 shown, a realization scheme of the authenticity identification of the voice in the audio information involved in this specification in the conventional technology is shown. As Figure 1 shown, the authenticity identification of the voice in the audio information can be performed by a computing platform. This computing platform can be implemented by hardware (such as a computer module, etc.) or software (such as an application deployed on an operating system). Referring to Figure 1 shown, the computing platform can receive the audio information whose authenticity is to be identified, extract the features of the audio information, encode the extracted features, and then perform the authenticity discrimination of the voice according to the encoding result, and output the discrimination result.

[0038] Considering the complexity of the actual situation, the traditional binary classification model may be limited by the sample source, sample quality, etc. In view of this, a technical solution for describing the latent space of the authenticity of the voice through prototype vectors is proposed. This technical solution is equivalent to expanding the number of classification categories, from two classification categories of real voice and forged voice to multiple hidden categories in the latent space, and a single hidden category can be described by a single prototype vector. Each of the real voice and the forged voice can correspond to one or more hidden categories. In this way, the differences in the data distribution between the forged voice and the real voice can be mined from multiple dimensions, and the generalization ability of identifying the authenticity of the voice can be improved.

[0039] Figure 2 shown is a schematic diagram of a specific implementation architecture under the concept of expanding classification categories through the latent space. Referring to Figure 2 shown, compared with the implementation architecture example shown in Figure 1 shown, under the concept of expanding classification categories through the latent space, after the feature extraction (the dashed box indicates that the corresponding processing can be omitted) and encoding processing of the audio information to be recognized, the encoding vector is compared with each prototype vector for similarity, so as to judge whether the voice in the audio information is a real voice or a forged voice according to the comparison result.

[0040] Figure 3 shown is a schematic diagram of a process for identifying audio information under the concept of expanding classification categories through the latent space. As Figure 3 shown, this process may include the following steps: Step 301, obtain the current audio information to be identified; Step 302, encode the current audio information to obtain the current encoding vector; Step 303, use the similarity comparison between the current encoding vector and each prototype vector to determine the identification result corresponding to the current audio information.

[0041] In Step 301, obtain the current audio information to be identified.

[0042] Here, the current audio information can be any piece of audio information to be authenticated. In this specification, limitations such as "first", "second", "current", etc. are for the convenience of description and to align or distinguish related concepts in terms of the subject, and do not constitute a substantial limitation to the related concepts. For example, "the current audio information" can be any piece of audio information to be authenticated, and "the current encoding vector" is the encoding vector corresponding to "the current audio information", for the alignment of the processing subject.

[0043] In an optional embodiment, the current audio information can also be preprocessed audio information. The preprocessing can, for example, make the length of the audio meet a predetermined condition, filter noise, etc. Taking making the length of the audio meet a predetermined condition as an example, it can be limited by passing the length of the audio through a predetermined duration or the duration falling within a predetermined range. In the case where the audio information is too short, it can be copied to reach a length that meets the predetermined condition. In the case where the audio information is too long, it can be truncated to meet the predetermined condition.

[0044] In step 302, the current audio information is encoded to obtain the current encoding vector.

[0045] The current encoding vector is a vector obtained by encoding the current audio information through an embedding network or an encoding network. The embedding network or the encoding network can be implemented by an existing network or a network constructed by a preset architecture (such as a network based on architectures such as the Transformer network, fully connected network, attention network, convolutional neural network, etc. based on the attention mechanism), which is not limited here.

[0046] In one embodiment, encoding the current audio information can be directly encoding it using an encoding network, for example, processing it using a multimodal large model, a pre-constructed Transformer architecture model based on the attention mechanism, etc. to obtain the current encoding vector.

[0047] In another embodiment, it may be to extract sound features from the current audio information, and then encode or embed the extracted sound features to obtain the current encoded vector. The sound features may include, for example, but are not limited to one of the following: loudness, pitch, frequency, timbre, musical sound, duration, harmonic structure, and so on. Among them, loudness describes the size of the sound and is determined by the amplitude, pitch describes the high or low of the sound (high pitch, low pitch) and is determined by the "frequency", frequency is the number of sound waves passing through a given point per second, timbre is also called tone quality and is determined by the waveform, musical sound is a regular and pleasant sound, noise can be the sound outside the main audio emitter (such as a person), such as the sound emitted by other emitters, and the harmonic structure can be the Fourier series decomposition of a periodic alternating quantity to obtain components with frequencies that are integer multiples greater than 1 of the fundamental frequency. The audio feature network can be various pre-trained sound feature extraction networks, which can use conventional sound feature extractors, such as the wav2vec[xlsr2_300m] model, the WavLM pre-trained model, or can also use a feature extraction network determined by at least one neural network such as a fully connected layer, a convolutional neural network, etc., which is not limited here.

[0048] In other embodiments, the current audio information can also be encoded by other reasonable means, which will not be elaborated here.

[0049] In step 303, the discrimination result corresponding to the current audio information is determined by comparing the similarity between the current encoded vector and each prototype vector.

[0050] It can be understood that each prototype vector can be a vector of each dimension constituting the latent space. Here, the latent space can be understood as a multi-dimensional hidden space. A single prototype vector is used to describe a more refined hidden category within the latent space of one of the two classification categories of real sound category or forged sound category, and can also be used as the central coordinate of the hidden category. As described above, the number of prototype vectors can be set in advance, and the dimension of the prototype vector is usually the same as that of the encoded vector.

[0051] The similarity between the current encoding vector and each prototype vector can be measured by one of the following: cosine similarity, Manhattan distance, Euclidean distance, Chebyshev distance, Pearson correlation coefficient, Jaccard coefficient, Hamming distance, variance, cross entropy, angle, etc. Specifically, the vector similarity can be negatively correlated with at least one of the Manhattan distance, Euclidean distance, Chebyshev distance, Hamming distance, variance, angle, etc., or positively correlated with at least one of the cosine similarity, Pearson correlation coefficient, Jaccard coefficient, cross entropy, etc. In other words, the higher the similarity between a single prototype vector and the current encoding vector, the closer the current encoding vector approaches the single prototype vector. Considering the vectors as spatial point coordinates, the distance between the corresponding spatial points of the two is closer.

[0052] Regarding a single prototype vector as the central coordinate of a hidden category, each prototype vector represents the center of a more refined hidden category. In this way, a single similarity is equivalent to the representation of the distance between the encoding feature of the audio information to be recognized and the center. The higher the similarity, the smaller the distance; conversely, the lower the similarity, the larger the distance. Therefore, the authenticity classification of the sound in the audio information to be recognized can be determined based on each similarity, that is, classified into the classification category of real sound or the classification category of forged sound.

[0053] In some alternative implementation manners, the current classification category to which the prototype vector corresponding to the maximum similarity belongs can be detected, and the authenticity of the sound in the current audio information can be determined according to the current classification category. Among them, the current classification category is one of the real sound classification category or the forged sound classification category. It can be understood that the prototype vector corresponding to the maximum similarity can be regarded as the center of the hidden category closest to the current audio information, and the classification category of real sound or forged sound corresponding to this hidden category center is the classification category to which the current audio information can be classified.

[0054] As a specific example, assume that the prototype vectors corresponding to real sounds are marked by numbers 1 to n1 (from 1 to n1), and the prototype vectors corresponding to forged sounds are marked by numbers n1 to n2 (from n1 to n2). If the number of the prototype vector corresponding to the maximum similarity falls within the range of 1 to n1, then the current audio information can be recognized as a real sound; if the number of the prototype vector corresponding to the maximum similarity falls within the range of n1 to n2, then the current audio information can be recognized as a forged sound.

[0055] In some other alternative implementations, the first fusion similarity and the second fusion similarity corresponding to the two classification categories of the real voice and the forged voice can be determined for the current audio information by using each similarity, and the current authenticity prediction result can be determined for the current audio information according to the magnitudes of the first fusion similarity and the second fusion similarity. Among them, the first fusion similarity or the second fusion similarity can be obtained by performing a fusion operation on each similarity corresponding to the corresponding classification category. In this way, it can be further clarified which prototype vector (or hidden classification) the current audio information tends to approach in the latent space.

[0056] When the number of prototype vectors corresponding to the two classification categories of the real voice and the forged voice is equal, the fusion operation of each similarity can be, for example, one of summation, weighted summation, mean value calculation, maximum value taking, etc. Among them, in the weighted summation process, the weights can be preset or determined by using the number of prototypes, etc., such as the reciprocal of the number of prototypes. Since the number of prototype vectors (i.e., the number of similarities) corresponding to the two classification categories is the same, summation is also equivalent to weighted summation with a weight of 1 for each similarity.

[0057] When the number of prototype vectors corresponding to the real voice classification category and the forged voice classification category is inconsistent, the fusion operation of each similarity can include at least one of weighted summation, mean value calculation, and maximum value taking. Among them, the weight in the weighted summation operation is negatively correlated with the number of prototype vectors corresponding to the corresponding classification category. At this time, since the number of similarities corresponding to the two classification categories is different, directly summing the similarities may lead to unreasonable results. Therefore, the summation method is usually not used to determine the first fusion similarity and the second fusion similarity. In the weighted summation stage, this imbalance in the number of prototype vectors can be balanced by the weights. For example, the weight is negatively correlated with the number of prototype vectors corresponding to the corresponding classification category. As an example, assuming that the number of prototype vectors corresponding to the real voice classification category is n1 and the number of prototype vectors corresponding to the forged voice classification category is n2, the weights of the similarities of each prototype vector used to determine the first fusion similarity can be positively correlated with 1 / n1, n2 / n1, n2 / (n1 + n2), etc. (these values are all negatively correlated with n1).

[0058] In more embodiments, other reasonable methods can also be used to compare the similarity between the current encoded vector and each prototype vector to determine the corresponding voice authenticity discrimination result for the current audio information, which will not be elaborated here.

[0059] Generally, prototype vectors can be divided into prototype vectors corresponding to genuine sound categories and prototype vectors corresponding to forged sound categories. A single classification category can correspond to at least one prototype vector. The number of prototype vectors can be preset, and generally can include at least one prototype vector corresponding to genuine sound and multiple prototype vectors corresponding to forged sound.

[0060] In one embodiment, considering that the features of genuine sound are relatively unified, the number of prototype vectors corresponding to genuine sound is small, such as n1, and the number of prototype vectors corresponding to forged sound is n2. Both n2 and n1 are natural numbers, and n2 > n1. Optionally, n1 can be set to 1.

[0061] In another embodiment, the number of prototype vectors corresponding to both genuine sound and forged sound is n. n is a natural number greater than 1. In other embodiments, the number of other prototype vectors can also be set according to experience, which is not limited here.

[0062] To make the above technical concept more clear, Figure 4 shows Figure 3 a comparison schematic diagram of the audio information discrimination scheme in Figure 4 and the audio information discrimination scheme of the conventional technology in terms of principle or effect. In Figure 1 , above the dotted line is equivalent to Figure 4 the technical scheme shown, and below the dotted line corresponds to the technical concept of this specification. Among them, the audio information represented by the small circle ○, and the triangle ▲ symbol represents the classification space (or classification center). As Figure 4 shown, under the binary classification model, the audio information can be divided into two categories. The left side represents that the sound is genuine, and the right side represents that the sound is fake, corresponding to two triangle ▲ symbols respectively. At this time, the audio information of the forged sound on the edge in the circles 401 and 402 is more likely to be classified into the category of genuine sound, resulting in misclassification. And under the technical concept of expanding the classification category in the hidden space, the audio information is mapped to multiple hidden dimensions through multiple prototype vectors, represented by multiple ▲ symbols. Considering that the features of genuine sound are relatively unified and the sounds in the synthetic audio have their own characteristics,

[0063] It should be noted that Figure 4 is only an example. In practice, there can also be multiple prototype vectors for genuine sound classification in the hidden space, and the number of prototype vectors for forged sound classification in the hidden space can also be other numbers, which are not limited here. Among them, under the above technical concept, before performing the authenticity identification of the sound, the prototype vectors can be determined and used as fixed parameters.

[0064] In order to enable the prototype vectors to better describe the categories in the latent space, the prototype vectors can be adjusted and optimized during the model training process, so as to effectively divide the real voice space and the forged voice space in the latent space. The adjustment process of the prototype vectors is distinguished according to different settings of the prototype vectors.

[0065] In one embodiment, the values of each dimension in the prototype vector can be used as undetermined parameters. At this time, each dimension in the prototype vector can be initialized to a random value, and the values of each dimension will be directly adjusted as undetermined parameters in subsequent parameter cycles.

[0066] In another embodiment, each prototype vector can be obtained by embedding the initial representation vector using an embedding network. At this time, the undetermined parameters can be the model parameters in the embedding network. The initial representation tensor can be set manually or determined by various coding methods. For example, the initial representation vector is a one-hot representation. A single initial representation vector has one dimension of 1 and other dimensions of 0, and the one-hot representations corresponding to each prototype vector are different. At this time, since the representation of the initial representation vector is too single and not conducive to describing the data distribution, the corresponding prototype vector can be obtained by embedding the one-hot representation describing the hidden category using an embedding network. The model parameters in the embedding network can be used as undetermined parameters, so that the prototype vector can be affected by adjusting the model parameters in the embedding network.

[0067] In other embodiments, the prototype vectors can also be adjusted by other means, which will not be elaborated here.

[0068] According to the usual supervised model training idea, the prototype vectors can be adjusted and optimized through multiple optimization cycles by using the sample audio information and its classification label of whether it is a real voice or a forged voice. In a single optimization cycle, a batch of sample audio information can be processed to obtain the predicted classification category, and the predicted classification category is compared with the corresponding classification label to obtain the current model loss. According to the model loss, the update gradient of each undetermined parameter can be determined, and then the undetermined parameters can be updated using a gradient update method such as the gradient descent method to adjust each prototype vector. The prototype vectors adjusted during the training process can be directly used to compare the similarity with the encoded vectors to identify the authenticity of the voice in the audio information.

[0069] However, considering that deepfake audio usually comes from various different audio generation models or tools and covers multiple fields, the hidden classes in the latent space need to be effectively differentiated. In order to enable the prototype vectors to better describe the hidden classes in the latent space, this specification also provides a technical concept of refining and optimizing the prototype vectors, so that they can more generally and effectively describe the latent space. This optimized technical solution introduces the differences between prototype vectors into the model loss, and based on the classification category to which the sample audio information belongs, introduces the adjustment of the relationship between the encoded vector of the sample audio information and the prototype vector, so as to use the differentiation between prototype vectors and the supervision of the classification category provided by the label of the sample as a common guide to optimize the prototype vectors, thereby refining the latent space and enhancing the generalization of the model.

[0070] The following describes the technical concept of this specification in detail with reference to the accompanying drawings.

[0071] Figure 5 The figure shows a flowchart of an optimization method for the latent space for identifying the authenticity of voices in audio information proposed in an embodiment of this specification. The execution subject of this process can be any computer, device, or server with a certain computing power. Before describing the optimization process for the latent space for identifying the authenticity of voices in audio information, it should be noted that the latent space can include at least one hidden class corresponding to real voices and forged voices respectively, and a single hidden class is described by a single prototype vector. This optimization process can include multiple optimization cycles. Figure 5 Taking the first training sample as an example, the optimization process of a single optimization cycle is shown. Among them, the first training sample can be any sample in the sample set, and it can correspond to the first sample audio information and the first classification label.

[0072] Refer to Figure 5 As shown, the optimization process for the latent space for identifying the authenticity of voices in audio information can include the following steps: Step 501, obtain the first sample audio information and its corresponding first classification category; Step 502, encode the first sample audio information to obtain the first encoded vector; Step 503, compare the first encoded vector with each prototype vector to determine the model loss, where the model loss includes: the first loss determined by comparing the similarity between the first encoded vector and each prototype vector according to the first classification category, the second loss determined based on the comparison between the prototype vectors corresponding to the same classification category, and the third loss determined based on the comparison between the prototype vectors of different classification categories; Step 504, optimize each prototype vector according to the model loss.

[0073] First, in step 501, obtain the first sample audio information and its corresponding first classification category.

[0074] Here, the first sample audio information is the audio information in the first training sample. Among them, in the current optimization cycle, at least one training sample can be used to adjust the undetermined parameters. The first training sample can be any training sample obtained in the current batch according to the sample acquisition rule. The sample acquisition rule is, for example, to obtain samples in the sample set in sequence, or randomly obtain samples in the sample set, and so on. The first sample audio information is the audio information obtained during the sample acquisition process. In an alternative embodiment, before using the collected audio information as the sample audio information, it can also be preprocessed, and the preprocessed audio information can be used as the sample audio information. The preprocessing can, for example, make the length of the audio meet a predetermined condition, filter out noise, etc.

[0075] The first training sample can also correspond to a first classification label. The first classification label can be a predetermined value (such as using 1 to represent a real voice and 0 to represent a fake voice, etc.), or it can be a vector (such as a two-dimensional vector composed of 1 and 0, where the dimension corresponding to 1 represents the classification category to which it belongs). In this way, according to the first classification label, it can be determined whether the voice in the audio information corresponding to the first training sample is a real voice or a forged voice.

[0076] Then, through step 502, the first sample audio information is encoded to obtain a first encoded vector.

[0077] The first encoded vector can be a vector obtained by encoding the first sample audio information through an embedding network or an encoding network. The embedding network or the encoding network can be implemented through existing networks, or can be implemented through a network constructed by a preset architecture (such as a network based on the Transformer network, fully connected network, attention network, convolutional neural network, etc. architectures based on the attention mechanism). It is not limited here. Among them, in the case of a network constructed by a preset architecture, the architecture in the network can also be used as an undetermined parameter to be adjusted in the current optimization cycle.

[0078] In one embodiment, encoding the first sample audio information can be directly encoding it using an encoding network, for example, using a multimodal large model, a pre-constructed Transformer architecture model based on the attention mechanism, etc. to process it to obtain the first encoded vector.

[0079] In another embodiment, the encoding of the first sample audio information may be to extract sound features from the first sample audio information, and then encode or embed the extracted sound features to obtain a first encoded vector. The sound features may include, for example, but are not limited to, one of the following: loudness, pitch, frequency, timbre, musical tone, duration, harmonic structure, and so on. The sound features may be extracted by an audio feature extraction network. The audio feature extraction network is a network that can extract sound features from the waveform of the audio.

[0080] It can be understood that encoding the sound features is a process of further processing the sound features, performing feature fusion or extracting higher-order features. The encoded vector may be, for example, m-dimensional (m is a natural number greater than 1).

[0081] In an alternative implementation, before or after encoding or embedding the sound features, the encoded vectors of each dimension may also be subjected to normalization (feature scaling) processing. This is because different features may have their own measurement criteria. For example, the loudness of sound is usually described in decibels (dB), usually ranging from 0 to 130 dB, and audio is usually described in vibration hertz (Hz), such as 20 to 2000 Hz, and so on. The numerical ranges of different features may vary greatly. Then, in the process of feature data processing, features with larger numerical values may play a decisive role, while features with smaller numerical values are ignored. Therefore, before encoding the sound features or after encoding (such as embedding or fusion processing), the sound features or encoded vectors may be normalized. The normalization methods may be, for example, one of the following: rescaling (min-max normalization, range scaling), mean normalization, standardization (Z-score normalization), unit length normalization (scaling to unit length), and so on. According to the actual business situation, appropriate normalization operations may be performed on the encoded vectors, which will not be elaborated here. It should be noted that the mean used in mean normalization may be determined by the mean of multiple sample audio data in the corresponding dimension.

[0082] In other embodiments, other suitable methods may also be used to encode the first sample audio information, which will not be elaborated here.

[0083] Next, through step 503, the model loss is determined by comparing the first encoded vector with each prototype vector.

[0084] It can be understood that each prototype vector can be a vector of each dimension constituting the latent space. Here, the latent space can be understood as a multi-dimensional hidden space. A single prototype vector can be regarded as a dimension of the latent space or as the central coordinates of a more refined hidden category in the m-dimensional space. As described above, the number of prototype vectors can be preset, and the dimension of the prototype vector is usually the same as that of the encoded vector, such as m dimensions.

[0085] It can be understood that if a single prototype vector is regarded as the central coordinates of a hidden category, then each prototype vector can represent the center of a more refined hidden category (such as Figure 3 the triangle in). In this way, a single similarity is equivalent to the representation of the distance between the encoded features of the audio information to be recognized and the center. The higher the similarity, the smaller the distance; conversely, the lower the similarity, the larger the distance. Therefore, the authenticity classification of the sound in the audio information to be recognized can be determined according to each similarity, that is, classified into the classification category of real sound or the classification category of forged sound.

[0086] The model loss can be the difference generated from the expected result during the process of processing the first sample audio information to determine the authenticity discrimination result. Usually, the direct expected result includes, for example, the first classification label, etc. Under the technical concept of this specification, in order to better utilize the prototype vector to refine the latent space and make it tend to a fixed value to effectively describe the hidden categories of real sound and forged sound, and the predicted category of the first sample audio information is determined based on its similarity to the prototype vector, so the expected result can be set for the prototype vector. Generally: on the one hand, it is hoped that the encoded vector of the sample audio information (such as the first encoded vector) is as close as possible to the prototype vector corresponding to its classification category (such as determined by the first classification label), and at the same time, as different as possible from the prototype vectors corresponding to other classification categories; on the other hand, it is also hoped that the prototype vectors of the same classification category have a large distance from each other to prevent multiple centers from gathering together; on the other hand, it is also hoped that the prototype vectors of different classification categories have a large distance from each other. In this way, the model loss can include three types, which are introduced one by one below.

[0087] The first loss is the classification loss, denoted as the first loss for example. The first loss can be determined by comparing the similarity between the first encoded vector and each prototype vector according to the first classification category. Among them, the similarity between vectors can be measured by one of the following: cosine similarity, Manhattan distance, Euclidean distance, Chebyshev distance, Pearson correlation coefficient, Jaccard coefficient, Hamming distance, variance, cross entropy, vector angle, etc. Specifically, the vector similarity can be negatively correlated with at least one of the Manhattan distance, Euclidean distance, Chebyshev distance, Hamming distance, variance, vector angle, etc., or positively correlated with at least one of the cosine similarity, Pearson correlation coefficient, Jaccard coefficient, cross entropy, etc. In other words, the higher the similarity between a single prototype vector and the first encoded vector, the closer the first encoded vector approaches the single prototype vector. Considering the vectors as the coordinates of spatial points, the distance between the corresponding spatial points of the two is closer.

[0088] To make the encoded vector of the sample audio information consistent with its actual classification category, the encoded vector can be made as close as possible to the prototype vector corresponding to its actual classification category. At the same time, it should be as different as possible from the prototype vectors corresponding to other classification categories. Specifically, it can be set that: the encoded vector of the sample audio information corresponding to the real sound is as close as possible to the prototype vector corresponding to the real sound (maximizing the similarity or minimizing the distance, angle, etc.), and at the same time, it is as different as possible from the prototype vector corresponding to the forged sound (such as minimizing the similarity or minimizing the distance, angle, etc.). Vice versa, it is hoped that the encoded vector of the sample audio information corresponding to the forged sound is as close as possible to the prototype vector corresponding to the forged sound, and at the same time, it is as different as possible from the prototype vector corresponding to the real sound.

[0089] For the first training sample, if the classification category to which the first sample audio information belongs is the first classification category, then the respective first similarities corresponding to the first encoded vector and each prototype vector corresponding to the first classification category can be calculated and maximized. Additionally, the respective second similarities corresponding to the first encoded vector and each prototype vector corresponding to other classification categories can be calculated and minimized. Therefore, the first loss can be determined based on each first similarity and each second similarity. Specifically, the first loss can be negatively correlated with each first similarity and positively correlated with each second similarity. In this way, when minimizing the first loss, each first similarity can be maximized and each second similarity can be minimized.

[0090] In one embodiment, the first loss may be proportional to the ratio of the sum of the second similarities to the sum of the first similarities, or proportional to the ratio of the sum of the second similarities to the total obtained by adding the sum of the first similarities and the sum of the second similarities, or proportional to the difference obtained by subtracting the maximum first similarity from the maximum second similarity, and so on. For example, in a specific example, the first loss may be determined according to the exponentially weighted values of the first similarities and the second similarities. Let the first parameter be the sum of the exponents obtained by taking the first similarities as exponents, and the second parameter be the sum of the exponents obtained by taking the second similarities as exponents. The ratio of the first parameter to the sum of the first parameter and the second parameter.

[0091] In other embodiments, the first loss may also be determined in other ways via the first similarity and the second similarity, which will not be elaborated here.

[0092] In an alternative implementation, before calculating the similarity between the first encoded vector and each prototype vector, each prototype vector may be separately subjected to normalization (feature scaling) processing. The normalization processing of a single prototype vector is, for example, at least one of rescaling, standardization, mean normalization, etc., which will not be elaborated here. It can be understood that the normalization of the prototype vector generally keeps the data distribution within the vector unchanged.

[0093] In this way, when the first encoded vector is closer to the prototype vector corresponding to the first classification category, the model loss is smaller, and when the first encoded vector is closer to the prototype vector corresponding to other classification categories, the model loss is larger.

[0094] The second loss is the within-class loss, denoted as the second loss, which can be determined based on the comparison between the prototype vectors corresponding to the same classification category. It can be understood that to prevent multiple hidden centers corresponding to the same classification category from clustering together, the prototype vectors corresponding to the same classification category should be made as different as possible. That is, the distance between any two prototype vectors corresponding to the same classification category should be as large as possible, or the similarity should be as small as possible.

[0095] As an example, the similarity between any two prototype vectors corresponding to the same classification category is denoted as the third similarity. Then, the similarity between any two prototype vectors corresponding to the forged sound, or the similarity between any two prototype vectors corresponding to the real sound can both be referred to as the third similarity. Generally, there are various forged sounds, and the prototype vectors corresponding to the real sound can be as few as 1, while there are multiple prototype vectors corresponding to the forged sound. Therefore, the second loss is determined at least via the third similarity between any two prototype vectors corresponding to the forged sound.

[0096] The similarity between pairwise prototype vectors can be measured by one of cosine similarity, Manhattan distance, Euclidean distance, Chebyshev distance, Pearson correlation coefficient, Jaccard coefficient, Hamming distance, variance, cross-entropy, vector angle, etc. As a specific example, in the case where the third similarity is measured by cosine similarity, matrix multiplication can be used to determine each third similarity. Specifically, for example, the prototype vectors corresponding to the forged voices can be arranged in sequence to form a forged center prototype tensor (matrix), and the product of this matrix and itself (i.e., taking the dot product of its transpose matrix and itself) is calculated to obtain a result matrix. Taking the values above or below the diagonal of this result matrix (excluding the diagonal values) gives each third similarity.

[0097] In order to effectively distinguish the prototype vectors of the same classification category, the second loss can be positively correlated with each third similarity. That is, under the same conditions, the greater the third similarity, the greater the second loss, and the smaller the third similarity, the smaller the second loss. For example, the second loss can be positively correlated with the reciprocal of the sum of each third similarity, or negatively correlated with the maximum value among each third similarity, etc. In this way, when minimizing the second loss, the similarity between the prototype vectors of the same classification category can be minimized.

[0098] In one embodiment, to avoid the loss being negative or zero, the second loss can be divided into a first loss term corresponding to the true voice classification category and a second loss term corresponding to the forged voice classification category. The first loss term is obtained by adding a predetermined first balance value to each third similarity corresponding to the true voice classification category, and the second loss term is obtained by adding a predetermined second balance value to each third similarity corresponding to the forged voice classification category. The second loss is the weighted sum of the first loss term and the second loss term, and the weighting weights can be determined in advance and can be equal or both 1 first. Optionally, before the weighted summation, the first loss term and the second loss term are also normalized, and the normalization method can be carried out by using conventional techniques and will not be elaborated here. By normalization, the norm of the vector composed of each third similarity can be limited within a predetermined range, which is convenient for data processing and optimizes the convergence speed.

[0099] The third loss is the inter-class loss, denoted as the third loss, which can be determined based on the comparison between the prototype vectors corresponding to different classification categories in the real voice and the forged voice. Those skilled in the art can understand that the prototype vector corresponding to the real voice and the prototype vector corresponding to the forged voice should distinguish the two classification categories in the latent space. Therefore, the prototype vector corresponding to the real voice should be as far as possible from the prototype vector corresponding to the forged voice, or the similarity should be as small as possible. Thus, the single prototype vector corresponding to the real voice can be compared with each prototype vector corresponding to the forged voice respectively to obtain the corresponding similarities, such as denoted as the fifth similarity. It can be understood that in the case where the real voice corresponds to n1 prototype vectors and the forged voice corresponds to n2 prototype vectors, n1×n2 fifth similarities can be obtained. The third loss can be positively correlated with the fifth similarity, for example, being proportional to the sum of each fifth similarity.

[0100] In one embodiment, the maximum value among each fifth similarity can be selected to determine the third loss. For example, the third loss includes the maximum value among each fifth similarity, or the sum of the maximum value among the fifth similarities and a predetermined value, etc. Here, the predetermined value can be a preset positive number (such as 0.1) to avoid the maximum value among the fifth similarities being negative. Optionally, the maximum value here can be the smoothed maximum value of each fifth similarity. The smoothed maximum value can convert the non-differentiable maximum value determination scheme into a differentiable approximate maximum value function, which is convenient for backpropagation. For example, using ln(e (k·x) +e (k ·y) ) / k as the function for determining the maximum value of x and y, the larger k is, the closer the whole function will be to the maximum value function, and the smaller k is, the smoother the function will become.

[0101] It can be understood that the above first loss, second loss, and third loss all contain expectations for the prototype vector. In other possible designs, the model loss may also include other losses. For example, the loss determined by comparing the first predicted classification category predicted for the first sample audio information with the first classification label, etc., which is not limited here.

[0102] In a single optimization cycle, the model loss can include the weighted sum of the first loss, the second loss, and the third loss, and accumulate on the training samples of the current batch. Among them, the weighting weights can be preset in advance. Optionally, in the case of using multiple batches of training samples in a single optimization cycle, it can also be accumulated or averaged on each batch as the model loss of a single optimization cycle, which will not be elaborated here.

[0103] Furthermore, in step 504, optimize each prototype vector according to the model loss.

[0104] It can be understood that, aiming to minimize the model loss, adjusting the undetermined parameters can complete the parameter optimization for the current cycle. During the process of minimizing the model loss, the following objectives can be achieved: making the encoded vector of the audio information close to the hidden center of the corresponding classification category and far from other hidden centers; dispersing the hidden centers of the same classification category; expanding the distance between the hidden centers of different classification categories, thereby expanding the gap between categories.

[0105] It is worth noting that the undetermined parameters used to optimize each prototype vector can include: the values of each dimension of the prototype vector, or the model parameters in the embedding network that embeds the one-hot representation of each hidden category to obtain the prototype vector, and so on. In an alternative embodiment, the undetermined parameters can also include the model parameters in the encoding network that encodes the audio information, which will not be elaborated here.

[0106] According to the adjusted undetermined parameters, each prototype vector can be updated. When the end condition of model training is not satisfied, the updated prototype vectors can be used to determine the model loss and further update in the next optimization cycle. When the end condition of model training is satisfied, the updated prototype vectors can be used as fixed values for Figure 3 The process of using the true / false sound hidden space to identify audio information is shown. Here, the end condition of model training can include, for example, at least one of the following: the model loss approaches 0, the undetermined parameters tend to converge, the number of optimization cycles reaches a predetermined number (such as 1000), and so on.

[0107] In this way, determining the model loss by comparing the first classification label with the corresponding true / false identification result can be decomposed into comparing the similarity between the first encoded vector and each prototype vector, and comparing the similarity between the prototype vectors, so that the optimized prototype vectors can effectively distinguish the classification categories of audio information in the hidden space, improving the classification accuracy and the generalization ability for identifying forged audio through different channels.

[0108] Reviewing the above process, the optimization method for the hidden space of authenticating the authenticity of voices in audio information provided under the technical concept of this specification can refine the classification categories. Specifically, the classification categories are refined into multiple hidden categories within the hidden space. A single hidden category corresponds to one of the two classification categories of real voices or forged voices, and is described by a single prototype vector. During the optimization process, after encoding the sample audio information to obtain the corresponding encoded vector, the encoded vector can be compared with each prototype vector respectively to obtain the corresponding similarities. On this basis, in order to enable the prototype vectors to effectively distinguish the classification categories, during the optimization of the prototype vectors, the within-class loss can be determined by comparing the similarities between the prototype vectors of the same class, the between-class loss can be determined by comparing the similarities between the prototype vectors of different classes, and the classification loss can be determined by comparing the similarities between the encoded vector of the audio information and the prototype vectors of its belonging classification category and other classification categories. Thus, the optimized prototype vectors can not only effectively distinguish within a class, but also effectively distinguish between classes, and can also be effectively classified in combination with the encoded vector, improving the classification effectiveness and generalization ability.

[0109] According to an embodiment of another aspect, an optimization device for the hidden space of authenticating the authenticity of voices in audio information is further provided. The device can be provided in a computer, a terminal, or a server with certain computing capabilities. More specifically, as Figure 1 shown in the computing platform. Figure 6 FIG. shows an optimization device 600 for the hidden space of authenticating the authenticity of voices in audio information according to an embodiment. As Figure 6 shown, the device 600 may include:

[0110] An acquisition unit 601, configured to acquire sample audio information and its corresponding classification category, where the classification category is one of the two classification categories of real voice and forged voice;

[0111] An encoding unit 602, configured to encode the sample audio information to obtain a corresponding encoded vector;

[0112] A determination unit 603, configured to determine a model loss by comparing the encoded vector with each prototype vector, where the model loss includes a first loss, a second loss, and a third loss. The first loss is determined by comparing the similarities between the corresponding encoded vector and each prototype vector according to the corresponding classification category, the second loss is determined based on the comparison between the prototype vectors corresponding to the same classification category, and the third loss is determined based on the comparison between the prototype vectors of different classification categories;

[0113] An adjustment unit 604, configured to optimize each prototype vector according to the model loss.

[0114] According to an embodiment of another aspect, there is also provided a device for authenticating audio information by using a sound authenticity hidden space. The device can be disposed in a computer, a terminal, or a server with a certain computing power. More specifically, it is, for example, Figure 1 the computing platform shown. Figure 7 FIG. shows a device 700 for authenticating audio information by using a sound authenticity hidden space according to an embodiment. As shown in Figure 7 FIG., the device 700 may include:

[0115] An acquisition unit 701, configured to acquire current audio information to be authenticated;

[0116] An encoding unit 702, configured to encode the current audio information to obtain a second encoded vector;

[0117] An authentication unit 703, configured to determine an authentication result corresponding to the current audio information by comparing the similarity between the current encoded vector and each prototype vector.

[0118] Among them, each prototype vector can be optimized via the Figure 6 device 600 shown.

[0119] It should be noted that Figure 6 , Figure 7 the devices 600 and 700 shown respectively correspond to the Figure 3 , Figure 5 described methods. The corresponding descriptions in the method embodiments shown in Figure 3 , Figure 5 also apply to the devices 600 and 700, and will not be repeated here.

[0120] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the methods described in combination with Figure 3 , Figure 5 and so on.

[0121] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the methods described in combination with Figure 3 , Figure 5 and so on are implemented. Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0122] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above description is only the specific embodiments of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of this specification shall be included within the protection scope of the technical concept of this specification.

Claims

1. An optimization method for an implicit space for identifying the authenticity of sounds in audio information. In the implicit space, each of the two classification categories of real sounds and forged sounds corresponds to at least one hidden category, and a single hidden category is described by a single prototype vector. The method includes multiple optimization cycles. In a single optimization cycle: Obtain first sample audio information and its corresponding first classification category, where the first classification category is one of the two classification categories of real sounds and forged sounds; Encode the first sample audio information to obtain a first encoded vector; Determine a model loss by comparing the first encoded vector with each prototype vector, where The model loss includes: a first loss determined by comparing the similarity between the first encoded vector and each prototype vector according to the first classification category; a second loss determined based on the comparison between prototype vectors corresponding to the same classification category; and a third loss determined based on the comparison between prototype vectors corresponding to different classification categories; Optimize each prototype vector according to the model loss.

2. The method according to claim 1, wherein The determining of the model loss by comparing the first encoded vector with each prototype vector includes: Comparing each first similarity corresponding to the comparison between the first encoded vector and each prototype vector corresponding to the first classification category, and each second similarity corresponding to the comparison between the first encoded vector and each prototype vector corresponding to other classification categories; Determine the first loss according to each first similarity and each second similarity, where the first loss is negatively correlated with each first similarity and positively correlated with each second similarity.

3. The method according to claim 1, wherein The determining of the model loss by comparing the first encoded vector with each prototype vector includes: Comparing each third similarity between every two prototype vectors corresponding to the same classification category; Determine the second loss according to each third similarity, and the second loss is negatively correlated with each third similarity.

4. The method according to claim 3, wherein The second loss includes: Determine a first loss term and a second loss term. The first loss term is obtained by adding a predetermined first balance value to each third similarity corresponding to the real sound classification category, and the second loss term is obtained by adding a predetermined second balance value to each third similarity corresponding to the forged sound classification category; Perform a weighted sum of the first loss term and the second loss term to obtain the second loss. Before the weighted sum, normalize the first loss term and the second loss term.

5. The method according to claim 1, wherein, The determining of the model loss by comparing the first encoded vector with each prototype vector includes: Compare the similarity between n1 prototype vectors corresponding to real sounds and n2 prototype vectors corresponding to forged sounds respectively, to obtain corresponding n1×n2 fifth similarities; Determine the third loss according to the maximum value among each fifth similarity.

6. The method according to claim 5, wherein, The maximum value is the smoothed maximum value of each fifth similarity.

7. The method according to claim 1, wherein The model loss includes the weighted sum of the first loss, the second loss, and the third loss, and the weighted weights of each loss are preset in advance.

8. The method according to claim 1, wherein The optimizing of each prototype vector according to the model loss includes: Adjust each undetermined parameter in the direction of reducing the model loss, where the undetermined parameters include: the values of each dimension in the prototype vector, or the model parameters in the embedding network that embeds the one-hot representations of each hidden category to obtain the prototype vector; Update each prototype vector according to the adjusted undetermined parameters.

9. A method for identifying audio information using a voice authenticity latent space, in the latent space, each of the two classification categories of real voice and forged voice corresponds to at least one hidden category, and a single hidden category is described by a single prototype vector, and each prototype vector is optimized in the manner described in claim 1; the method includes: Obtain the current audio information to be identified; Encode the current audio information to obtain a current encoded vector; Determine the identification result corresponding to the current audio information by comparing the similarity between the current encoded vector and each prototype vector.

10. The method according to claim 9, wherein, Determining the identification result corresponding to the current audio information by comparing the similarity between the current encoded vector and each prototype vector includes: Detect each similarity between the current encoded vector and each prototype vector; Determine the current classification category to which the prototype vector corresponding to the maximum value among the similarities belongs, and the current classification category is one of real voice and forged voice; Determine the voice authenticity identification result in the current audio information according to the current classification category.

11. The method according to claim 9, wherein, Determining the identification result corresponding to the current audio information by comparing the similarity between the current encoded vector and each prototype vector includes: Detect each similarity between the current encoded vector and each prototype vector; Determine a first fusion similarity and a second fusion similarity corresponding to the two classification categories of real voice and forged voice respectively, where the first fusion similarity or the second fusion similarity is obtained by performing a fusion operation on each similarity corresponding to the corresponding classification category; Determine the voice authenticity identification result in the current audio information according to the larger value between the first fusion similarity and the second fusion similarity.

12. An optimization device for a latent space for identifying voice authenticity in audio information, in the latent space, each of the two classification categories of real voice and forged voice corresponds to at least one hidden category, and a single hidden category is described by a single prototype vector; the device includes: An acquisition unit configured to acquire sample audio information and its corresponding classification category, and the classification category is one of the two classification categories of real voice and forged voice; An encoding unit configured to encode the sample audio information to obtain a corresponding encoded vector; A determination unit configured to determine a model loss by comparing the encoded vector with each prototype vector, where the model loss includes: a first loss determined by comparing the similarity between the corresponding encoded vector and each prototype vector according to the corresponding classification category; a second loss determined based on the comparison between the prototype vectors corresponding to the same classification category; a third loss determined based on the comparison between the prototype vectors corresponding to different classification categories; An adjustment unit configured to optimize each prototype vector according to the model loss.

13. A device for identifying audio information using the authenticity hidden space of sound. In the hidden space, each of the two classification categories of real sound and forged sound corresponds to at least one hidden category, and a single hidden category is described by a single prototype vector. Each prototype vector is optimized by the optimization device described in claim 12. The device for identifying audio information includes: An acquisition unit configured to acquire current audio information to be identified; An encoding unit configured to encode the current audio information to obtain a current encoded vector; An identification unit configured to determine the identification result corresponding to the current audio information by comparing the similarity between the current encoded vector and each prototype vector.

14. A computer-readable storage medium having a computer program stored thereon. When the computer program is executed on a computer, the computer is caused to execute the method according to any one of claims 1-11.

15. A computing device, comprising a memory and a processor, characterized in that, Executable code is stored in the memory. When the processor executes the executable code, the method according to any one of claims 1-11 is implemented.