Artificial intelligence-based voice detection server and method

The AI-based voice detection server enhances voice modulation detection by preprocessing and learning from diverse data sets, effectively identifying voice phishing and deepfake attacks through a sophisticated voice detection model.

WO2025150627A1PCT designated stage expired Publication Date: 2025-07-17DEEPBRAIN AI INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/006986
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2024-05-23
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing voice detection systems struggle to effectively identify voice modulation, including voice phishing, as they are not adequately trained to recognize subtle differences between original and modulated voices, especially in the context of deepfake technologies.

Method used

An AI-based voice detection server and method that utilizes a detection model learning unit to preprocess and learn from diverse data sets, including voice synthesis techniques, deepfake data, and speaker-specific information, employing convolutional and fully connected layers to generate a voice detection model capable of determining modulation probability.

Benefits of technology

Improves the detection performance of voice modulation by accurately distinguishing between original and modulated voices, enhancing the ability to detect and prevent voice phishing and deepfake attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024006986_17072025_PF_FP_ABST
    Figure KR2024006986_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an artificial intelligence-based voice detection server and method. The voice detection server according to an embodiment of the present invention comprises: a detection model training unit for generating an artificial intelligence-based voice detection model by performing pre-training using a training data set configured by one or more different formats including original voice data and modulated data of the original voice data, wherein a voice detection model for extracting speaker-unique information from an input value and determining authenticity is generated by performing pre-training through first model training and second model training by using a pre-processed training data set; and a detection processing unit for determining whether input voice data is modulated, on the basis of the modulation probability of the voice data by using the voice detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Artificial intelligence-based voice detection server and method

[0001] The disclosed embodiments relate to an artificial intelligence-based voice detection server and method.

[0002] Voice phishing is a combination of the words voice, private data, and fishing, and can mean deceiving or threatening victims to request personal information or financial transaction information or to transfer money.

[0003] As technology advances, voice phishing is evolving from simply having people manually alter their voices to more sophisticated methods that use various technologies to manipulate the voice.

[0004] Accordingly, various technologies are being proposed to prevent damage from voice manipulation, including voice phishing.

[0005] The disclosed embodiments seek to provide an artificial intelligence-based voice detection server and method for improving the detection performance of voice modulation.

[0006] An artificial intelligence-based voice detection server according to one embodiment includes a detection model learning unit that performs pre-learning using a learning data set composed of at least one different format including original voice data and modulation data of the original voice data to generate an artificial intelligence-based voice detection model, wherein the pre-learning is performed through first model learning and second model learning using the learning data set on which preprocessing has been performed to extract speaker-specific information from input voice data and generate the voice detection model for determining authenticity; and a detection processing unit that determines whether or not there is modulation based on a modulation probability of input voice data using the voice detection model.

[0007] The above learning data set may include data sets in different formats, including a first data set generated by mixing preset first voice synthesis techniques according to preset criteria, a second data set that is a voice deepfake data set, and a third data set generated based on the preset second voice synthesis technique.

[0008] The above detection model learning unit may perform a preprocessing operation according to data selection conditions for distinguishing original voice data and modulated data from each voice file of the training data set, and deleting a voice file that includes at least one of the following: if the original voice data corresponding to the speaker's voice is missing from each voice file, if the quality of the voice file is below a standard value, and if the noise included in the voice file is above a standard value, from the training data set.

[0009] The above detection model learning unit can perform a preprocessing task of adjusting the length of each voice file of the above learning data set to a preset length.

[0010] The above detection model learning unit generates an audio embedding model for extracting audio embedding data representing original characteristics and modulation characteristics to distinguish the original and modulation from audio information of the learning data set through the first model learning, and the first model learning may be convolution (1D CONV)-based audio embedding model learning.

[0011] The above detection model learning unit generates a genuineness determination model for outputting a forgery probability using the audio embedding data as an input value through the second model learning, and the second model learning may be a genuineness determination model learning based on a fully connected layer.

[0012] The above detection processing unit reads the voice data input through the voice detection model to determine whether there is any modulation, and after reading the voice data based on a plurality of voice extensions and sample rates, performs preprocessing on the read voice data according to a preset format, outputs a modulation probability, and determines whether there is any modulation based on the modulation probability.

[0013] According to another embodiment, an artificial intelligence-based voice detection method is provided, which is performed by a voice detection server, comprising: collecting a training data set having at least one different format including original voice data and modulation data of the original voice data; performing preprocessing on the training data set; performing pre-training through first model learning and second model learning using the preprocessed training data set to extract speaker-specific information from input voice data and generate an artificial intelligence-based voice detection model for determining authenticity; and determining whether or not there is modulation based on a modulation probability of input voice data using the voice detection model.

[0014] The above learning data set may include data sets in different formats, including a first data set generated by mixing preset first voice synthesis techniques according to preset criteria, a second data set that is a voice deepfake data set, and a third data set generated based on the preset second voice synthesis technique.

[0015] In addition, a computer-readable recording medium recording a computer program for executing a method for implementing the disclosed embodiment may be further provided.

[0016] According to the disclosed embodiments, a voice detection model is created through learning processing using various formats of learning data sets including original voice data and modulation data of the original voice data, and by using the same to detect whether input voice data is modulated, it is expected that the performance of voice modulation detection can be improved.

[0017] Figures 1 and 2 are exemplary diagrams schematically illustrating a voice detection method according to one embodiment.

[0018] Figure 3 is a block diagram showing the configuration of a voice detection server according to one embodiment.

[0019] Figures 4 to 7 are exemplary diagrams for explaining a voice detection method according to one embodiment.

[0020] Figure 8 is a flowchart for explaining a voice detection method according to one embodiment.

[0021] FIG. 9 is a block diagram illustrating a computing environment including a computing device according to one embodiment.

[0022] Hereinafter, specific embodiments of the present invention will be described with reference to the drawings. The following detailed description is provided to facilitate a comprehensive understanding of the methods, devices, and / or systems described herein. However, these are merely examples and the present invention is not limited thereto.

[0023] In describing embodiments of the present invention, if a detailed description of a known technology related to the present invention is judged to unnecessarily obscure the gist of the present invention, the detailed description will be omitted. In addition, the terms described below are terms defined in consideration of their functions in the present invention, and this may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification. The terminology used in the detailed description is only for the purpose of describing embodiments of the present invention and should not be limited in any way. Unless clearly used otherwise, the singular form includes the plural form. In this description, expressions such as "comprises" or "having" are intended to indicate certain features, numbers, steps, operations, elements, parts or combinations thereof, and should not be construed to exclude the presence or possibility of one or more other features, numbers, steps, operations, elements, parts or combinations thereof other than those described.

[0024] FIG. 1 and FIG. 2 are exemplary diagrams for schematically explaining a voice detection method according to one embodiment, and FIG. 3 is a block diagram showing the configuration of a voice detection server according to one embodiment.

[0025] Hereinafter, a description will be given with reference to FIGS. 4 to 7, which are exemplary diagrams for explaining a voice detection method according to one embodiment.

[0026] Referring to FIG. 1, a voice detection method according to one embodiment may include a model learning step for generating an artificial intelligence-based voice detection model for determining whether input voice data is tampered with, and a service step for determining whether input voice data is tampered with using the generated voice detection model.

[0027] Specifically, the model training step can generate a voice detection model by learning through the collection of a training data set, preprocessing voice data from the training data set, and training a voice detection model. The service stage can perform data inference and data post-processing procedures based on the voice detection model trained in the model training step. A detailed description of these steps will be provided below.

[0028] Referring to FIG. 2, a voice detection server (100) according to one embodiment can extract voice data transmitted when a call connection is established, perform preprocessing on the extracted voice data, and then detect whether voice is altered using a voice detection model. At this time, if the result of detecting whether voice data is altered is an original voice that has not been altered (Real: original voice in FIG. 2), the voice detection server (100) can continue the call without any additional processing, and if it is an altered voice (Fake: deep voice) that has been altered, the voice detection server can output a warning sound notifying that an altered voice is included and then terminate the call.

[0029] In Fig. 2, the disclosed voice detection technology is described as an example of being applied when a call is connected, but it is not limited thereto and can be applied to any case where it is necessary to determine whether the original voice has been altered.

[0030] Referring to FIG. 3, the voice detection server (100) includes a detection model learning unit (110) and a detection processing unit (130).

[0031] The components illustrated in FIG. 3 are not essential for implementing a voice detection server (100) according to the present disclosure, and thus, the voice detection server (100) described in this specification may have more or fewer components than the components listed above.

[0032] The components illustrated in FIG. 3 may be communicatively connected to one another via a communications network (not shown). In some embodiments, the communications network may include the Internet, one or more local area networks, wire area networks, a cellular network, a mobile network, other types of networks, or a combination of these networks.

[0033] Referring to FIG. 3, the detection model learning unit (110) can collect a training data set composed of at least one different format, including original voice data and modulation data of the original voice data. In the present embodiment, the configuration for collecting the training data set is disclosed as the detection model learning unit (110), but is not limited thereto and may be implemented as a separate configuration (e.g., a data collection unit) from the detection model learning unit (110).

[0034] The above learning data set may include data sets in different formats, including a first data set generated by mixing preset first voice synthesis technologies according to preset criteria, a second data set which is a voice deepfake data set, and a third data set generated based on the preset second voice synthesis technology. For example, the first data set may be the AsvSpoof 2019 data set, and the second data set may be the Fraunhofer data set. The third data set may include modulated data generated by modulating original voice data according to criteria preset by an operator. For example, the third data set may be a data set including modulated data (e.g., I am a boy) in the form of modulated data in which only the voice is modulated for the same sentence as the original voice data (e.g., I am a boy). In this case, the third data set may reflect voice characteristics that may occur for each person when modulating the voice (e.g., tone, interjection, breathing interval during pronunciation, habits, etc.). Additionally, the third data set may include data sets collected from various social media such as YouTube and voice benchmark data.

[0035] The above-described training data set can be comprised primarily of data sets that can help improve generalization detection performance, and the included training data sets are not limited to the cases described above. The above-described generalization detection performance in a voice detection model can refer to the detection model's ability to detect forgery or alteration of voice data for previously unseen individuals and languages, beyond its discriminatory accuracy for the training data set.

[0036] The criteria and reasons for configuring the learning data set according to this embodiment may be as follows. The description of the learning data set described below may be applied to the first to third data sets described above.

[0037] For example, when selecting a data set for learning, the detection model learning unit (110) can collect a data set for learning in which original voice data and altered data of the original voice data exist simultaneously, and since the voice detection model is a model that applies learning of subtle differences between the original and altered data, the original and altered data for a specific person can exist simultaneously so that learning can be performed.

[0038] As another example, when selecting a training data set, the detection model learning unit (110) may select a training data set comprised of multiple speakers. The voice detection model applied in this embodiment aims to improve performance for previously unseen speakers. Accordingly, the detection model learning unit (110) may enable learning to be performed based on data reflecting the voices of various people during the training stage of the voice detection model.

[0039] As another example, when selecting a training data set, the detection model learning unit (110) may select a training data set that includes continuous training data that satisfies a preset time (e.g., more than a preset time) including original voice data and modulated data. The more training data sets there are, the more data there is to compare between the original and modulated data from the voice detection model's perspective, which may improve generalization performance. The third training data set may use data that satisfies a length greater than the average length per voice, so as to provide sufficient training data.

[0040] The first dataset mentioned above is the AsvSpoof 2019 dataset, which may be the official dataset for the voice deepfake detection competition hosted by InterSpeech, a speech synthesis society. As it is data generated by a highly credible speech synthesis society, it is likely a well-organized dataset with a high level of understanding of the creation of falsified datasets. This dataset was created using 19 different speech synthesis techniques, combining Text-to-Speech (TTS) and Vocoder, allowing the voice detection model to learn a wide range of falsification methods by applying various synthesis techniques.

[0041] The second dataset, Fraunhofer's "in the wild" dataset, is a deepfake voice dataset collected by the German speech research institute. This dataset may consist of famous politicians and celebrities. Considering that many voice phishing attacks are based on celebrities, this example applies a training dataset containing the voices of celebrities as input for training a voice detection model to prepare for such cases. This is expected to further improve the detection performance of celebrities against deepfake attacks.

[0042] The third dataset may include modulated data generated by modulating original voice data according to voice synthesis criteria preset by the operator. The dataset may be comprised of voice data from multiple speakers (e.g., 30 Korean speakers) and may be a dataset in which the original and modulated sentences perfectly match. Considering that the first and second datasets are datasets comprised of a first language (e.g., English), the third dataset may be comprised of a second language (e.g., Korean). Therefore, the present embodiment is expected to improve language detection performance for both the first and second languages. Furthermore, the third dataset may be comprised of voices that are identical or similar between the original voice data and the modulated data, enabling the voice detection model to learn subtle differences in forgery and modulation during training. In this case, the fact that the voices are identical or similar between the original voice data and the modulated data may mean that the voices are identical within a margin of error. The margin of error may be set by the operator. That is, the detection model learning unit (110) of the present embodiment performs learning so that even relatively minute differences in voice modulation can be determined by using a third data set in which sentences and voices between the original voice data and the modulated data are identical during learning for generating a voice detection model.

[0043] The detection model learning unit (110) can perform data preprocessing, including data selection and voice length adjustment, on the collected learning data set.

[0044] For example, the detection model learning unit (110) can perform a preprocessing operation according to data selection conditions that distinguish between original voice data and altered data from each voice file of the learning data set, and delete a voice file that includes at least one of the following: if the original voice data corresponding to the speaker's voice is missing from each voice file, if the quality of the voice file is below a standard value, and if the noise included in the voice file is above a standard value, from the learning data set.

[0045] That is, the detection model learning unit (110) can perform a procedure to check whether the speaker's voice exists in each voice file of the training data set and to check for voice noise that is irrelevant to the determination. The detection model learning unit (110) can first delete the data if the training data set contains a voice of a person other than the speaker or only contains noise, as this may cause confusion in the voice detection model. Fig. 4 can show an example of voices deleted during data selection.

[0046] As another example, the detection model learning unit (110) may perform a preprocessing task of adjusting the length of each voice file of the learning data set to a preset length.

[0047] The detection model learning unit (110) can perform a voice length adjustment step of cutting or expanding (padding) each voice file included in the collected training data set to a certain length before training. When the voice detection model receives voice audio of variable length in operation, it may not be able to detect the presence or absence of modulation, or the reliability of the result of detecting the presence or absence of modulation may be low. Accordingly, the detection model learning unit (110) can adjust the different lengths of the voice files (audios) of the collected training data set. At this time, the detection model learning unit (110) can set in advance a length that can best reflect the characteristics of different voice data of different lengths so that the result of detecting the presence or absence of modulation can be better than a standard value, and can adjust the length of the voice files (audios) according to the set length.

[0048] For example, the detection model learning unit (110) can adjust the length of the audio data (audio) to be used for learning to 80,000 length (e.g., audio data of 5 seconds in length). That is, audio data for learning can be extended by adding silence (an inaudible sound made up of 0) to the end of audio shorter than 5 seconds based on 5 seconds, and audio files longer than 5 seconds can be cut to fit 5 seconds and created as multiple pieces. The detection model learning unit (110) can adjust audio of 45,000 length (approximately 2.8 seconds) to 80,000 length by adding silence (an inaudible sound consisting of 0) of 35,000 length (approximately 2.2 seconds), or can cut the signal of the latter 11,000 length (approximately 0.7 seconds) from audio of 91,000 length (approximately 5.7 seconds) while leaving only the audio corresponding to the first 80,000 length (5 seconds). At this time, the section cut by the detection model learning unit (110) is not limited to the above-described case, and can be changed to the middle, the latter, etc. of the entire voice file (audio) section according to the needs of the operator. Fig. 5 may show an example of a voice file before and after the length has been adjusted.

[0049] The various preprocessing tasks described above (e.g., preprocessing tasks for data selection, preprocessing tasks for length adjustment, etc.) may be applied alone or in combination of at least two or more during the preprocessing of the learning data set to be disclosed.

[0050] The deepfake data set disclosed in this embodiment consists of three folders: training (train), verification (dev), and test (eval). Training (train) may refer to a data set used for model training, verification (dev) may refer to a data set used for model weight selection, and test (eval) may refer to a data set used for final inspection.

[0051] The first dataset, the ASVSpoof2019 dataset, is structured in the format described above, but the data within each folder is relatively uneven. The first dataset can be restructured by integrating the data within each folder while maintaining the folder structure for training (train), validation (dev), and testing (eval).

[0052] For example, the training and validation (dev) data sets before the change include 20 speakers and 6 synthesis methods, and the test (eval) folder includes 68 speakers and 13 synthesis methods. The detection model learning unit (110) of the present embodiment can integrate the data of the three folders and then rearrange them so that the 19 synthesis methods are evenly distributed to each folder by dividing the number of speakers in the ratio of training (train): validation (dev): test (eval) = 8:1:1. In the present embodiment, because the data distribution used for training (train) becomes relatively more diverse due to this data rearrangement, detection model learning with improved generalization effect can be expected.

[0053] In addition, in the case of the first data set, the data set is composed of a logical access data set that generates modulated data using a voice synthesis method such as TTS or Vocoder, and a physical access data set that is a forged data set created by imitating the original voice using a physical recording device. The detection model learning unit (110) may exclude the physical access data set in advance, considering that the detection model of the present embodiment is a model that detects voice.

[0054] In addition, because the first data network may include files with low sound quality and muffled pronunciation, the detection model learning unit (110) may exclude files below the standard by conducting data inspection based on pronunciation and sound quality.

[0055] In the case of the second data set, Fraunhofer's in the wild dataset, the original composition of the data set was a single folder named release_in_the_wild in which the original and altered data sets for each speaker were mixed without distinction, but the detection model learning unit (110) can additionally perform preprocessing work to unify the data set composition with the first data set (ASVSpoof2019) into a learning (train), verification (dev), and test (eval) structure.

[0056] First, the detection model learning unit (110) can classify original alterations by first creating an alteration original folder within a speaker folder using metadata (data information about data) provided by Fraunhofer to divide the original alteration data set by speaker.

[0057] Next, the detection model learning unit (110) can proceed with the task of dividing the speakers classified in the first stage into three folders: training (train), verification (dev), and test (eval) by adjusting the gender ratio evenly. The second data set includes short audio files shorter than a preset time (e.g., less than 2 seconds) and files with severe background noise, so that data selection can be performed based on background noise and audio length.

[0058] In the case of the third data set, the data of the third data set can be generated by dividing the original modulation folder by speaker at the time of creation, and performing a preprocessing task of dividing them into training, verification (dev), and testing (eval). The detection model learning unit (110) considers that the third data set is difficult data consisting of pairs of the original script content and the synthesized content, although the background noise or noise is relatively mild, and if the speech speed or specific pronunciation of the synthesized modulation original is abnormal, it can be deleted from the data set if an error exceeds the standard when compared with the original. Since the present embodiment includes only modulation data that is as similar as possible to the original in learning through this task, it can be expected to have the effect of improving the learning performance of the voice detection model.

[0059] The detection model learning unit (110) can perform pre-learning using a learning data set composed of at least one different format including original voice data and modulation data of the original voice data to create an artificial intelligence-based voice detection model.

[0060] The detection model learning unit (110) can perform pre-learning through first model learning and second model learning using the preprocessed learning data set to extract speaker-specific information from the input voice data and generate a voice detection model for determining authenticity. The first model learning may be convolution (1D CONV)-based audio embedding model learning, and the second model learning may be fully connected layer-based authenticity determination model learning.

[0061] The speaker-specific information described above may refer to features that enable speaker recognition from speech data. For example, speaker-specific information may include the speaker's vocal characteristics (e.g., tone, interjections, breathing intervals during speech, habits, etc.). Extracting speaker-specific information from speech data may refer to extracting audio embeddings from the first model during training. At this time, speaker-specific information may be any feature that can distinguish between the original and the altered characteristics.

[0062] In the case of the first data set (ASVSpoof2019 data set), the detection model learning unit (110) can extract speaker-specific information by considering the characteristics of the original modulation when the data sound quality is below the standard or the pronunciation is slurred. That is, in the case of the first data set (ASVSpoof2019 data set), the detection model learning unit (110) can extract speaker-specific information by considering the data sound quality and speech characteristics (particularly, pronunciation). Although the first data set excludes low sound quality below the standard and slurred pronunciation during preprocessing, it may be considered that unique data characteristics (e.g., low sound quality, slurred pronunciation, etc.) arising from the synthesis method and the original data collection method may still exist in the data.

[0063] In the case of the second data set (In the wild data set), the detection model learning unit (110) can extract speaker information during learning by focusing on noise or intonation unique to the synthesis. That is, in the case of the second data set (In the wild data set), the detection model learning unit (110) can extract speaker-specific information by considering noise and speech characteristics (particularly, intonation). In the case of the second data set, cases where background noise exceeds the standard level during preprocessing were excluded, but background noise that may remain throughout the data due to the characteristics of the data may be taken into consideration.

[0064] For the third data set, the detection model learning unit (110) can extract voice information that distinguishes between the original and the altered sound by taking hints from various vocalization characteristics, such as differences in subtle breathing intervals between the synthesized and original sounds, and differences in pitch between the beginning and end of speech. That is, for the third data set, the detection model learning unit (110) can extract speaker-specific information by considering vocalization characteristics.

[0065] Additionally, the detection model learning unit (110) can learn a first model to distinguish between features of data appearing in alteration and features appearing in the original, based on various features that humans cannot distinguish audibly as well as the aforementioned distinction points from the first to third data sets described above. The detection model learning unit (110) can transfer information extracted through the first model learning to a second model, which is a discrimination model.

[0066] The detection model learning unit (110) can extract audio-specific information (audio embedding) by performing parameter learning of a band pass filter on the voice data (audio) of the preprocessed learning data set in a 1d CNN-based deep learning model called SincNet.

[0067] Specifically, the detection model learning unit (110) can be achieved by merging two models: convolution (1D CONV)-based audio embedding model learning and fully connected layer-based authenticity discrimination model learning. The detection model learning unit (110) extracts speaker-specific information from voice data (audio information) input as audio embedding, and transmits the output value as an input value to the authenticity discrimination model of the corresponding information, and can ultimately learn to determine whether the corresponding voice is original or altered.

[0068] The detection model learning unit (110) can generate an audio embedding model for extracting audio embedding data representing original and modulation characteristics to distinguish original and modulated audio information from the training data set through first model learning. At this time, the audio embedding data can be applied as an input value during second model learning.

[0069] An audio embedding model may be a model that extracts data that well represents the characteristics of the original and altered input voice data (audio information). At this time, audio embedding may mean interpreting voice features that can distinguish the original and altered voice data. At this time, the detection model learning unit (110) may learn a deep learning model to best extract the original and altered voice features that can maximize the performance of the voice detection model in the audio embedding learning step, but is not limited thereto.

[0070] Figure 6 is a diagram showing an example of an audio embedding extraction structure.

[0071] Referring to FIG. 6, from a deep learning structural perspective, the detection model learning unit (110) inputs the training data set based on the embedding model into a filtering structure formed by one-dimensional convolution (1D CONV) to perform filtering, and can further refine the filtered training data set through two residual block (Resblock) layers. Thereafter, the detection model learning unit (110) can input the refined training data set to a neural network formed by a Gated Recurrent Unit (GRU) and a fully connected layer to interpret the data divided into voice frame units as one voice data considering continuity and output the final audio embedding result value. The voice audio embedding extraction method using the deep learning model applied in the present disclosure can be expected to have the effect of extracting information suitable for the characteristics of voice data by overcoming the limitations of methods such as general Mel-Frequency Cepstral Coefficient.

[0072] The detection model learning unit (110) can create a genuineness determination model for outputting the probability of forgery using audio embedding data as an input value through second model learning.

[0073] Figure 7 shows an example of a full-connection-based truth-discrimination model structure.

[0074] The authenticity discrimination model may refer to a discrimination model composed of a fully connected layer. Referring to Fig. 7, the detection model learning unit (110) may apply the audio embedding data extracted in the previous step as an input value to the authenticity discrimination model composed of a single fully connected layer, and output the probability of forgery as a result value. Compared to a discrimination model with added synthetic holes, the authenticity discrimination model can be trained and read relatively quickly by optimizing the parameter values ​​of the model used.

[0075] The present disclosure generates a voice detection model that includes an audio embedding model and a true-false discrimination model. However, by placing greater emphasis on the audio embedding model to efficiently detect the characteristics of the original modulation, the model can be implemented with a structure capable of achieving high accuracy with relatively few parameters. In other words, the voice detection model of the present disclosure can provide high accuracy and fast discrimination performance when applied to a service.

[0076] The detection processing unit (130) can determine whether or not there is modulation based on the modulation probability of the input voice data using a voice detection model.

[0077] The detection processing unit (130) reads voice data input through a voice detection model to determine whether there is any modulation, and after reading voice data based on multiple voice extensions and sample rates, performs preprocessing on the read voice data according to a preset format, outputs a modulation probability, and can determine whether there is any modulation based on the modulation probability.

[0078] In the disclosed embodiment, referring to FIG. 1, the service step of determining whether voice data is tampered with may mean a task of optimizing a pre-learned voice detection model and data processing code and converting them into a service code.

[0079] The detection processing unit (130) can read the input voice data using a voice detection model and output as a final value whether the voice data is original or altered.

[0080] When reading voice data, the detection processing unit (130) can read data of various voice extensions (e.g., wav, flac) and sample rates (e.g., 16000 hz, 48000 hz) taking compatibility into consideration.

[0081] The detection processing unit (130) can, when preprocessing voice data, convert stereo (2-channel) voice into mono (1-channel) voice and perform length adjustment. Thereafter, the detection processing unit (130) can input preprocessed voice data into a pre-trained voice detection model to obtain a modulation probability as an output value.

[0082] The detection processing unit (130) can store the final result of voice detection in the form of a file (e.g., JSON format). Specifically, the detection processing unit (130) can store the result of voice detection matched with the name of the voice file and the modulation probability (modulation (counterfeiting) probability value between 0 and 100) in the form of a file. At this time, the closer the modulation probability is to 0, the higher the probability that it is a modulated voice (fake in FIG. 2) (interpreted as modulation), and the closer it is to 100, the higher the probability that it is an original voice (real in FIG. 2) (interpreted as original). Through this process, the detection processing unit (130) can process the provision of the final detection result to the user by building a detection record database and transmitting information to the front server.

[0083] The detection processing unit (130) can process the final results of voice detection and provide them to the operator for confirmation. For example, the detection processing unit (130) can calculate the average of all final results of voice detection, calculate the average of the final results of voice detection for each voice file, or generate a voice detection report that distinguishes and displays the original and modified sections for each voice file, and provide the report to the operator terminal. At this time, the voice detection report can include forms such as graphs, diagrams, and narrative explanations, depending on the operator's needs.

[0084] Although not shown, a voice detection server (100) according to one embodiment includes a processor (110). At this time, the processor may be composed of one or more cores, and may include a processor for data analysis and deep learning, such as a central processing unit, a general purpose graphics processing unit, and a tensor processing unit of a computing device. The processor may be the processor (14) of FIG. 9, and may read a computer program stored in a computer-readable storage medium (16) to perform data processing for machine learning according to the present disclosure. According to the present embodiment, the processor may perform operations for learning a neural network. The processor may perform calculations for learning a neural network, such as processing input data for learning in deep learning, extracting features from input data, calculating errors, and updating weights of a neural network using backpropagation.

[0085] The above neural network model may be a deep neural network. In the present disclosure, neural network, network function, and neural network may be used interchangeably. A deep neural network (DNN) may refer to a neural network that includes multiple hidden layers in addition to an input layer and an output layer. Using a deep neural network, it is possible to identify latent structures of data. That is, it is possible to identify latent structures of photos, text, videos, voices, and music (e.g., what objects are in the photo, what the content and emotion of the text are, what the content and emotion of the voice are, etc.). A deep neural network may include a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a Q network, a U network, a Siamese network, etc.

[0086] Convolutional neural networks (CNNs) are a type of deep neural network that include neural networks containing convolutional layers. CNNs are a type of multilayer perceptron designed to use minimal preprocessing. CNNs can be composed of one or more convolutional layers and artificial neural network layers combined with them. CNNs can additionally utilize weight and pooling layers. This structure allows CNNs to fully utilize two-dimensional input data. CNNs can be used to recognize objects in images. CNNs can process image data by representing it as a matrix with dimensions. For example, in the case of image data encoded in RGB (red-green-blue), each of the R, G, and B colors can be represented as a two-dimensional (for example, in a two-dimensional image) matrix. That is, the color value of each pixel of the image data can be an element of a matrix, and the size of the matrix can be the same as the size of the image. Therefore, the image data can be represented as three two-dimensional matrices (a three-dimensional data array).

[0087] In a convolutional neural network, a convolutional process (input and output of a convolutional layer) can be performed by moving the convolutional filter and multiplying the matrix elements at each location of the image with the convolutional filter. The convolutional filter can be composed of an n*n matrix. The convolutional filter can generally be composed of a fixed-shape filter that is smaller than the total number of pixels in the image. That is, when an m*m image is input to a convolutional layer (for example, a convolutional layer whose convolutional filter has a size of n*n), a matrix representing n*n pixels containing each pixel of the image can be component-wise multiplied with the convolutional filter (i.e., each element of the matrix is ​​multiplied). By multiplying with the convolutional filter, a component matching the convolutional filter can be extracted from the image. For example, a 3*3 convolutional filter for extracting up and down straight line components from an image can be configured as [[0,1,0], [0,1,0], [0,1,0]]. When a 3*3 convolutional filter for extracting up and down straight line components from an image is applied to an input image, up and down straight line components matching the convolutional filter from the image can be extracted and output. A convolutional layer can apply a convolutional filter to each matrix for each channel representing an image (i.e., R, G, B colors in the case of an R, G, B coded image). A convolutional layer can extract features matching the convolutional filter from the input image by applying a convolutional filter to the input image. The filter value of the convolutional filter (i.e., the value of each element of the matrix) can be updated by backpropagation during the learning process of a convolutional neural network.

[0088] A subsampling layer can be connected to the output of a convolutional layer to simplify the output of the convolutional layer, thereby reducing memory usage and computational complexity. For example, when the output of a convolutional layer is input to a pooling layer having a 2*2 max pooling filter, the image can be compressed by outputting the maximum value included in each patch for each 2*2 patch from each pixel of the image. The above-described pooling may be a method of outputting the minimum value in a patch or the average value of a patch, and any pooling method may be included in the present disclosure.

[0089] A convolutional neural network may include one or more convolutional layers and subsampling layers. A convolutional neural network can extract features from an image by repeatedly performing convolutional and subsampling processes (e.g., the aforementioned max pooling). Through repeated convolutional and subsampling processes, the neural network can extract global features of the image.

[0090] The output of a convolutional layer or a subsampling layer can be input to a fully connected layer. A fully connected layer is a layer in which all neurons in one layer are connected to all neurons in the neighboring layer. A fully connected layer can refer to a structure in a neural network in which all nodes in each layer are connected to all nodes in other layers.

[0091] At least one of the CPU, GPGPU, and TPU of the processor can process network function learning. For example, the CPU and GPGPU can jointly process network function learning and data classification using the network function. Furthermore, in one embodiment of the present disclosure, processors of multiple computing devices can be used together to process network function learning and data classification using the network function. Furthermore, a computer program executed on a computing device according to one embodiment of the present disclosure may be a CPU, GPGPU, or TPU executable program.

[0092] FIG. 8 is a flowchart illustrating a voice detection method according to one embodiment. The method illustrated in FIG. 8 may be performed, for example, by the aforementioned voice detection server (100). While the illustrated flowchart divides the method into multiple steps, at least some of the steps may be performed in a different order, combined with other steps and performed together, omitted, divided into substeps, or additionally performed with one or more steps not illustrated.

[0093] The voice detection method disclosed below can equally apply the role of the voice detection server (100) disclosed through the above-described FIGS. 1 to 7, and for the convenience of explanation, redundant descriptions will be omitted.

[0094] At step 1100, the voice detection server (100) can collect a training data set consisting of at least one different format including voice original data and modulation data of the voice original data.

[0095] The above learning data set may include data sets in different formats, including a data set generated by mixing preset first voice synthesis techniques according to preset criteria, a voice deepfake data set, and a data set generated based on a preset second voice synthesis technique.

[0096] At step 1200, the voice detection server (100) can perform preprocessing of the learning data set.

[0097] At step 1300, the voice detection server (100) can perform pre-learning through first model learning and second model learning using the preprocessed learning data set to extract speaker-specific information from the input voice data and generate an artificial intelligence-based voice detection model for determining authenticity.

[0098] At step 1400, the voice detection server (100) can determine whether or not there is any modulation based on the modulation probability of the input voice data using a voice detection model.

[0099] FIG. 9 is a block diagram illustrating a computing environment including a computing device according to one embodiment. In the illustrated embodiment, each component may have different functions and capabilities other than those described below, and may include additional components other than those described below.

[0100] The illustrated computing environment (10) includes a computing device (12). The computing device (12) may be one or more components included in a voice detection server (100) according to one embodiment.

[0101] A computing device (12) includes at least one processor (14), a computer-readable storage medium (16), and a communication bus (18). The processor (14) may cause the computing device (12) to operate according to the exemplary embodiments mentioned above. For example, the processor (14) may execute one or more programs stored in the computer-readable storage medium (16). The one or more programs may include one or more computer-executable instructions, which, when executed by the processor (14), may be configured to cause the computing device (12) to perform operations according to the exemplary embodiments.

[0102] A computer-readable storage medium (16) is configured to store computer-executable instructions or program code, program data, and / or other suitable forms of information. A program (20) stored in the computer-readable storage medium (16) includes a set of instructions executable by the processor (14). In one embodiment, the computer-readable storage medium (16) may be a memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, any other form of storage medium that can be accessed by the computing device (12) and store desired information, or a suitable combination thereof.

[0103] A communication bus (18) interconnects various other components of the computing device (12), including the processor (14) and computer-readable storage media (16).

[0104] The computing device (12) may also include one or more input / output interfaces (22) that provide interfaces for one or more input / output devices (24) and one or more network communication interfaces (26). The input / output interfaces (22) and the network communication interfaces (26) are connected to the communication bus (18). The input / output devices (24) may be connected to other components of the computing device (12) via the input / output interfaces (22). Exemplary input / output devices (24) may include input devices such as pointing devices (such as a mouse or a trackpad), a keyboard, a touch input device (such as a touchpad or a touchscreen), a voice or sound input device, various types of sensor devices and / or photographing devices, and / or output devices such as display devices, printers, speakers and / or network cards. The exemplary input / output devices (24) may be included within the computing device (12) as a component constituting the computing device (12), or may be connected to the computing device (12) as a separate device distinct from the computing device (12).

[0105] The disclosed embodiments may be implemented in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, and when executed by a processor, may generate program modules to perform the operations of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.

[0106] While representative embodiments of the present invention have been described in detail above, those skilled in the art will appreciate that various modifications to the above-described embodiments are possible without departing from the scope of the present invention. Therefore, the scope of the present invention should not be limited to the described embodiments, but should be determined not only by the claims set forth below but also by equivalents thereof.

Claims

1. A detection model learning unit that generates an artificial intelligence-based voice detection model by performing pre-learning using a learning data set consisting of at least one different format including original voice data and modulation data of the original voice data, and extracts speaker-specific information from input voice data and generates the voice detection model for determining authenticity by performing pre-learning through first model learning and second model learning using the learning data set on which preprocessing has been performed; and An artificial intelligence-based voice detection server, comprising a detection processing unit that determines whether or not voice data is tampered with based on the probability of tampering using the voice detection model.

2. In claim 1, An artificial intelligence-based voice detection server, wherein the above learning data set comprises data sets in different formats: a first data set generated by mixing preset first voice synthesis techniques according to preset criteria, a second data set which is a voice deepfake data set, and a third data set generated based on the preset second voice synthesis technique.

3. In claim 1, The above detection model learning unit, An artificial intelligence-based voice detection server that performs a preprocessing task according to data selection conditions for distinguishing original voice data and modulated data from each voice file of the above learning data set, and deleting a voice file that includes at least one of the following: if the original voice data corresponding to the speaker's voice is missing from each voice file, if the quality of the voice file is below a standard value, and if the noise included in the voice file is above a standard value, from the above learning data set.

4. In claim 3, The above detection model learning unit, An artificial intelligence-based voice detection server that performs a preprocessing task of adjusting the length of each voice file in the above learning data set to a preset length.

5. In claim 1, The above detection model learning unit, Through the above first model learning, an audio embedding model is created to extract audio embedding data representing original characteristics and modulation characteristics to distinguish the original and modulation from the audio information of the above learning data set, The above first model learning is an artificial intelligence-based voice detection server that is a convolution (1D CONV)-based audio embedding model learning.

6. In claim 5, The above detection model learning unit, Through the above second model learning, a genuineness determination model is created to output the forgery probability using the audio embedding data as an input value, The above second model learning is an artificial intelligence-based voice detection server that is a truth-discrimination model learning based on a fully connected layer.

7. In claim 1, The above detection processing unit, An artificial intelligence-based voice detection server that reads the voice data input through the voice detection model to determine whether there is any modulation, reads the voice data based on a plurality of voice extensions and sample rates, performs preprocessing on the read voice data according to a preset format, outputs a modulation probability, and determines whether there is any modulation based on the modulation probability.

8. A method performed by a voice detection server, A step of collecting a training data set comprising at least one different format including original voice data and modulation data of the original voice data; A step of performing preprocessing of the above learning data set; A step of performing pre-learning through first model learning and second model learning using the above preprocessed learning data set to extract speaker-specific information from input voice data and create an artificial intelligence-based voice detection model for determining authenticity; and An artificial intelligence-based voice detection method, comprising a step of determining whether or not voice data is modulated based on a modulation probability of the input voice data using the above voice detection model.

9. In claim 8, The above learning data set is an artificial intelligence-based voice detection method, which includes data sets in different formats: a first data set generated by mixing preset first voice synthesis techniques according to preset criteria, a second data set which is a voice deepfake data set, and a third data set generated based on the preset second voice synthesis technique.

10. A computer-readable recording medium having recorded thereon a program for executing the method of claim 8.

Citation Information

Patent Citations

  • Voice synthesis learning device, method, and program

    JP2018036413A

  • Fire proof clothes having fire response system

    KR1020250062196A

  • Plant cultivation equipment

    KR102226925B1

  • Sentence writing system

    KR102370729B1

  • System for providing adjacent building pre-survey service usign 360 degree virtual reality camera

    KR102388777B1