Voice file false alteration detection device and method

By amplifying the transition zone in the voice file and analyzing the blowing phenomenon, the pseudo-mutation problem of the voice file after mixed pasting and compression in the prior art is solved, and effective detection of pseudo-mutation and provision of objective evidence is achieved.

CN120164488APending Publication Date: 2025-06-17SSMM INC +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311730185.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing methods of pseudo-variant detection of voice files have limitations in judging small differences, especially when it is difficult to effectively detect pseudo-variant after mixing and pasting and compression operations.

Method used

By amplifying the transition zone in the voice file recorded by the user terminal, using the blowing phenomenon generated in the amplified transition zone, it is determined whether the air sound or the gingival palate friction sound disappears, or whether the blowing phenomenon occurs in other sounds, thereby detecting the pseudo-change of the voice file.

Benefits of technology

Even in a voice file that is mixed and pasted and compressed, pseudo-change can be effectively detected, and since edited clips of the voice file can be objectively detected, it can be used as objective evidence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164488A_ABST
    Figure CN120164488A_ABST
Patent Text Reader

Abstract

The invention relates to a device and a method for detecting false alteration of a voice file using specific pronunciation. The apparatus for detecting false alteration of a voice file using a specific pronunciation according to the present invention comprises: a receiving unit for receiving a voice file recorded by a user terminal; a pre-processing unit that outputs the received audio file in the form of a graph composed of a time axis and a frequency axis, and independently amplifies a specific transition zone in the output graph; a signal detection unit that extracts a specific signal generated due to a blowing phenomenon in the specific transition zone, and extracts a voice waveform corresponding to a vocal sound or a gingival palate friction sound; the false alteration judgment part is used for analyzing the correlation between the specific signal and the voice waveform and judging whether the voice file is subjected to false alteration or not according to an analysis result; and a control unit that, if it is determined that the audio file has been pseudo-altered, displays a pseudo-altered segment and outputs the pseudo-altered segment to a screen. As described above, according to the present invention, a transition zone is independently amplified from a voice file output in a graphic form, and whether or not false alteration is performed is determined using a blowing phenomenon occurring in the amplified transition zone, so that false alteration detection can be performed even for a voice file that is subjected to mixed pasting and compression, and the false alteration detection accuracy is improved. And the editing segments of the voice file can be objectively detected, so that the editing segments can be used as objective evidence data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a voice file pseudo-alteration detection device and method, and more specifically, to a voice file pseudo-alteration detection device and method for amplifying a transition band in a voice file recorded by a user terminal, and using the amplified transition band to determine whether the blowing phenomenon disappears in a breathed sound or an alveolar-palatal fricative sound, or whether the blowing phenomenon occurs in sounds other than breathed sounds or alveolar-palatal fricative sounds to detect the pseudo-alteration of the voice file. Background Art

[0002] Due to the nature of digital voice files, sophisticated forgeries can be made through large-scale copying of the original files and digital editing programs, so actual cases of manipulation of digital voice files are increasing.

[0003] In addition, recently, various sound editors with sound visualization functions such as spectrogram have been provided. For example, Photoshop can select, cut, copy, and mix and paste syllables, units, and other fragments of digital voice files with just a few clicks, making it possible for not only experts but also ordinary people to edit sounds easily and accurately.

[0004] In the sound editing method, if the digital voice files to be edited are not recorded in the same environment, compression is performed in order to eliminate noise and provide the same sense of space and clean sound after mixing and pasting.

[0005] The hybrid paste includes a function that automatically reduces breathing sounds, breathy sounds or alveolar palatal fricative sounds, and also reduces ambient noise, thereby reducing unnatural empathy. In particular, the compressor is an effector of the mode envelope that processes the volume changes inherent in the sound. When this technology is applied, according to the ADSR (Attack, Decay, Sustain, Release) principle, a frequency shift phenomenon will occur in the attack of the front end of the compressor application.

[0006] On the other hand, when mixing, pasting and compressors are used, subtle differences appear in the spectrogram. However, existing artefact detection methods have limitations in determining whether small differences occur naturally or are the result of human editing.

[0007] Specifically, in the past, the forgery of voice files has been detected by using the Electrical Network Frequency (ENF). However, when using mixed paste editing for audio, synthesis can be performed without touching the audible frequency band. In addition, when there is an overlap of voices between speakers, it is usually judged that there is no editing, but this can be used to create an overlapping voice between speakers.

[0008] In particular, in the call recordings of recent user terminals, low-frequency band signals are removed through preprocessing, so there is a problem that it is almost impossible to detect the ENF signal from the actually recorded digital voice files.

[0009] In addition, since the automatic detection method for forged segments of audio signals only considers the characteristics of background noise caused by voice insertion, it is difficult to apply it to voice files in which voice segments are deleted or mixed and pasted. Therefore, there are limitations in judging whether forgery has occurred.

[0010] The background art of the present invention has been disclosed in Korean Patent No. 10-1382356 (announced on January 10, 2014). Summary of the Invention

[0011] According to the present invention as described above, an object is to provide a voice file forgery detection device and method that amplify a transition band in a voice file recorded by a user terminal, and use the amplified transition band to judge whether the blowing phenomenon disappears in aspirated sounds or alveopalatal fricatives, or whether a blowing phenomenon occurs in sounds other than aspirated sounds or alveopalatal fricatives to detect forgery of the voice file.

[0012] To solve the above technical problems, according to an embodiment of the present invention, in a voice file forgery detection device using specific pronunciations, it includes: a receiving unit for receiving a voice file recorded by a user terminal; a preprocessing unit for outputting the received voice file in the form of a graph composed of a time axis and a frequency axis, and independently amplifying a specific transition band in the output graph; a signal detection unit for extracting a specific signal generated due to the blowing phenomenon in the specific transition band, and extracting a voice waveform corresponding to an aspirated sound or an alveopalatal fricative; a forgery judgment unit for analyzing the correlation between the specific signal and the voice waveform, and judging whether the voice file is forged according to the analysis result; and a control unit for, if it is judged that the voice file is forged, displaying the forged segment and outputting it to the screen.

[0013] The forgery judgment unit judges whether forgery has occurred by judging whether a specific signal is detected in a segment matching the voice waveform corresponding to the aspirated sound or the alveopalatal fricative.

[0014] In the pseudo-forgery determination unit, if no specific signal is detected, the spectrum corresponding to the segment occurring in the aspirated sound or alveolar palatal fricative included in the same voice file or another voice file is mixed and pasted into the corresponding segment, thereby determining that it has been pseudo-forged.

[0015] In the pseudo-forgery determination unit, if a specific signal is detected but the spectrum of the specific signal is irregularly output, the spectrum corresponding to the segment occurring in the aspirated sound or alveolar palatal fricative included in the same voice file or another voice file is mixed and pasted into the corresponding segment, and then a compressor is applied to determine that it has been pseudo-forged.

[0016] In the pseudo-forgery determination unit, when a specific signal is detected, the magnitude of the specific signal corresponding to the same aspirated sound or alveolar palatal fricative and the difference value of the specific signal are compared with a reference value. If the difference value is greater than the reference value, the spectrum corresponding to the segment occurring in the aspirated sound or alveolar palatal fricative uttered by others included in the same voice file or another voice file is mixed and pasted into the corresponding segment, thereby determining that it has been pseudo-forged.

[0017] In the pseudo-forgery determination unit, it is determined whether a specific signal is detected in a segment that matches the voice waveform corresponding to a sound other than the aspirated sound or alveolar palatal fricative. If a specific signal is detected, the spectrum corresponding to the segment where the sound other than the aspirated sound or alveolar palatal fricative occurs in the same voice file or another voice file is mixed and pasted into the corresponding segment, thereby determining that it has been pseudo-forged.

[0018] In addition, a pseudo-forgery detection method using a voice file pseudo-forgery detection device according to an embodiment of the present invention includes: a step of receiving a voice file recorded through a user terminal; a step of outputting the received voice file in the form of a graph composed of a time axis and a frequency axis, and independently magnifying a specific transition band in the output graph; a step of extracting a specific signal generated due to a blowing phenomenon in the specific transition band and extracting a voice waveform corresponding to an aspirated sound or alveolar palatal fricative; a step of analyzing the correlation between the specific signal and the voice waveform, and determining whether the voice file has been pseudo-forged according to the analysis result; and a step of displaying the segment where pseudo-forgery occurs and outputting it to the screen if it is determined that the voice file has been pseudo-forged.

[0019] Advantages of the Invention

[0020] According to the present invention, the transition band is independently enlarged from a voice file output in a graphical form, and the blowing phenomenon generated in the enlarged transition band is used to determine whether there is forgery. Therefore, even for a voice file that has been mixed and pasted and compressed, forgery detection can be performed. Moreover, since the edited segment of the voice file can be objectively detected, it can also be used as objective evidence.

[0021] In addition, according to an embodiment of the present invention, aspirated sounds or palato-alveolar fricatives vary from person to person according to the oral cavity structure or the amount of air exhaled from the lungs. Accordingly, the output of the blowing phenomenon generated also varies. It is possible to determine whether the voice of another person has been synthesized or whether the recording was made in another space, and it has the effect of reducing the physical time required to analyze the voice file and reducing the cost incurred due to the involvement of experts. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a structural diagram of a voice file forgery detection device according to an embodiment of the present invention.

[0023] Figure 2 It is a flowchart for explaining a forgery detection method using a voice file forgery detection device according to an embodiment of the present invention.

[0024] Figure 3 It shows in Figure 2 An example diagram of the state of graphically outputting a voice file graph in step S220 shown.

[0025] Figure 4 It is for explaining Figure 2 An example diagram of a specific signal detected in step 230 shown.

[0026] Figure 5 An example diagram showing the state of deleting a specific signal by mixing and pasting a voice file according to an embodiment of the present invention.

[0027] Figure 6 An example diagram showing the state of detecting an irregular specific signal using a compressor after mixing and pasting an aspirated sound or a palato-alveolar fricative in a voice file according to an embodiment of the present invention.

[0028] Figure 7 An example diagram showing the state of detecting an irregular specific signal using a compressor after mixing and pasting sounds other than aspirated sounds or palato-alveolar fricatives in a voice file according to an embodiment of the present invention.

[0029] Figure 8 An example diagram showing the state of mixing and pasting the voice of another person in a voice file according to an embodiment of the present invention.

[0030] Description of Reference Numerals:

[0031] 100: Voice file forgery detection device; 110: Receiving unit; 120: Preprocessing unit; 130: Signal detection unit; 140: Forgery judgment unit; 150: Control unit. Detailed implementation manner

[0032] Next, the preferred embodiments according to the present invention will be described in detail with reference to the accompanying drawings. In this process, for clarity and convenience of explanation, the thickness of the lines or the dimensions of the components shown in the drawings may be exaggerated.

[0033] In addition, the terms described below are terms defined in consideration of the functions in the present invention and may be changed according to the intention or habit of the user or operator. Therefore, these terms should be defined based on the content of the entire specification.

[0034] Next, the Figure 1 voice file forgery detection device according to an embodiment of the present invention will be described in more detail.

[0035] Figure 1 FIG. is a structural diagram for explaining the voice file forgery detection device according to an embodiment of the present invention.

[0036] As Figure 1 shown, the voice file forgery detection device 100 according to an embodiment of the present invention includes: a receiving unit 110, a preprocessing unit 120, a signal detection unit 130, a forgery judgment unit 140, and a control unit 150.

[0037] First, the receiving unit 110 receives a voice file for detecting forgery. At this time, the voice file shows a file recorded using a user terminal.

[0038] Then, the preprocessing unit 120 converts the received voice file into a graphical form and independently magnifies the transition band in the converted graph.

[0039] Specifically, the preprocessing unit 120 outputs the voice file in the form of a graph combining the characteristics of a waveform and a spectrum. At this time, the waveform represents the amplitude axis that changes according to the time axis, the spectrum represents the amplitude axis that changes according to the frequency axis, and the change of the amplitude axis is displayed by the difference in concentration or display color.

[0040] Then, the preprocessing unit 120 extracts the transition band from the output graph and magnifies the extracted transition band to output the spectrum in the transition band.

[0041] The signal detection unit 130 is configured to detect speech waveforms corresponding to aspirated sounds and alveolo-palatal fricatives. Additionally, the signal detection unit 130 is configured to detect specific signals generated by the blowing phenomenon in the amplified transition band.

[0042] Then, the forgery determination unit 140 uses the detected speech waveforms and specific signals to determine whether the speech file has been forged.

[0043] Specifically, aspirated sounds or alveolo-palatal fricatives generate a blowing phenomenon. Therefore, the forgery determination unit 140 determines whether a specific signal is generated in a segment that matches the speech waveform corresponding to an aspirated sound or an alveolo-palatal fricative. When a specific signal is generated, it is determined that no forgery has occurred in the corresponding speech file.

[0044] On the other hand, if no specific signal is generated in a segment that matches the speech waveform corresponding to an aspirated sound or an alveolo-palatal fricative, the forgery determination unit 140 determines that forgery has occurred in the corresponding speech file.

[0045] In addition, when a specific signal is generated in a segment that matches the speech waveform corresponding to a sound other than an aspirated sound or an alveolo-palatal fricative, the forgery determination unit 140 determines that forgery has occurred in the corresponding speech file.

[0046] Finally, the control unit 140 displays the forged segment and outputs it on the screen.

[0047] Next, Figures 2 to 8 More specifically, a forgery detection method using the speech file forgery detection device 100 according to an embodiment of the present invention will be described.

[0048] Figure 2 It is a flowchart for explaining a forgery detection method using a speech file forgery detection device according to an embodiment of the present invention.

[0049] As Figure 2 shown, the speech file forgery detection device 100 according to an embodiment of the present invention receives a speech file to be detected for forgery (S210).

[0050] Specifically, the receiving unit 110 receives a speech file recorded by a user terminal.

[0051] The purpose of the speech file forgery detection device 100 according to an embodiment of the present invention is to determine whether a corresponding speech file has been forged by using the blowing phenomenon that occurs when an aspirated sound or an alveolo-palatal fricative is made.

[0052] When recording sound using a user terminal, a low-pass filter or a band-pass filter is applied, and within the frequency range where the low-pass filter or the band-pass filter is applied, the blowing phenomenon corresponding to the production of aspirated sounds or palato-alveolar fricatives is recorded.

[0053] At this time, an aspirated sound, which is a sound produced with a strong puff of air when a plosive sound is made, is called an aspirated sound ("c", "k", "t", "p") in Korean phonology. A palato-alveolar fricative is a sound produced by the friction generated when air passes through a narrow gap when a consonant is made, and a blowing phenomenon occurs when "s" or "ss" appears before / i, j / .

[0054] In addition, the blowing phenomenon also occurs when making tense consonants such as "gg", "dd", "bb", "zz" according to the oral structure or characteristics of the speaker. Therefore, the voice file forgery detection device 100 according to an embodiment of the present invention can also use the blowing phenomenon generated by tense consonants other than aspirated sounds or palato-alveolar fricatives to determine whether a corresponding voice file has been forged according to user needs.

[0055] Therefore, in an embodiment of the present invention, a voice file recorded by a user terminal is used to analyze forgery.

[0056] Then, the preprocessing unit 120 outputs the received voice file in the form of a graph, and independently magnifies the transition band in the output graph (S220).

[0057] Specifically, the preprocessing unit 120 outputs the sound or waveform included in the voice file in the form of a graph for visualization.

[0058] Figure 3 For showing Figure 2 An example diagram of the state in which the voice file is output as a graph in step S220 shown.

[0059] As Figure 3 shown, the X-axis of the graph represents the time axis, and the Y-axis represents the frequency axis. The graph represents the amplitude difference generated according to the changes in the time axis and the frequency axis using concentration and color differences.

[0060] On the other hand, since the low-pass filter or the band-pass filter applied to the voice file cannot completely block the signal, a transition band is formed between the low-pass filter or the band-pass filter and the stop band, and the blowing phenomenon generated by aspirated sounds or palato-alveolar fricatives is recorded in the transition band.

[0061] However, the transition band is narrow, so there are limitations in distinguishing the specific signals generated by the blowing phenomenon.

[0062] Therefore, the preprocessing unit 120 independently amplifies a specific transition band.

[0063] Then, the signal detection unit 130 detects the speech waveform corresponding to aspirated sounds or alveolar fricative sounds and the specific signals in the specific transition band (S230).

[0064] First, the signal detection unit 130 is used to detect the speech waveform corresponding to aspirated sounds or palato-alveolar fricative sounds in the speech waveform. Then, the signal detection unit 130 detects the specific signals in the transition band.

[0065] Figure 4 For the purpose of illustration in Figure 2 An example diagram showing the specific signals detected in step 230 shown.

[0066] As Figure 4 shown, the signal detection unit 130 amplifies the transition band to generate specific signals corresponding to aspirated sounds such as "settap" and "bolteu" and specific signals generated corresponding to palato-alveolar fricative sounds such as "same".

[0067] When step S230 is completed, the forgery determination unit 140 uses the detected speech waveform and specific signals to determine whether the corresponding speech file has been forged (S240).

[0068] In the case where the speech file has not been forged, specific signals are detected in the transition band that matches the speech waveform corresponding to aspirated sounds or palato-alveolar fricative sounds.

[0069] Therefore, the forgery determination unit 140 determines whether there is forgery by using whether specific signals are detected in the transition band that matches the speech waveform corresponding to aspirated sounds or palato-alveolar fricative sounds.

[0070] Next, Figures 5 to 8 More specifically, a forgery determination method according to an embodiment of the present invention will be described.

[0071] Figure 5 An example diagram showing the state of deleting specific signals by mixing and pasting a speech file according to an embodiment of the present invention.

[0072] As Figure 5 shown, segment A represents the spectrum corresponding to the segment generating aspirated sounds or palato-alveolar fricative sounds, and segment B represents the spectrum corresponding to the segment generating sounds other than aspirated sounds or palato-alveolar fricative sounds.

[0073] In addition, assume that a forger copies segment A, mixes and pastes it into segment B, and then performs forgery. At the same time, the mixing and pasting is performed by copying only the spectrum within a low-pass filter or a band-pass filter.

[0074] Therefore, since the forgery determination unit 140 does not detect a specific signal in the transition band corresponding to segment B, it determines that the corresponding voice file has been forged.

[0075] Figure 6 An example diagram showing a state in which an irregular specific signal is detected by a compressor after mixing and pasting aspirated sounds or palato-alveolar fricatives in a voice file according to an embodiment of the present invention.

[0076] As Figure 6 shown, segment A is a spectrum corresponding to a segment that generates aspirated sounds or palato-alveolar fricatives, and segment B is a spectrum corresponding to a segment that generates sounds other than aspirated sounds or palato-alveolar fricatives.

[0077] In addition, assume that a forger copies segment A, mixes and pastes it into segment B, and then performs forgery using a compressor.

[0078] As Figure 6 shown, specific signals are detected in both the transition band corresponding to segment A and the transition band corresponding to segment B. However, if the specific signal in the transition band corresponding to segment B is uneven, the forgery determination unit 140 determines that forgery has occurred in segment B.

[0079] Figure 7 An example diagram showing a state in which an irregular specific signal is detected by a compressor after mixing and pasting sounds other than aspirated sounds or palato-alveolar fricatives in a voice file according to an embodiment of the present invention.

[0080] As Figure 7 shown, when encoding is performed using a compressor, blowing may also occur in sounds other than aspirated sounds or palato-alveolar fricatives. Therefore, if the detected specific signal has a higher spectral output than other specific signals, the forgery determination unit 140 determines that forgery using a compressor has occurred in the corresponding voice file.

[0081] Figure 8 An example diagram showing a state in which someone else's voice is mixed and pasted into a voice file according to an embodiment of the present invention.

[0082] As Figure 8As shown, the forgery determination unit 140 sets a reference value for the specific signal. Then, when a plurality of specific signals corresponding to "bolteu" are detected, the forgery determination unit 140 calculates the difference between the specific signal in the A segment and the specific signal in the B segment, and compares the calculated difference with the reference value.

[0083] Furthermore, if the calculated difference is greater than the reference value, the forgery determination unit 140 determines that the speaker of the A segment is different from the speaker of the B segment, and determines that the corresponding voice file has been forged.

[0084] If it is determined in step S240 that forgery has occurred in the corresponding voice file, the control unit 150 displays the forged segment and outputs it on the screen.

[0085] Specifically, as described in step S240, forgery may occur in a specific segment. Then, the control unit 150 displays the specific segment and outputs it to the screen, so that the user can visually identify whether the voice file has been forged.

[0086] As described above, according to the voice file forgery detection device of the embodiment of the present invention, the transition band is independently enlarged from the voice file output in the form of a graph, and the forgery is determined by using the blowing phenomenon generated in the enlarged transition band. Therefore, forgery can be detected even in a voice file that has been mixed, pasted, and compressed. Moreover, since the edited segment of the voice file can be objectively detected, it can also be used as objective evidence material.

[0087] In addition, according to the voice file forgery detection device of the embodiment of the present invention, different aspirated sounds or alveolar palatal fricatives are emitted according to each person's oral structure or the amount of air exhaled from the lungs. Therefore, the corresponding blowing phenomenon is output differently, so that it is possible to determine whether the voice of another person has been synthesized or whether the recording has been made in another space, etc., and it has the effect of reducing the physical time required to analyze the voice file and reducing the cost caused by the involvement of experts.

[0088] Although the present invention has been described with reference to the embodiments shown in the drawings, these are only examples, and those of ordinary skill in the art should understand that various modifications and other equivalent embodiments can be made therefrom. Therefore, the true technical protection scope of the present invention should be determined by the technical idea of the appended claims.

Claims

1. A voice file forgery detection device, which is a voice file forgery detection device that uses voice waveforms for specific pronunciations. Among them, The voice file forgery detection device includes: a receiving unit configured to receive a voice file recorded through a user terminal; a preprocessing unit configured to output the received voice file in the form of a graph composed of a time axis and a frequency axis, and independently magnify a specific transition band in the output graph; a signal detection unit configured to extract a specific signal generated due to a blowing phenomenon in the specific transition band, and extract a voice waveform corresponding to aspirated sound or alveolo-palatal fricative; a forgery determination unit configured to analyze a correlation between the specific signal and the voice waveform, and determine whether the voice file has been forged based on an analysis result; and a control unit configured to, if it is determined that the voice file has been forged, display a forged segment and output it to a screen.

2. The voice file forgery detection device according to claim 1, wherein, The forgery determination unit determines whether a specific signal is detected in a segment matching the voice waveform corresponding to the aspirated sound or alveolo-palatal fricative, so as to determine whether it has been forged.

3. The voice file forgery detection device according to claim 2, wherein, In the forgery determination unit, if no specific signal is detected, the spectrum corresponding to the segment where the aspirated sound or alveolo-palatal fricative occurs in the same voice file or another voice file is mixed and pasted into the corresponding segment, so as to determine that it has been forged.

4. The voice file forgery detection device according to claim 2, wherein, In the forgery determination unit, if a specific signal is detected but the spectrum of the specific signal is irregularly output, the spectrum corresponding to the segment where the aspirated sound or alveolo-palatal fricative occurs in the same voice file or another voice file is mixed and pasted into the corresponding segment, and then a compressor is applied to determine that it has been forged.

5. The voice file forgery detection device according to claim 2, wherein, In the forgery determination unit, when a specific signal is detected, the magnitude of the specific signal corresponding to the same aspirated sound or alveolo-palatal fricative and the difference value of the specific signal are compared with a reference value. If the difference value is greater than the reference value, the spectrum corresponding to the segment where the aspirated sound or alveolo-palatal fricative made by others in the same voice file or another voice file is mixed and pasted into the corresponding segment, so as to determine that it has been forged.

6. The voice file forgery detection device according to claim 2, wherein, In the forgery determination unit, it is determined whether a specific signal is detected in a segment matching the voice waveform corresponding to a sound other than the aspirated sound or alveolo-palatal fricative. If a specific signal is detected, the spectrum corresponding to the segment where the sound other than the aspirated sound or alveolo-palatal fricative occurs in the same voice file or another voice file is mixed and pasted into the corresponding segment, so as to determine that it has been forged.

7. A voice file forgery detection method, which is a forgery detection method that uses a voice file forgery detection device. Among them, The voice file forgery detection method includes: a step of receiving a voice file recorded through a user terminal; a step of outputting the received voice file in the form of a graph composed of a time axis and a frequency axis, and independently magnifying a specific transition band in the output graph; a step of extracting a specific signal generated due to a blowing phenomenon in the specific transition band, and extracting a voice waveform corresponding to aspirated sound or alveolo-palatal fricative; a step of analyzing a correlation between the specific signal and the voice waveform, and determining whether the voice file has been forged based on an analysis result; and A step of displaying a forged segment and outputting it to a screen if it is determined that the voice file has been forged.

8. The voice file forgery detection method according to claim 7, wherein, The step of determining whether the voice file has been forged is to determine whether a specific signal is detected in a segment that matches the voice waveform corresponding to the aspirated sound or the alveolar palatal fricative sound, so as to determine whether it has been forged.

9. The voice file forgery detection method according to claim 8, wherein, The step of determining whether the voice file has been forged is that if no specific signal is detected, the spectrum corresponding to the segment occurring in the aspirated sound or the alveolar palatal fricative sound included in the same voice file or another voice file is mixed and pasted into the corresponding segment, so as to determine that it has been forged.

10. The voice file forgery detection device according to claim 9, wherein, The step of determining whether the voice file has been forged is that if a specific signal is detected but the spectrum of the specific signal is output irregularly, the spectrum corresponding to the segment occurring in the aspirated sound or the alveolar palatal fricative sound included in the same voice file or another voice file is mixed and pasted into the corresponding segment, and then a compressor is applied to determine that it has been forged.

11. The voice file forgery detection method according to claim 8, wherein, The step of determining whether the voice file has been forged is that when a specific signal is detected, the magnitude of the specific signal corresponding to the same aspirated sound or alveolar palatal fricative sound and the difference value of the specific signal are compared with a reference value. If the difference value is greater than the reference value, the spectrum corresponding to the segment occurring in the aspirated sound or the alveolar palatal fricative sound included in the same voice file or another voice file and made by others is mixed and pasted into the corresponding segment, so as to determine that it has been forged.

12. The method for detecting forged speech files according to claim 8, wherein, The step of determining whether the voice file has been forged is to determine whether a specific signal is detected in a segment that matches the voice waveform corresponding to a sound other than the aspirated sound and the alveolar palatal fricative sound. If a specific signal is detected, the spectrum corresponding to the segment occurring in a sound other than the aspirated sound or the alveolar palatal fricative sound included in the same voice file or another voice file is mixed and pasted into the corresponding segment, so as to determine that it has been forged.

Citation Information

Patent Citations

  • Audio file forgery detection apparatus

    KR101382356B1