Voice detection device, voice detection method, and recording medium

The voice detection device dynamically adjusts thresholds based on tentative voice segment characteristics to enhance the accuracy and efficiency of post-processing operations by ensuring detected voice segments contain sufficient information and reduce computational load.

JP7718578B2Active Publication Date: 2025-08-05NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024508838
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-22
Publication Date
2025-08-05
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

Existing voice detection technologies struggle to accurately determine the start and end of voice segments in audio signals, leading to inefficient and inaccurate post-processing operations such as speech recognition, voice authentication, or emotion recognition.

Method used

A voice detection device and method that dynamically adjusts the threshold for determining the end of a voice segment based on the characteristics of the tentative voice segment, such as its length, to ensure the detected segment is appropriate for subsequent processing.

Benefits of technology

The solution enhances the accuracy of post-processing operations by ensuring the detected voice segments contain sufficient information while minimizing the computational load, thus improving the efficiency and effectiveness of speech recognition, voice authentication, and emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007718578000001
    Figure 0007718578000001
  • Figure 0007718578000002
    Figure 0007718578000002
  • Figure 0007718578000003
    Figure 0007718578000003
Patent Text Reader

Abstract

This voice detection device 1 comprises a start point determination means 112 which determines a start point of a voice section including a voice contained in a voice signal, an end point determination means 112 which determines the end point of the voice section by determining whether or not a length Lb of a non-voice section contained after the start point is determined to be longer than or equal to a threshold value TH, and a setting means 113 which sets the threshold value TH on the basis of a characteristic of a preliminary voice section which starts from the start point.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the technical field of a voice detection device, a voice detection method, and a recording medium that can detect a voice section that appears in an audio signal, for example. [Background technology]

[0002] An example of a speech detection device capable of detecting speech segments appearing in an audio signal is described in Patent Document 1. Other prior art documents related to this disclosure include Patent Documents 2 to 4 and Non-Patent Document 1. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2021 / 014612 Brochure [Patent Document 2] International Publication No. 2016 / 143125 Brochure [Patent Document 3] Japanese Patent Application Publication No. 2017-097330 [Patent Document 4] International Publication No. 2015 / 059947 Brochure [Non-patent literature]

[0004] [Non-Patent Document 1] Takenori Yoshimura et.al, “END-TO-END AUTOMATIC SPEECH RECOGNITION INTEGRATED WITH CTC-BASED VOICE ACTIVITY DETECTION”, arXiv 2002.00551, February 14, 2020 Summary of the Invention [Problem to be solved by the invention]

[0005] The present disclosure aims to provide a voice detection device, a voice detection method, and a recording medium that aim to improve upon the techniques described in prior art documents. [Means for solving the problem]

[0006] One aspect of the voice detection device disclosed herein comprises a start determination means for determining the start of a voice section that includes voice that appears in a voice signal, an end determination means for determining the end of the voice section by determining whether the length of a non-voice section that appears after the start is determined is equal to or greater than a threshold, and a setting means for setting the threshold based on the characteristics of a tentative voice section starting from the start.

[0007] One aspect of the speech detection method of this disclosure includes determining the start of a speech section that includes speech that appears in a speech signal, determining the end of the speech section by determining whether the length of a non-speech section that appears after the start is determined is equal to or greater than a threshold, and setting the threshold based on characteristics of a tentative speech section that begins from the start.

[0008] One aspect of a recording medium of this disclosure is a recording medium having recorded thereon a computer program that causes a computer to execute a voice detection method, the voice detection method including: determining the start of a voice section that includes voice appearing in a voice signal; determining the end of the voice section by determining whether the length of a non-voice section that appears after the start is determined is equal to or greater than a threshold; and setting the threshold based on characteristics of a tentative voice section starting from the start. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing the configuration of the voice detection device according to the first embodiment. [Figure 2] FIG. 2 shows the relationship between a speech signal, a speech section, and a non-speech section. [Figure 3] FIG. 3 is a block diagram showing the configuration of the voice detection device according to the second embodiment. [Figure 4]FIG. 4 is a block diagram showing the configuration of the symbol generating unit. [Figure 5] FIG. 5(a) shows a method for determining the start of a speech section together with symbol data, and FIG. 5(b) shows a method for determining the end of a speech section together with symbol data. [Figure 6] FIG. 6 is a flowchart showing the flow of the voice detection operation performed by the voice detection device in the second embodiment. [Figure 7] FIG. 7 shows symbol data in which the start of the voice section has been determined. [Figure 8] FIG. 8 is a graph showing an example of the relationship between the length of a tentative voice segment and a threshold value. [Figure 9] FIG. 9(a) shows a voice segment detected by the voice detection device of the comparative example, and FIG. 9(b) shows a voice segment detected by the voice detection device of the second embodiment. [Figure 10] FIG. 10(a) shows a voice segment detected by the voice detection device of the comparative example, and FIG. 10(b) shows a voice segment detected by the voice detection device of the second embodiment. [Figure 11] Each of FIGS. 11(a) to 11(c) is a graph showing an example of the relationship between the length of a tentative voice segment and a threshold value. [Figure 12] FIG. 12 is a block diagram showing the configuration of the voice detection device according to the third embodiment. [Figure 13] FIG. 13 is a block diagram showing the configuration of a voice detection device according to the fourth embodiment. [Figure 14] FIG. 14 is a block diagram showing the configuration of a voice detection device according to the fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of a voice detection device, a voice detection method, and a recording medium will be described with reference to the drawings.

[0011] (1) First embodiment First, a first embodiment of a voice detection device, a voice detection method, and a recording medium will be described. Hereinafter, the first embodiment of a voice detection device, a voice detection method, and a recording medium will be described with reference to Fig. 1 using a voice detection device 1000 to which the first embodiment of the voice detection device, the voice detection method, and a recording medium is applied. Fig. 1 is a block diagram showing the configuration of the voice detection device 1000 in the first embodiment.

[0012] As shown in FIG. 1, the speech detection device 1000 includes a start determination unit 1001, an end determination unit 1002, and a setting unit 1003. As shown in FIG. 2, the start determination unit 1001 determines the start of a speech section that appears in a speech signal. As shown in FIG. 2, the end determination unit 1002 determines the end of a speech section by determining whether the length Lb of a non-speech section that appears after the start is determined is equal to or greater than a threshold TH. For example, the end determination unit 1002 may determine the time corresponding to the end of a speech section based on the time at which the length Lb of the non-speech section becomes equal to or greater than the threshold TH. The setting unit 1003 sets the threshold TH based on the characteristics of a tentative speech section starting from the start (i.e., a tentative speech section whose end has not yet been determined). For example, as shown in FIG. 2, the setting unit 1003 may set the threshold TH so that the threshold TH is changed from a first candidate value TH1 to a second candidate value TH2 when the characteristics of the tentative speech section (length in the example shown in FIG. 2) change.

[0013] As described above, the speech detection device 1000 of the first embodiment can set (i.e., change) the threshold value TH based on the tentative speech segment length Lt. Therefore, the speech detection device 1000 can detect a speech segment having a length appropriate for a post-processing operation (e.g., a speech recognition operation, a voice authentication operation, or an emotion recognition operation) that is performed after the speech segment is detected.

[0014] (2) Second embodiment Next, a second embodiment of the voice detection device, the voice detection method, and the recording medium will be described. Hereinafter, the second embodiment of the voice detection device, the voice detection method, and the recording medium will be described using the voice detection device 1 to which the second embodiment of the voice detection device, the voice detection method, and the recording medium is applied.

[0015] The voice detection device 1 is a device that performs voice activity detection (VAD). The voice activity detection operation is an operation that detects a voice activity from a voice signal that indicates a voice uttered by a speaker. In other words, the voice activity detection operation is an operation that distinguishes a voice activity that appears in a voice signal from a non-voice activity that appears in the voice signal. A voice activity is an activity that includes a voice uttered by a speaker. In other words, a voice activity is an activity that occurs when a speaker is uttering a voice. On the other hand, a non-voice activity is an activity that differs from a voice activity. Typically, a non-voice activity is an activity that occurs when a speaker is not uttering a voice.

[0016] The following describes the voice activity detection device 1 that performs such a voice activity detection operation.

[0017] (2-1) Configuration of the voice detection device 1 First, the configuration of a voice detection device 1 in the second embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of a voice detection device 1 in the second embodiment.

[0018] 3, the voice detection device 1 includes a calculation device 11, a storage device 12, and a communication device 13. The voice detection device 1 may further include an input device 14 and an output device 15. However, the voice detection device 1 does not necessarily have to include at least one of the input device 14 and the output device 15. The calculation device 11, the storage device 12, the communication device 13, the input device 14, and the output device 15 may be connected via a data bus 16.

[0019] The arithmetic device 11 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic device 11 loads a computer program. For example, the arithmetic device 11 may load a computer program stored in the storage device 12. For example, the arithmetic device 11 may load a computer program stored in a computer-readable, non-transitory storage medium using a storage medium reading device (not shown) included in the voice detection device 1. The arithmetic device 11 may acquire (i.e., download or load) the computer program from a device (not shown) located outside the voice detection device 1 via the communication device 13 (or another communication device). The arithmetic device 11 executes the loaded computer program. As a result, logical functional blocks for executing operations to be performed by the voice detection device 1 (for example, the above-mentioned voice activity detection operation) are realized within the arithmetic device 11. In other words, the arithmetic device 11 can function as a controller for realizing logical functional blocks for executing the operations (in other words, processing) that the voice detection device 1 should perform.

[0020] Fig. 3 shows an example of logical functional blocks implemented in the calculation device 11 for executing the voice activity detection operation. As shown in Fig. 3, implemented in the calculation device 11 are a symbol generation unit 111, which is a specific example of "generation means" described in the appendix to be described later, a voice activity detection unit 112, which is a specific example of "starting edge determination means" and "ending edge determination means" described in the appendix to be described later, and a threshold setting unit 113, which is a specific example of "setting means" described in the appendix to be described later.

[0021] The symbol generation unit 111 generates symbol data from the audio signal. Specifically, the symbol generation unit 111 outputs a symbol for each audio frame SF (e.g., audio frame SF of several tens of milliseconds) obtained by subdividing the audio signal. The symbols may include character symbols that represent the audio spoken by the speaker in the audio frame SF as characters. One character symbol may represent one character (e.g., one alphabet, one hiragana, one Hangul character, or one Chinese character). For example, one character symbol may represent a single alphabet, such as "a." One character symbol may represent multiple characters (e.g., multiple alphabets, multiple hiragana, multiple Hangul characters, or multiple Chinese characters). For example, one character symbol may represent multiple alphabets, such as "pat." The symbols may include a blank symbol that represents the absence of audio spoken by the speaker in the audio frame SF. As a result, the symbol generation unit 111 generates symbol data in which the output symbols are arranged in chronological order. The character symbols may be symbols that represent characters (for example, hiragana or alphabet) themselves, or may be symbols that represent phonemes, which are the smallest structural units of characters.

[0022] In the second embodiment, the symbol generation unit 111 generates symbol data from a speech signal using a Connectionist Temporal Classification (CTC) model. A method for generating symbol data from a speech signal using the CTC model is described in Non-Patent Document 1. Therefore, a detailed description of the method for generating symbol data from a speech signal using the CTC model will be omitted, but a brief overview will be provided below with reference to FIG. 4. The symbol generation unit 111, which generates symbol data from a speech signal using the CTC model, may be implemented by a recurrent neural network, as shown in FIG. 2. Specifically, the symbol generation unit 111 divides a speech signal into multiple speech frames SF and inputs the multiple speech frames SF to multiple long short-term memories (LTSMs), respectively. The neural network including the multiple LTSMs outputs a posterior probability that each of multiple types of characters corresponds to the speech uttered by the speaker in each speech frame SF. The symbol generation unit 111 then generates symbol data from sequence data of multiple symbols constituting a character string with the highest posterior probability. FIG. 4 shows an example of symbol data generated by the symbol generating unit 111 when the posterior probability of the character string "GO--" is the highest.

[0023] In Fig. 4, the symbol "-" means a blank symbol. A blank symbol is output when there is a low possibility that speech has been uttered in a certain speech frame SF. In other words, a blank symbol is output when there is no character corresponding to a certain speech frame.

[0024] 3 again, the voice activity detection unit 112 detects a voice activity using the symbol data generated by the symbol generation unit 111. An overview of the operation of the voice activity detection unit 112 to detect a voice activity will be described with reference to FIGS. 5(a) and 5(b).

[0025] First, as shown in FIG. 5(a), the voice activity detection unit 112 determines the start of a voice activity. Specifically, the voice activity detection unit 112 detects a character symbol by searching symbol data in chronological order while the start of the voice activity has not yet been determined (i.e., has not yet been detected). Thereafter, the voice activity detection unit 112 determines, as the start of the voice activity, the voice frame SF that is a predetermined number of frames MS before the voice frame SF in which the character symbol was detected. In the example shown in FIG. 5(a), the predetermined number of frames MS is 2. However, the predetermined number of frames MS may be 0, 1, or 3 or more.

[0026] After that, the voice activity detection unit 112 determines the end of the voice activity. Specifically, as shown in FIG. 5(b), the voice activity detection unit 112 searches symbol data in time series under the condition where the start of the voice activity has been determined, and determines whether the length Lb of the non-voice activity that appears after the start of the voice activity is determined is equal to or greater than a predetermined threshold value TH. The non-voice activity is a period in which a blank symbol is output. In this case, the number of blank symbols output consecutively in time series (i.e., the number of voice frames SF in which blank symbols are output) may be used as the length Lb of the non-voice activity. In the following description, an example will be given in which the number of blank symbols output consecutively in time series (hereinafter referred to as the "number of blank symbols BSN") is used as the length Lb of the non-voice activity. If it is determined that the number of blank symbols BSN is equal to or greater than the predetermined threshold value TH (i.e., the length Lb of the non-voice activity is equal to or greater than the predetermined threshold value TH), the voice activity detection unit 112 determines the voice activity frame that is a predetermined number ME of frames after the voice frame in which the last character symbol was detected as the end of the voice activity. In the example shown in FIG. 5(b), the predetermined number ME is 2. However, the predetermined number of frames ME may be 0, 1, or 3 or more.

[0027] 3 again, the threshold setting unit 113 sets a threshold value TH that is used by the voice activity detection unit 112 to determine the end of a voice activity. Note that the method by which the threshold setting unit 113 sets the threshold value TH will be described in detail later with reference to FIG. 6 etc.

[0028] The storage device 12 can store desired data. For example, the storage device 12 may temporarily store a computer program executed by the arithmetic device 11. The storage device 12 may temporarily store data that the arithmetic device 11 temporarily uses when the arithmetic device 11 is executing a computer program. The storage device 12 may store data that the voice detection device 1 stores long-term. The storage device 12 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. In other words, the storage device 12 may include a non-temporary recording medium.

[0029] The communication device 13 is capable of communicating with devices external to the voice detection device 1 .

[0030] The input device 14 is a device that accepts information input to the voice detection device 1 from outside the voice detection device 1. For example, the input device 14 may include an operation device (for example, at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of the voice detection device 1. For example, the input device 14 may include a reading device that can read information recorded as data on a recording medium that can be externally attached to the voice detection device 1.

[0031] The output device 15 is a device that outputs information to the outside of the voice detection device 1. For example, the output device 15 may output information as an image. That is, the output device 15 may include a display device (a so-called display) that can display an image showing the information to be output. For example, the output device 15 may output information as sound. That is, the output device 15 may include an audio device (a so-called speaker) that can output sound. For example, the output device 15 may output information on paper. That is, the output device 15 may include a printing device (a so-called printer) that can print desired information on paper.

[0032] (2-2) Voice detection operation performed by the voice detection device 1 Next, the voice detection operation performed by the voice detection device 1 will be described with reference to Fig. 6. Fig. 6 is a flowchart showing the flow of the voice detection operation performed by the voice detection device 1 in the second embodiment.

[0033] 6, the symbol generation unit 111 generates symbol data from an audio signal (step S100). For example, the symbol generation unit 111 may acquire an audio signal generated by an audio sensor such as a microphone and generate symbol data from the acquired audio signal. In this case, the symbol generation unit 111 may continue to acquire the audio signal and generate symbol data as long as the audio signal continues to be generated. Alternatively, for example, the symbol generation unit 111 may read an audio signal recorded on a recording medium and generate symbol data from the read audio data.

[0034] Thereafter, the speech section detection unit 112 determines the start of the speech section based on the symbol data generated in step S100 (step S101). Thereafter, the speech section detection unit 112 determines the end of the speech section based on the symbol data generated in step S100 (steps S103 to S104). That is, the speech section detection unit 112 determines whether the number of blank symbols BSN is equal to or greater than a threshold value TH (step S103). The speech section detection unit 112 determines the end of the speech section based on the determination result in step S103 (step S104).

[0035] Particularly in the second embodiment, the threshold setting unit 113 sets the threshold TH to be used in step S103 from when the start point of the speech section is determined until the end point of the speech section is determined (step S102). Specifically, the threshold setting unit 113 sets (i.e., changes) the threshold TH based on the characteristics of the tentative speech section starting from the start point determined in step S101.

[0036] A tentative voice segment refers to a tentative voice segment whose end has not been determined. Specifically, as shown in FIG. 7 showing symbol data in which the start point of the voice segment has been determined, in the second embodiment, until the end point of the voice segment is determined, a voice segment starting from the start point determined in step S101 is referred to as a tentative voice segment whose end has not been definitively determined. A voice frame SF currently being focused on for voice segment detection (hereinafter referred to as a "frame of interest") may be used as the tentative end point of the tentative voice segment. The frame of interest may refer to the voice frame SF corresponding to the last symbol currently searched when searching symbol data in chronological order for voice segment detection.

[0037] In the second embodiment, an example will be described in which the length Lt of a tentative voice segment is used as a characteristic of the tentative voice segment. In the following, an example will be described in which the number of voice frames SF included in the tentative voice segment Lt (i.e., the number of symbols included in the tentative voice segment Lt) is used as the length Lt of the tentative voice segment. In this case, the threshold setting unit 113 sets the threshold TH based on the length Lt of the tentative voice segment. Specifically, the threshold setting unit 113 changes the threshold TH based on the length Lt of the tentative voice segment. For example, the threshold setting unit 113 may set the threshold TH so that the threshold TH set when the length Lt of the tentative voice segment is a first length is different from the threshold TH set when the length Lt of the tentative voice segment is a second length different from the first length.

[0038] Particularly in the second embodiment, the threshold setting unit 113 may set the threshold TH so that the threshold TH set when the tentative voice segment length Lt is a first length is greater than the threshold TH set when the tentative voice segment length Lt is a second length longer than the first length. For example, as shown in FIG. 8, the threshold setting unit 113 may set the threshold TH to a first candidate value TH11 when the tentative voice segment length Lt is shorter than Lt11. Furthermore, the threshold setting unit 113 may set the threshold TH to a second candidate value TH12 smaller than the first candidate value TH11 when the tentative voice segment length Lt is longer than Lt11 and shorter than Lt12 (where Lt12 is longer than Lt11). Furthermore, the threshold setting unit 113 may set the threshold TH to a third candidate value TH13 smaller than the second candidate value TH12 when the tentative voice segment length Lt is longer than Lt12. That is, in the example shown in FIG. 8, the threshold value setting unit 113 sets the threshold value TH to one candidate value selected from three different candidate values based on the length Lt of the tentative voice segment.

[0039] Referring back to FIG. 6, the voice detection device 1 repeats the same operation to detect the voice section until it has been completed for all sections of the symbol data generated in step S100 (step S105).

[0040] (2-3) Technical Effects of Voice Detection Device 1 As described above, the speech detection device 1 of the second embodiment can set (i.e., change) the threshold value TH based on the length Lt of the tentative speech segment. This allows the speech detection device 1 to detect a speech segment having a length appropriate for post-processing operations (e.g., speech recognition operation, voice authentication operation, or emotion recognition operation) performed after the speech segment is detected. Specific reasons why a speech segment having a length appropriate for post-processing operations can be detected will be described below with reference to FIGS. 9(a) to 9(b) and 10(a) to 10(b).

[0041] First, the speech detection device of the comparative example, which has a fixed threshold TH regardless of the length Lt of the tentative speech section, may detect a speech section that is shorter than necessary. For example, if a speaker speaks for only a short time and then takes a short pause, the speech detection device of the comparative example is likely to detect a short speech section containing speech uttered during the short utterance. For example, FIG. 9(a) shows a speech section detected by the speech detection device of the comparative example, which has a fixed threshold TH of 5 (i.e., 5 frames). In the example shown in FIG. 9(a), the speech detection device of the comparative example detects the character symbol "a" when the Nth speech frame SF becomes the frame of interest. Therefore, the speech detection device of the comparative example determines that the (N-2)th speech frame SF, which is a predetermined number of frames MS (in this case, 2 frames) before the Nth speech frame SF, is the start of the speech section. Then, when the (N+5)th speech frame SF becomes the frame of interest, the speech detection device of the comparative example determines that the length Lb of the non-speech section (i.e., the number of blank symbols BSN) is equal to or greater than the threshold TH. Therefore, the speech detection device of the comparative example determines the (N+2)th speech frame SF, which is a predetermined number of frames ME (two frames in this case) after the Nth speech frame SF in which the last character symbol was detected, as the end of the speech segment. As a result, the speech detection device of the comparative example detects a relatively short speech segment having a length of five frames. In this case, as shown in FIG. 9(a), the number of character symbols contained in the detected speech segment is not necessarily large. This is because the shorter the speech segment, the fewer character symbols it contains. In other words, the speech detection device of the comparative example may detect a speech segment that does not contain sufficient information. As a result, the accuracy of post-processing operations performed after speech segment detection may be reduced. For example, the context of the sentence representing the speech uttered by the speaker may not be properly understood by the speech recognition operation.

[0042] However, in the second embodiment, the voice detection device 1 determines the threshold value TH based on the length Lt of the tentative voice segment. Therefore, compared to the voice detection device of the comparative example, the voice detection device 1 is less likely to detect a voice segment that is shorter than necessary. For example, FIG. 9(b) shows a voice segment detected by the voice detection device 1 that sets the threshold value TH to 7 (7 frames) when the tentative voice segment length Lt is 10 frames or less, sets the threshold value TH to 5 (5 frames) when the tentative voice segment length Lt is 11 frames or more but 15 frames or less, and sets the threshold value TH to 3 (3 frames) when the tentative voice segment length Lt is 16 frames or more. In the example shown in FIG. 9(b), the voice detection device 1 determines the (N-2)th voice frame SF as the start of the voice segment, similar to the voice detection device of the comparative example shown in FIG. 9(a). At this stage, the length Lt of the tentative voice segment is 3 frames, so the threshold value TH is set to 7. Furthermore, even when the (N+5)th speech frame SF is the frame of interest, the length Lt of the tentative speech section is 8 frames, so the threshold value TH is set to 7. As a result, unlike the speech detection device of the comparative example, the speech detection device 1 does not determine that the length Lb of the non-speech section is equal to or greater than the threshold value TH when the (N+5)th speech frame SF is the frame of interest. Thereafter, when the (N+13)th speech frame SF is the frame of interest, the length Lt of the tentative speech section is equal to or greater than 16 frames, so the threshold value TH is set to 3. As a result, the speech detection device 1 determines that the length Lb of the non-speech section is equal to or greater than the threshold value TH when the (N+13)th speech frame SF is the frame of interest. Therefore ... end of the speech section is the (N+12)th speech frame SF, which is a predetermined number of frames ME (two frames in this case) after the (N+10)th speech frame SF in which the last character symbol was detected. As a result, the speech detection device 1 detects a speech section that is longer than the speech section detected by the speech detection device of the comparative example. In other words, the voice detection device 1 can solve the technical problem of detecting a voice section that is shorter than necessary. Therefore, compared to the voice detection device of the comparative example, the voice detection device 1 is more likely to detect a voice section that contains sufficient information.As a result, the accuracy of the post-processing operation performed after the voice activity is detected by the voice activity detection device 1 is higher than the accuracy of the post-processing operation performed after the voice activity is detected by the voice activity detection device of the comparative example.

[0043] On the other hand, the voice detection device of the comparative example, in which the threshold value TH is fixed regardless of the tentative voice segment length Lt, may detect an unnecessarily long voice segment in addition to or instead of an unnecessarily short voice segment. For example, when a speaker speaks rapidly and continuously, the voice detection device of the comparative example is more likely to detect an unnecessarily long voice segment. For example, FIG. 10(a) shows a voice segment detected by the voice detection device of the comparative example, in which the threshold value TH is fixed at 5 (5 frames). In the example shown in FIG. 10(a), the voice detection device of the comparative example detects the character symbol "a" when the Mth voice frame SF becomes the frame of interest. Therefore, the voice detection device of the comparative example determines that the (M-2)th voice frame SF, which is a predetermined number of frames MS (in this case, 2 frames) before the Mth voice frame SF, is the start of the voice segment. Then, when the (M+23)th voice frame SF becomes the frame of interest, the voice detection device of the comparative example determines that the length Lb of the non-voice segment (i.e., the number of blank symbols BSN) is equal to or greater than the threshold value TH. Therefore, the voice detection device of the comparative example determines the (M+20)th voice frame SF, which is a predetermined number of frames ME (two frames in this case) after the (M+18)th voice frame SF in which the last character symbol was detected, as the end of the voice activity. As a result, the voice detection device of the comparative example detects a relatively long voice activity with a length of 24 frames. In this case, the amount of calculation required for the post-processing operation performed after the voice activity is detected may be excessively large. This is because the longer the voice activity, the greater the amount of calculation required for the post-processing operation performed after the voice activity is detected. This may result in a long delay time between when the voice signal is input to the voice detection device of the comparative example and when the result of the post-processing operation is output.

[0044] However, in the second embodiment, the voice detection device 1 determines the threshold value TH based on the length Lt of the tentative voice segment. Therefore, compared to the voice detection device of the comparative example, the voice detection device 1 is less likely to detect a voice segment that is longer than necessary. For example, FIG. 10(b) shows a voice segment detected by the voice detection device 1 that sets the threshold value TH to 7 (7 frames) when the tentative voice segment length Lt is 10 frames or less, sets the threshold value TH to 5 (5 frames) when the tentative voice segment length Lt is 11 frames or more but 15 frames or less, and sets the threshold value TH to 3 (3 frames) when the tentative voice segment length Lt is 16 frames or more. In the example shown in FIG. 10(b), the voice detection device 1 determines the (M-2)th voice frame SF as the start of the voice segment, similar to the voice detection device of the comparative example shown in FIG. 10(a). Thereafter, when the (M+13)th voice frame SF is the frame of interest, the length Lt of the tentative voice segment is 16 frames or more, so the threshold value TH is set to 3. As a result, the voice detection device 1 determines that the length Lb of the non-voice section is equal to or greater than the threshold value TH when the (M+13)th voice frame SF becomes the frame of interest. Therefore, the voice detection device 1 determines that the (M+12)th voice frame SF, which is a predetermined number of frames ME (two frames in this case) after the (M+10)th voice frame SF in which the last character symbol was detected, is the end of the voice section. As a result, the voice detection device 1 detects a voice section that is shorter than the voice section detected by the voice detection device of the comparative example. In other words, the voice detection device 1 can solve the technical problem of detecting a voice section that is longer than necessary. Therefore, compared to the voice detection device of the comparative example, the voice detection device 1 is less likely to detect a voice section that requires an excessively large amount of calculation for post-processing. As a result, the amount of calculation required for post-processing performed after the voice detection by the voice detection device 1 is less than the amount of calculation required for post-processing performed after the voice detection by the voice detection device of the comparative example.

[0045] As such, compared to the voice detection device of the comparative example, the voice detection device 1 is less likely to detect a voice segment that is too short or too long for the post-processing operation that is performed after the voice detection. In other words, the voice detection device 1 can detect a voice segment that has a length appropriate for the post-processing operation that is performed after the voice detection.

[0046] Considering the technical effects described above, it is preferable that the speech detection device 1 sets the threshold value TH based on the length Lt of the provisional speech section so as to achieve both the effect of detecting a speech section long enough to grasp the context of the sentence indicated by the speech spoken by the speaker and the effect of making the amount of calculation required for post-processing operations appropriate.

[0047] Additionally, the voice activity detection device 1 detects the voice activity using symbol data generated using the CTC model, which allows the voice activity detection device 1 to appropriately detect the voice activity.

[0048] (2-4) Modification In the example shown in FIG. 8 , the threshold setting unit 113 sets the threshold TH to one candidate value selected from three different candidate values based on the tentative voice segment length Lt. However, the method of setting the threshold TH shown in FIG. 8 is just an example, and the method of setting the threshold TH is not limited to the setting method shown in FIG. 8. For example, as shown in FIG. 11( a), the threshold setting unit 113 may set the threshold TH to one candidate value selected from two different candidate values based on the tentative voice segment length Lt. For example, as shown in FIG. 11( b), the threshold setting unit 113 may set the threshold TH to one candidate value selected from four or more different candidate values based on the tentative voice segment length Lt. For example, as shown in FIG. 11( c), the threshold setting unit 113 may continuously change the threshold TH based on the tentative voice segment length Lt, in addition to or instead of gradually changing the threshold TH based on the tentative voice segment length Lt as shown in FIG. 8 and from FIG. 11( a) to FIG. 11( b).

[0049] In the above description, the voice activity detection unit 112 determines the end of a voice activity based on symbol data including multiple symbols constituting a character string with the highest posterior probability. However, the voice activity detection unit 112 may also determine the end of a voice activity based on symbol data including multiple symbols constituting a character string with a relatively high, but not the highest, posterior probability. For example, the voice activity detection unit 112 may determine the end of a voice activity based on symbol data including multiple symbols constituting a character string with the Nth highest posterior probability (where N is an integer equal to or greater than 1). In other words, the voice activity detection unit 112 may determine whether the length Lb of a non-voice activity is equal to or greater than a predetermined threshold TH using symbol data including multiple symbols constituting a character string with the Nth highest posterior probability. Even in this case, the voice activity detection unit 112 can appropriately set the end of a voice activity.

[0050] (3) Third embodiment Next, a third embodiment of the voice detection device, voice detection method, and recording medium will be described. Hereinafter, the third embodiment of the voice detection device, voice detection method, and recording medium will be described with reference to Fig. 12, using a voice detection device 1b to which the third embodiment of the voice detection device, voice detection method, and recording medium is applied. Fig. 12 is a block diagram showing the configuration of voice detection device 1b in the third embodiment.

[0051] 12, the voice detection device 1b of the third embodiment differs from the voice detection device 1 of the second embodiment in that it includes a threshold setting unit 113b instead of the threshold setting unit 113. Other features of the voice detection device 1b may be the same as other features of the voice detection device 1.

[0052] The threshold setting unit 113b differs from the above-described threshold setting unit 113 in that it uses a characteristic of the tentative voice segment other than the length Lt as a characteristic used to set the threshold TH. Other features of the threshold setting unit 113b may be the same as other features of the threshold setting unit 113.

[0053] For example, the threshold setting unit 113b may use the number of characters included in the tentative voice section (e.g., the number of characters represented by a character symbol) as a characteristic of the tentative voice section. Here, the longer the length Lt of the tentative voice section, the more likely it is that the number of characters included in the tentative voice section will be. Therefore, the number of characters included in the tentative voice section has a correlation with the length Lt of the tentative voice section. Therefore, the operation of setting the threshold value TH based on the number of characters included in the tentative voice section may be considered to be substantially equivalent to the operation of setting the threshold value TH based on the length Lt of the tentative voice section. In this case, the threshold setting unit 113b may set the threshold value TH based on the number of characters included in the tentative voice section in the same manner as when setting the threshold value TH based on the length Lt of the tentative voice section. For example, the threshold setting unit 113b may set the threshold value TH such that the threshold value TH set when the number of characters included in the tentative voice section is a first number is greater than the threshold value TH set when the number of characters included in the tentative voice section is a second number greater than the first number.

[0054] For example, the threshold setting unit 113b may use the number of words included in the tentative speech section as a characteristic of the tentative speech section. Because a word is a combination of characters, the speech detection device 1 can detect words based on the character symbols included in the symbol data. Specifically, the threshold setting unit 113b can detect words by performing morphological analysis on the character symbols included in the symbol data. Therefore, the threshold setting unit 113b can calculate the number of words included in the tentative speech section. Here, the longer the length Lt of the tentative speech section, the more likely it is that the number of words included in the tentative speech section will be. Therefore, the number of words included in the tentative speech section is correlated with the length Lt of the tentative speech section. Therefore, the operation of setting the threshold value TH based on the number of words included in the tentative speech section may be considered to be substantially equivalent to the operation of setting the threshold value TH based on the length Lt of the tentative speech section. In this case, the threshold setting unit 113b may set the threshold value TH based on the number of words included in the tentative speech section in the same manner as when setting the threshold value TH based on the length Lt of the tentative speech section. For example, the threshold setting unit 113b may set the threshold TH so that the threshold TH that is set when the number of words included in the provisional speech section is a first number is greater than the threshold TH that is set when the number of words included in the provisional speech section is a second number that is greater than the first number.

[0055] For example, the threshold setting unit 113b may use the speaking speed of the speech appearing in the tentative speech section as a characteristic of the tentative speech section. The faster the speaking speed, the more likely it is that a speech section will contain a large number of character symbols. As a result, the greater the number of character symbols contained in a speech section, the greater the amount of calculation required for post-processing. Therefore, considering the amount of calculation required for post-processing, it is preferable that the faster the speaking speed, the shorter the length of the speech section (and, as a result, the fewer character symbols contained in the speech section). Therefore, the threshold setting unit 113b may set the threshold TH so that the faster the speaking speed, the shorter the threshold TH. For example, the threshold setting unit 113b may set the threshold TH so that the threshold TH set when the speaking speed in the tentative speech section is a first speed is smaller than the threshold TH set when the speaking speed in the tentative speech section is a second speed slower than the first speed.

[0056] Note that the faster the speaking speed, the greater the number of characters per unit time (i.e., the number of character symbols). Also, the faster the speaking speed, the greater the number of words per unit time. Also, the faster the speaking speed, the fewer the number of blank symbols per unit time. Therefore, the threshold setting unit 113b may calculate at least one of the number of characters per unit time (i.e., the number of character symbols), the number of words per unit time, and the number of blank symbols per unit time as an index value representing the speaking speed.

[0057] For example, the threshold setting unit 113b may use the number of character symbols included in the tentative voice section as a characteristic of the tentative voice section. Here, the longer the length Lt of the tentative voice section, the more likely it is that the number of character symbols included in the tentative voice section will be. Therefore, the number of character symbols included in the tentative voice section is correlated with the length Lt of the tentative voice section. Therefore, the operation of setting the threshold value TH based on the number of character symbols included in the tentative voice section may be considered to be substantially equivalent to the operation of setting the threshold value TH based on the length Lt of the tentative voice section. In this case, the threshold setting unit 113b may set the threshold value TH based on the number of character symbols included in the tentative voice section in the same manner as when setting the threshold value TH based on the length Lt of the tentative voice section. For example, the threshold setting unit 113b may set the threshold value TH such that the threshold value TH set when the number of character symbols included in the tentative voice section is a first number is greater than the threshold value TH set when the number of character symbols included in the tentative voice section is a second number greater than the first number.

[0058] The voice detection device 1b in the third embodiment described above can achieve the same effects as those that can be achieved by the voice detection device 1 in the second embodiment described above.

[0059] (4) Fourth embodiment Next, a fourth embodiment of the voice detection device, voice detection method, and recording medium will be described. Hereinafter, the fourth embodiment of the voice detection device, voice detection method, and recording medium will be described with reference to Fig. 13 using a voice detection device 1c to which the fourth embodiment of the voice detection device, voice detection method, and recording medium is applied. Fig. 13 is a block diagram showing the configuration of the voice detection device 1c in the fourth embodiment.

[0060] 13, the voice detection device 1c of the fourth embodiment differs from at least one of the voice detection devices 1 of the second embodiment to the voice detection device 1b of the third embodiment in that it includes a threshold setting unit 113c instead of the threshold setting unit 113. Furthermore, the voice detection device 1c of the fourth embodiment differs from at least one of the voice detection devices 1 of the second embodiment to the voice detection device 1b of the third embodiment in that the storage device 12 stores speaker information 121c. Other features of the voice detection device 1c may be the same as other features of at least one of the voice detection devices 1 and 1b.

[0061] The threshold setting unit 113c differs from at least one of the threshold setting units 113 and 113b in that it sets the threshold TH based on the speaker information 121c in addition to or instead of the characteristics of the tentative voice segment. Other features of the threshold setting unit 113c may be the same as other features of at least one of the threshold setting units 113 and 113b.

[0062] The speaker information 121c includes information about the characteristics of speech of the speakers. For example, the storage device 12 may include first speaker information including information about the characteristics of speech of the first speaker and second speaker information including information about the characteristics of speech of the second speaker.

[0063] The speaker information 121c may include information on the results of a voice detection operation performed based on a voice signal indicating a voice uttered by a certain speaker, as information on the characteristics of the voice uttered by the speaker. For example, the speaker information 121c may include at least one of information on the average length of the detected voice segments (or other calculated value, the same applies below), information on the average length of the detected non-voice segments, information on the average number of characters uttered per unit time, information on the average number of words uttered per unit time, and information on the speaking rate.

[0064] The threshold setting unit 113c may identify a speaker from which the voice signal input to the voice detection device 1c was acquired, acquire speaker information 121c corresponding to the identified speaker from the storage device 12, and set the threshold value TH based on the acquired speaker information 121c. For example, the threshold setting unit 113c may set the threshold value TH to a larger value as the average length of the voice segments indicated by the speaker information 121c increases, so that relatively long voice segments are detected. For example, the threshold setting unit 113c may set the threshold value TH to the average length of the non-voice segments indicated by the speaker information 121c or a value close to that average. For example, the threshold value TH may be set to a smaller value as the average number of characters indicated by the speaker information 121c increases, so that relatively short voice segments (as a result, voice segments that do not contain an excessively large number of characters) are detected. For example, the threshold value TH may be set to a smaller value so that the larger the average number of words indicated by the speaker information 121c, the shorter the speech interval (resulting in a speech interval that does not include an excessively large number of words) to be detected.For example, the faster the speaking rate indicated by the speaker information 121c, the smaller the threshold value TH may be set to a smaller value so that the larger the speech interval (resulting in a speech interval that does not include an excessively large number of characters) to be detected.

[0065] The voice detection device 1c in the fourth embodiment described above can achieve the same effects as those achieved by at least one of the voice detection device 1 in the second embodiment to the voice detection device 1b in the third embodiment described above. Furthermore, the voice detection device 1c can set a threshold value TH that matches the characteristics of the speech of the speaker. Therefore, the voice detection device 1c can more appropriately detect a voice section by taking into account differences in the characteristics of the speech of the speaker.

[0066] (5) Fifth embodiment Next, a fifth embodiment of the voice detection device, voice detection method, and recording medium will be described. Hereinafter, the fifth embodiment of the voice detection device, voice detection method, and recording medium will be described with reference to Fig. 14, using a voice detection device 1d to which the fifth embodiment of the voice detection device, voice detection method, and recording medium is applied. Fig. 14 is a block diagram showing the configuration of a voice detection device 1d in the fifth embodiment.

[0067] 14, the speech detection device 1d of the fifth embodiment differs from at least one of the speech detection devices 1 of the second embodiment to 1c of the fourth embodiment in that it includes a text generation unit 111d instead of the symbol generation unit 111. Other features of the speech detection device 1d may be the same as other features of at least one of the speech detection devices 1, 1b, and 1c.

[0068] The text generation unit 111d differs from the symbol generation unit 111, which generates symbol data using a CTC model, in that it generates text data representing speech uttered by a speaker as characters from a speech signal without using a CTC model. For example, the text generation unit 111d calculates the posterior probability of a character string using an acoustic model, a pronunciation dictionary, and a language model, and generates, as text data, sequence data of multiple texts constituting the character string with the highest posterior probability. Even in this case, the voice activity detection unit 112 may determine the start of a voice activity from the generated text data, and then determine the end of the voice activity by comparing the length Lb of the non-voice activity with a threshold TH. Furthermore, the threshold setting unit 113 may set the threshold TH based on the length Lt of the tentative voice activity. As a result, the above-mentioned effects can be achieved even when a CTC model is not used.

[0069] When the text generating unit 111d generates text data using a pronunciation dictionary (i.e., dictionary data), the threshold setting unit 113 may set the threshold TH based on the characteristics of the pronunciation dictionary. For example, when the pronunciation dictionary has the characteristic of generating text data that includes many Chinese characters, the threshold setting unit 113 may set the threshold TH to a value smaller than the standard value so that a relatively short speech section (as a result, a speech section that does not include an excessively large number of characters) is detected.

[0070] The voice detection device 1d in the fifth embodiment described above can achieve the same effects as those achieved by at least one of the voice detection devices 1 in the second embodiment to 1c in the fourth embodiment. Furthermore, the voice detection device 1d can set the threshold value TH based on a pronunciation dictionary. This allows the voice detection device 1d to more appropriately detect voice segments, taking into account differences in the habits of operations for converting a voice signal into text data.

[0071] (6) Supplementary notes The following additional notes are provided regarding the above-described embodiment. [Appendix 1] a start point determining means for determining the start point of a voice section including a voice appearing in the voice signal; an end determination means for determining the end of a speech section by determining whether the length of a non-speech section that appears after the start point is determined is equal to or greater than a threshold; a setting means for setting the threshold value based on characteristics of a tentative voice segment starting from the beginning; A voice detection device comprising: [Appendix 2] The characteristics of the provisional speech segment include a length of the provisional speech segment. 2. The speech detection device of claim 1. [Appendix 3] The setting means sets the threshold value such that the threshold value set when the length of the tentative voice section is a first length is larger than the threshold value set when the length of the tentative voice section is a second length longer than the first length. 3. The speech detection device of claim 2. [Appendix 4] The characteristics of the provisional speech section include at least one of the number of characters of the speech included in the provisional speech section, the number of words of the speech included in the provisional speech section, and the speaking rate of the speech included in the provisional speech section. 4. A speech detection device according to any one of claims 1 to 3. [Appendix 5] the speech detection device further comprises a generating means for generating symbol data including character symbols and blank symbols from the speech signal using a CTC (Connectionist Temporal Classification) model; the starting point determining means determines the starting point based on the symbol data; the termination determining means determines the termination based on the symbol data; The non-voice section includes a section in which the blank symbols appear consecutively. 5. A speech detection device according to any one of claims 1 to 4. [Appendix 6] The characteristics of the provisional speech segment include the number of character symbols included in the provisional speech segment. 6. The speech detection device of claim 5. [Appendix 7] the voice detection device further comprises a storage means for storing, for each speaker, speaker information relating to characteristics of the voice uttered by the speaker; The setting means identifies a speaker from which the voice signal is acquired, and sets the threshold value based on the speaker information corresponding to the identified speaker. 7. A speech detection device according to any one of claims 1 to 6. [Appendix 8] the voice detection device further comprises a conversion means for converting the voice signal into text data by analyzing the voice signal using dictionary data; The setting means sets the threshold value based on the characteristics of the dictionary data. 8. A speech detection device according to any one of claims 1 to 7. [Appendix 9] determining the start of a speech interval containing speech present in the speech signal; determining whether the length of a non-voice section that appears after the start point is determined is equal to or greater than a threshold value, thereby determining the end point of the voice section; setting the threshold value based on characteristics of a tentative speech segment starting from the beginning; A speech detection method comprising: [Appendix 10] A recording medium on which a computer program for causing a computer to execute a voice detection method is recorded, The speech detection method includes: determining the start of a speech interval containing speech present in the speech signal; determining whether the length of a non-voice section that appears after the start point is determined is equal to or greater than a threshold value, thereby determining the end point of the voice section; setting the threshold value based on characteristics of a tentative speech segment starting from the beginning; A recording medium including:

[0072] At least some of the components of each of the above-described embodiments can be appropriately combined with at least some of the other components of each of the above-described embodiments. Some of the components of each of the above-described embodiments may not be used. Furthermore, to the extent permitted by law, the disclosures of all documents (e.g., published patent applications) cited in this disclosure are incorporated by reference as part of the description of this disclosure.

[0073] This disclosure may be modified as appropriate within the scope of the claims and the technical idea that can be read from the entire specification. Voice detection devices, voice detection methods, and recording media incorporating such modifications are also included in the technical idea of this disclosure. [Explanation of symbols]

[0074] 1. Voice detection device 11 Arithmetic unit 111 Symbol Generation Unit 112 Voice Activity Detector 113 Threshold setting unit 1000 Voice Detection Device 1001 Starting point determination unit 1002 Termination determination unit 1003 Settings

Claims

1. a start point determining means for determining the start point of a voice section including a voice appearing in the voice signal; an end determination means for determining the end of a speech section by determining whether the length of a non-speech section that appears after the start point is determined is equal to or greater than a threshold; a setting means for setting the threshold value based on characteristics of a tentative voice segment starting from the beginning; Equipped with the characteristics of the provisional speech segment include a length of the provisional speech segment; The setting means sets the threshold value such that the threshold value set when the length of the tentative voice segment is a first length is larger than the threshold value set when the length of the tentative voice segment is a second length longer than the first length. Voice detection device.

2. The characteristics of the provisional speech section include at least one of the number of characters of the speech included in the provisional speech section, the number of words of the speech included in the provisional speech section, and the speaking rate of the speech included in the provisional speech section. The voice detection device according to claim 1 .

3. the speech detection device further comprises a generating means for generating symbol data including character symbols and blank symbols from the speech signal using a CTC (Connectionist Temporal Classification) model; the starting point determining means determines the starting point based on the symbol data; the termination determining means determines the termination based on the symbol data; The non-voice section includes a section in which the blank symbols appear consecutively.

3. The voice detection device according to claim 1 or 2.

4. The characteristics of the provisional speech segment include the number of character symbols included in the provisional speech segment. The voice detection device according to claim 3 .

5. the voice detection device further comprises a storage means for storing, for each speaker, speaker information relating to characteristics of the voice uttered by the speaker; The setting means identifies a speaker from which the voice signal is acquired, and sets the threshold value based on the speaker information corresponding to the identified speaker. A voice detection device according to any one of claims 1 to 4.

6. the voice detection device further comprises a conversion means for converting the voice signal into text data by analyzing the voice signal using dictionary data; The setting means sets the threshold value based on the characteristics of the dictionary data. A voice detection device according to any one of claims 1 to 5.

7. determining the start of a speech interval containing speech present in the speech signal; determining whether the length of a non-voice section that appears after the start point is determined is equal to or greater than a threshold value, thereby determining the end point of the voice section; setting the threshold value based on characteristics of a tentative speech segment starting from the beginning; Including, the characteristics of the provisional speech segment include a length of the provisional speech segment; The setting includes setting the threshold value such that the threshold value set when the length of the tentative voice segment is a first length is greater than the threshold value set when the length of the tentative voice segment is a second length longer than the first length. Voice detection methods.

8. A computer program causing a computer to perform a speech detection method, comprising: The speech detection method includes: determining the start of a speech interval containing speech present in the speech signal; determining whether the length of a non-voice section that appears after the start point is determined is equal to or greater than a threshold value, thereby determining the end point of the voice section; setting the threshold value based on characteristics of a tentative speech segment starting from the beginning; Including, the characteristics of the provisional speech segment include a length of the provisional speech segment; The setting includes setting the threshold value such that the threshold value set when the length of the tentative voice segment is a first length is greater than the threshold value set when the length of the tentative voice segment is a second length longer than the first length. Computer program.

Citation Information

Patent Citations

  • Voice / non-voice determination compensation device, voice / non-voice determination compensation method, voice / non-voice determination compensation program and its recording medium, and voice mixing device, voice mixing method, voice mixing program and its recording medium

    JP2008134565A

  • Voice recognition method and voice recognition device

    JP2017097330A

  • Speech detection device, speech detection method, and program

    WO2015059947A1

  • Speech segment detection device and method for detecting speech segment

    WO2016143125A1

  • Utterance segment detection device, utterance segment detection method, and program

    WO2021014612A1