Information processing device, detection method, and recording medium

By segmenting the sound signal and calculating the change value, setting the detection threshold, and combining it with voice degree analysis, the problem of detection accuracy when the noise power rises sharply is solved, and high-precision voice interval detection is achieved.

CN114746939BActive Publication Date: 2025-09-30MITSUBISHI ELECTRIC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980102693.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-13
Publication Date
2025-09-30
Estimated Expiration
2039-12-13

AI Technical Summary

Technical Problem

The existing technology has difficulty in detecting the detection object, such as the speech interval, with high precision when the noise power increases sharply.

Method used

By dividing the sound signal into multiple intervals, calculating the variation value and power of each interval, setting the detection threshold to the interval above the maximum value as the detection target interval, using GMM or DNN to calculate the voice degree, and selecting the appropriate interval for detection.

Benefits of technology

The high-precision detection of speech intervals is achieved when the noise power rises sharply, reducing the false detection of noise intervals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114746939B_ABST
    Figure CN114746939B_ABST
Patent Text Reader

Abstract

The information processing device (100) comprises: an acquisition unit (110) for acquiring a sound signal; and a control unit (120) for dividing the sound signal into a plurality of intervals, calculating a variation amount per interval time of each of the plurality of intervals, i.e., a variation value, based on the sound signal, determining an interval in which the variation value is below a predetermined threshold value among the plurality of intervals, calculating the power of the sound signal in the determined interval based on the sound signal, determining a maximum value from the power of the sound signal in the determined interval, setting a value based on the maximum value as a detection threshold, and detecting an interval above the detection threshold as a detection target interval in the power of the sound signal over time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing device, a detection method, and a recording medium having a detection program recorded thereon. Background Art

[0002] Speech recognition technology is well known. For example, a technology for performing speech recognition on a speech section in a speech signal has been proposed (see Patent Document 1).

[0003] Prior art literature

[0004] Patent Literature

[0005] Patent Document 1: Japanese Patent Application Laid-Open No. 10-288994 Summary of the Invention

[0006] Problems to be solved by the invention

[0007] However, sometimes it's desirable to detect an object from a sound signal. For example, consider a method that uses a threshold based on noise power to detect an object. However, noise power can sometimes rise dramatically. If the noise power exceeds the threshold, this method cannot accurately detect the object.

[0008] An object of the present invention is to enable detection of a detection object with high accuracy.

[0009] Means for solving problems

[0010] An information processing device according to one embodiment of the present invention is provided. The information processing device includes: an acquisition unit that acquires a sound signal; and a control unit that divides the sound signal into a plurality of intervals, calculates a change value (i.e., an amount of change per interval time) in each of the plurality of intervals based on the sound signal, identifies an interval in which the change value is below a predetermined threshold value from the plurality of intervals, calculates the power of the sound signal in the identified interval based on the sound signal, determines a maximum value among the powers of the sound signal in the identified intervals, sets a value based on the maximum value as a detection threshold, and detects, as a detection target interval, intervals above the detection threshold value in the power of the sound signal over time.

[0011] Effects of the Invention

[0012] According to the present invention, a detection target can be detected with high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a diagram showing the hardware configuration of the information processing device according to the first embodiment.

[0014] Figure 2 It is a diagram showing a comparative example.

[0015] Figure 3 This is a block diagram showing the functions of the information processing device according to the first embodiment.

[0016] Figure 4 This is a flowchart showing an example of processing executed by the information processing apparatus according to the first embodiment.

[0017] Figure 5 A specific example of the processing executed by the information processing apparatus according to the first embodiment will be described.

[0018] Figure 6 This is a block diagram showing the functions of the information processing device according to the second embodiment.

[0019] Figure 7 This is a flowchart showing an example of processing executed by the information processing apparatus according to the second embodiment.

[0020] Figure 8 A specific example of processing executed by the information processing apparatus according to the second embodiment will be described.

[0021] Figure 9 This is a block diagram showing the functions of the information processing device according to the third embodiment.

[0022] Figure 10 This is a flowchart showing an example of processing executed by the information processing device according to the third embodiment.

[0023] Figure 11 This is a block diagram showing the functions of the information processing device according to the fourth embodiment.

[0024] Figure 12 This is a flowchart (part 1) showing an example of processing executed by the information processing apparatus according to the fourth embodiment.

[0025] Figure 13 This is a flowchart (part 2) showing an example of processing executed by the information processing apparatus according to the fourth embodiment.

[0026] Figure 14 A specific example (part one) of the processing executed by the information processing apparatus according to the fourth embodiment will be described.

[0027] Figure 15 A specific example (part 2) of the processing executed by the information processing apparatus according to the fourth embodiment is shown.

[0028] Figure 16 This is a flowchart showing a modified example of the fourth embodiment (part 1).

[0029] Figure 17 This is a flowchart showing a modified example of the fourth embodiment (part 2).

[0030] Figure 18This is a block diagram showing the functions of the information processing device according to the fifth embodiment.

[0031] Figure 19 This is a flowchart (part 1) showing an example of processing executed by the information processing apparatus according to the fifth embodiment.

[0032] Figure 20 This is a flowchart (part 2) showing an example of processing executed by the information processing apparatus according to the fifth embodiment.

[0033] Figure 21 A specific example (part one) of the processing executed by the information processing apparatus according to the fifth embodiment will be described.

[0034] Figure 22 A specific example (part 2) of the processing executed by the information processing apparatus according to the fifth embodiment is shown. DETAILED DESCRIPTION

[0035] The following embodiments are merely examples and various modifications are possible within the scope of the present invention.

[0036] Implementation Method 1

[0037] Figure 1 1 is a diagram showing the hardware configuration of an information processing device according to Embodiment 1. Information processing device 100 is a device that executes the detection method and includes a processor 101 , a volatile storage device 102 , and a nonvolatile storage device 103 .

[0038] Processor 101 controls the entire information processing device 100. For example, processor 101 may be a CPU (Central Processing Unit), an FPGA (Field Programmable Gate Array), or the like. Processor 101 may also be a multiprocessor. Information processing device 100 may be implemented using a processing circuit, software, firmware, or a combination thereof. Furthermore, the processing circuit may be a single circuit or a composite circuit.

[0039] The volatile storage device 102 is the main storage device of the information processing device 100. For example, the volatile storage device 102 is a RAM (Random Access Memory). The non-volatile storage device 103 is an auxiliary storage device of the information processing device 100. For example, the non-volatile storage device 103 is an HDD (Hard Disk Drive) or an SSD (Solid State Drive).

[0040] Figure 2 It is a diagram showing a comparative example. Figure 2 The upper part of the graph shows the waveform of the sound. Figure 2 The curve graph of the sound signal of the upper part of the sound is Figure 2 The lower part of . Figure 2 A range of 900 represents noise.

[0041] Sometimes it is desirable to detect the object from the sound signal. Figure 2 In this example, the detection target is speech. Here, the power of noise is often lower than that of speech. Therefore, a method using a threshold to detect speech is considered. Figure 2 A threshold value 901 is shown. For example, a section that is greater than or equal to the threshold value 901 is detected as a detection target section. In other words, the detection target section is detected as a speech section.

[0042] Here, sometimes the power of the noise rises sharply. For example, Figure 2 It shows that the power of the noise rises sharply after time t90. For example, Figure 2 The range 902 represents noise. When the power of the noise increases rapidly, the section after time t90 is detected as the detection target section. Figure 2 This shows a case where the power of noise exceeds a threshold value, and therefore noise is also a detection target in addition to speech.

[0043] In this way, Figure 2 Therefore, the following describes a method that can detect the detection object with high accuracy.

[0044] Figure 3 This is a block diagram illustrating the functions of the information processing device according to Embodiment 1. The information processing device 100 includes an acquisition unit 110 , a control unit 120 , and an output unit 130 .

[0045] Part or all of the acquisition unit 110, control unit 120, and output unit 130 may also be implemented by the processor 101. Part or all of the acquisition unit 110, control unit 120, and output unit 130 may also be implemented as modules of a program executed by the processor 101. For example, the program executed by the processor 101 is also called a detection program. For example, the detection program is recorded on a recording medium.

[0046] The acquisition unit 110 acquires a voice signal. For example, the voice signal may be the sound in a conference room where a meeting is being held, a telephone conversation, etc. Alternatively, the voice signal may be a signal based on recorded data, for example.

[0047] The control unit 120 calculates the power of the sound signal over time based on the sound signal. In other words, the control unit 120 calculates the power of the sound signal over time based on the sound signal. Hereinafter, the power of the sound signal will be referred to as the sound signal power. Alternatively, the sound signal power may be calculated by a device other than the information processing device 100.

[0048] The control unit 120 divides the sound signal into multiple intervals. The control unit 120 may divide the sound signal equally or unequally. The control unit 120 calculates a variation value for each of the multiple intervals based on the sound signal. The variation value represents the amount of variation per interval. Alternatively, the variation value represents the amount of variation in the power of the sound signal per interval. The interval time represents the time corresponding to one interval.

[0049] The control unit 120 determines an interval in which the variation value is below a predetermined threshold value from among a plurality of intervals. The control unit 120 calculates the power of the sound signal in the determined interval based on the sound signal. That is, the control unit 120 calculates the power of the sound signal in the determined interval based on the sound signal. The control unit 120 determines the maximum value from the power of the sound signal in the determined interval. The control unit 120 sets the value based on the maximum value as the detection threshold. In other words, the control unit 120 sets the value above the maximum value as the detection threshold. For example, the control unit 120 sets the value obtained by adding a predetermined value to the maximum value as the detection threshold. The control unit 120 detects an interval above the detection threshold in the sound signal power as a detection target interval.

[0050] The output unit 130 outputs information indicating the detection target section. For example, the output unit 130 outputs the information indicating the detection target section to a display. Furthermore, for example, the output unit 130 outputs the information indicating the detection target section to an external device connectable to the information processing device 100. Furthermore, for example, the output unit 130 outputs the information indicating the detection target section to a paper medium via a printer.

[0051] Next, the processing executed by the information processing apparatus 100 will be described using a flowchart.

[0052] Figure 4 This is a flowchart showing an example of processing executed by the information processing apparatus according to the first embodiment.

[0053] (Step S11 ) The acquisition unit 110 acquires a sound signal.

[0054] (Step S12) The control unit 120 divides the audio signal into frame units and calculates the power for each frame. For example, a frame is 10 msec.

[0055] That is, in the process of step S12 , the audio signal power is calculated, and thereby, for example, the audio signal power can be expressed in a graph.

[0056] (Step S13) The control unit 120 divides the audio signal into a plurality of intervals. For example, the control unit 120 may divide the audio signal power represented by a graph into a plurality of intervals. The plurality of frames in step S12 belong to one interval.

[0057] (Step S14) The control unit 120 calculates the variation value for each interval based on the audio signal. Alternatively, the control unit 120 may calculate the variance value for each interval based on the audio signal.

[0058] The calculation of the variance value is explained. First, the power m of the audio signal in the interval is calculated using equation (1). P is the power. i is the frame number. Also, i is a value from 1 to N.

[0059]

Mathematical formula 1

[0060]

[0061] Then, the variance value v is calculated using formula (2).

[0062]

Mathematical formula 2

[0063]

[0064] (Step S15) The control unit 120 identifies a section where the variation value is equal to or smaller than a preset threshold value. When the variance value is calculated, the control unit 120 identifies a section where the variance value is equal to or smaller than a preset threshold value.

[0065] (Step S16 ) The control unit 120 calculates the power of the audio signal in the specified section using equation (1).

[0066] (Step S17) The control unit 120 specifies the maximum power among the powers calculated for each interval and sets a value equal to or greater than the maximum power as a detection threshold.

[0067] (Step S18 ) The control unit 120 detects a section of the audio signal power that is equal to or greater than the detection threshold as a speech section.

[0068] (Step S19) The output unit 130 outputs information indicating the speech section. For example, the output unit 130 outputs the start time and the end time of the speech section.

[0069] Figure 5 A specific example of the processing executed by the information processing apparatus according to the first embodiment will be described. Figure 5 Graph showing the sound signal power 11 calculated by the control unit 120. For example, Figure 5 The vertical axis of the graph is dB. Figure 5 The horizontal axis of the graph is time. Figure 5 It is shown that the power of the noise increases sharply after time t1.

[0070] also, Figure 5 The graph shows the voice degree 12. The voice degree is described in the second embodiment.

[0071] For example, the control unit 120 divides the sound signal power 11 into multiple intervals. The control unit 120 calculates a variation value for each interval. The control unit 120 identifies intervals in which the variation value is below a predetermined threshold. For example, the control unit 120 identifies intervals 13a to 13e in which the variation value is below a predetermined threshold. Thus, for example, interval 14 is excluded. In addition, interval 14 is a speech interval. Thus, the control unit 120 identifies intervals other than the speech interval. In other words, the control unit 120 identifies noise intervals. Here, in the following description, it is assumed that intervals 13a to 13e are identified.

[0072] The control unit 120 calculates the power of the intervals 13a to 13e using equation (1). The control unit 120 determines the maximum power among the powers of the intervals 13a to 13e. The control unit 120 sets a value equal to or greater than the maximum power as a detection threshold. Figure 5 A detection threshold of 15 is shown.

[0073] The control unit 120 detects a section of the audio signal power 11 that is equal to or greater than the detection threshold 15 as a speech section. For example, the control unit 120 detects section 14. The output unit 130 outputs information indicating the speech section.

[0074] According to Embodiment 1, even when the noise power rises sharply, information processing device 100 sets the detection threshold to a value greater than the noise power. Therefore, information processing device 100 does not detect noise intervals as detection target intervals. For example, information processing device 100 does not detect intervals 13a to 13e. Furthermore, information processing device 100 detects speech intervals. This allows information processing device 100 to accurately detect speech as detection target.

[0075] Implementation Method 2

[0076] Next, the second embodiment will be described. In the second embodiment, the matters different from the first embodiment will be mainly described. In addition, in the second embodiment, the description of matters common to the first embodiment will be omitted. In the description of the second embodiment, reference will be made to Figure 1 、 3 .

[0077] Figure 6 : is a block diagram showing the functions of the information processing device according to the second embodiment. Figure 3 The same structure as shown Figure 6 Structural annotation and Figure 3 The same reference numerals are shown.

[0078] The information processing device 100a includes a control unit 120a. The control unit 120a will be described later.

[0079] Figure 7 This is a flowchart showing an example of processing executed by the information processing apparatus according to the second embodiment.

[0080] (Step S21 ) The acquisition unit 110 acquires a sound signal.

[0081] (Step S22) The control unit 120a divides the audio signal into frames and calculates the power of each frame. In other words, the control unit 120a calculates the audio signal power.

[0082] (Step S23) The control unit 120a divides the audio signal into frames and calculates the voice degree for each frame. The voice degree is the degree of voice similarity. For example, the control unit 120a calculates the voice degree using a Gaussian mixture model (GMM) or a deep neural network (DNN).

[0083] (Step S24) The control unit 120a divides the audio signal into a plurality of intervals. For example, the control unit 120a may divide the audio signal power into a plurality of intervals.

[0084] (Step S25) The control unit 120a calculates the variation value and the voice level for each interval based on the audio signal. For example, the control unit 120a calculates the variation value and the voice level for the first interval among the multiple intervals. In this way, the control unit 120a calculates the variation value and the voice level for the same interval.

[0085] Here, the calculation of the voice level of a section will be described. For example, the control unit 120a calculates the average of the voice levels of a plurality of frames belonging to a section as the voice level of the section. The control unit 120a similarly calculates the voice level for each section.

[0086] In this way, the control unit 120a calculates the voice degree of each of the plurality of intervals based on the audio signal. Specifically, the control unit 120a calculates the voice degree of each of the plurality of intervals based on the audio signal using a predetermined method such as GMM or DNN.

[0087] (Step S26) The control unit 120a specifies a section in which the variation value is equal to or less than a preset threshold value and the voice level is equal to or less than a voice level threshold value from among the plurality of sections.

[0088] (Step S27 ) The control unit 120 a calculates the power of the audio signal in the specified section using equation (1).

[0089] (Step S28) The control unit 120a specifies the maximum power among the powers calculated for each interval, and sets a value equal to or greater than the maximum power as a detection threshold.

[0090] (Step S29) The control unit 120a detects a section of the audio signal power that is equal to or greater than the detection threshold as a speech section.

[0091] (Step S30 ) The output unit 130 outputs information indicating the speech section.

[0092] Figure 8 A specific example of processing executed by the information processing apparatus according to the second embodiment will be described. Figure 8 Graph showing the audio signal power 21 calculated by the control unit 120a. Figure 8 The graph of the voice degree 22 calculated by the control unit 120a is shown. Figure 8 2 shows a state where the graph of the voice signal power 21 and the graph of the voice level 22 are mixed. The graph of the voice signal power 21 and the graph of the voice level 22 may be separated. Figure 8 The horizontal axis shows time.

[0093] Here, for example, with Figure 8 The 0 corresponding to the voice degree shown on the vertical axis means that the degree of voice similarity is about 50%. Therefore, for example, the interval corresponding to the voice degree value greater than 0 can also be considered as a voice interval. In addition, for example, the interval corresponding to the voice degree value less than 0 can also be considered as a noise interval.

[0094] The control unit 120a divides the audio signal power 21 into a plurality of intervals, calculates a variation value for each interval, and calculates a voice level for each interval.

[0095] The control unit 120a determines a section in which the variation value is less than or equal to a preset threshold value and the voice level is less than or equal to the voice level threshold value. Here, the section in which the voice level is less than or equal to the voice level threshold value will be described. Figure 8 The speech level threshold 23 is shown. For example, the intervals where the speech level is below the speech level threshold 23 are intervals 24a to 24e. For example, the intervals where the variation value is below this threshold and the speech level is below the speech level threshold 23 are intervals 25a to 25e. In the following description, it is assumed that intervals 25a to 25e are determined.

[0096] The control unit 120a calculates the power of the intervals 25a to 25e using equation (1). The control unit 120a determines the maximum power among the powers of the intervals 25a to 25e. The control unit 120a sets a value equal to or greater than the maximum value as a detection threshold. Figure 8 A detection threshold 26 is shown.

[0097] The control unit 120a detects, as a speech segment, a segment having a detection threshold value 26 or higher in the audio signal power 21. The output unit 130 outputs information indicating the speech segment.

[0098] According to the second embodiment, the information processing device 100a uses the voice degree, thereby preventing a section in which the voice signal power of a voice such as "ah" is constant from being mistakenly regarded as a noise section.

[0099] Implementation 3

[0100] Next, the third embodiment will be described. In the third embodiment, the matters different from the first and second embodiments will be mainly described. In the third embodiment, the matters similar to the first and second embodiments will be omitted. Figure 1 、 3 、7.

[0101] Figure 9 : is a block diagram showing the functions of the information processing device of embodiment 3. Figure 3 The same structure as shown Figure 9 Structural annotation and Figure 3 The same reference numerals are shown.

[0102] The information processing device 100b includes a control unit 120b. The control unit 120b will be described later.

[0103] Figure 10 This is a flowchart showing an example of processing executed by the information processing device according to the third embodiment.

[0104] exist Figure 10 In the processing, Figure 7 The difference in the processing is that steps S26a, 26b, 27a, and 28a are executed. Figure 10 In the following, steps S26a, 26b, 27a, and 28a are described. Figure 10 The other steps in Figure 7 The steps S21 to S25 and steps S29 and S30 are executed by the control unit 120b.

[0105] (Step S26a) The control unit 120b specifies an interval in which the change value is equal to or smaller than a preset threshold value among the plurality of intervals.

[0106] (Step S26b) The control unit 120b arranges the speech degrees of the specified interval in ascending order. In addition, the speech degrees of the specified interval are calculated in step S25.

[0107] The control unit 120b selects a predetermined number of intervals in ascending order. Hereinafter, the predetermined number is expressed as N. In addition, N is a positive integer.

[0108] In this way, the control unit 120b selects the upper N intervals in ascending order.

[0109] (Step S27a) The control unit 120b calculates the power of the audio signal in the top N intervals based on the audio signal. Specifically, the control unit 120b calculates the power of the audio signal in the top N intervals using equation (1).

[0110] (Step S28a) The control unit 120b determines the maximum value among the power values ​​of the audio signal in the top N intervals, and sets a value equal to or greater than the maximum value as the detection threshold.

[0111] Here, as in Implementation 2, a speech intensity threshold is set to detect one or more intervals. However, it is possible that one or more intervals may not be detected depending on the value of the speech intensity threshold or the speech intensity. In this case, Implementation 3 is effective. According to Implementation 3, N intervals are selected. Furthermore, information processing device 100b detects speech intervals in step S29. This allows information processing device 100b to detect the target speech with high accuracy.

[0112] Implementation 4

[0113] Next, the fourth embodiment will be described. In the fourth embodiment, the matters different from the first embodiment will be mainly described. In the fourth embodiment, the matters similar to the first embodiment will be omitted. Figure 1 、 3 .

[0114] Figure 11 : is a block diagram showing the functions of the information processing device of embodiment 4. Figure 3 The same structure as shown Figure 11 Structural annotation and Figure 3 The same reference numerals are shown.

[0115] The information processing device 100c includes a control unit 120c. The control unit 120c will be described later.

[0116] Figure 12 This is a flowchart (part 1) showing an example of processing executed by the information processing apparatus according to the fourth embodiment.

[0117] (Step S31 ) The acquisition unit 110 acquires a sound signal.

[0118] (Step S32) The control unit 120c divides the audio signal into frames and calculates the power of each frame. In other words, the control unit 120c calculates the audio signal power.

[0119] (Step S33) The control unit 120c divides the audio signal into a plurality of intervals. For example, the control unit 120c may divide the audio signal power into a plurality of intervals.

[0120] (Step S34 ) The control unit 120 c calculates a variation value for each interval based on the audio signal.

[0121] (Step S35 ) The control unit 120 c specifies an interval in which the variation value is equal to or smaller than a preset threshold value among the plurality of intervals.

[0122] (Step S36) The control unit 120c calculates the power of the audio signal in the specified section using equation (1). The control unit 120c then advances the process to step S41.

[0123] Figure 13 This is a flowchart (part 2) showing an example of processing executed by the information processing apparatus according to the fourth embodiment.

[0124] (Step S41 ) The control unit 120 c selects one section from the sections specified in step S35 .

[0125] (Step S42) The control unit 120c sets the power of the audio signal in the selected section to be equal to or higher than the power of the audio signal in the selected section as a temporary detection threshold. In addition, in step S36, the power of the audio signal in the selected section is calculated.

[0126] (Step S43) The control unit 120c detects the number of intervals in the audio signal power that are equal to or greater than the temporary detection threshold.

[0127] (Step S44) The control unit 120c determines whether all the sections specified in step S35 are selected. If all the sections are selected, the control unit 120c advances the process to step S45. If there are any unselected sections, the control unit 120c advances the process to step S41.

[0128] Thus, the control unit 120c sets a value based on the power of the audio signal in each section determined in step S35 as a temporary detection threshold, and detects the number of sections where the audio signal power is equal to or greater than the temporary detection threshold.

[0129] (Step S45 ) The control unit 120 c detects, as the detection threshold, the temporary detection threshold when the number of sections detected in step S43 is the largest, from among the temporary detection thresholds set for each section determined in step S35 .

[0130] (Step S46) The control unit 120c detects the interval detected using the temporary detection threshold detected in step S45 as the speech interval. In other words, the control unit 120c detects the interval detected using the detection threshold as the speech interval.

[0131] (Step S47 ) The output unit 130 outputs information indicating the speech section.

[0132] Figure 14 A specific example (part one) of the processing executed by the information processing apparatus according to the fourth embodiment will be described. Figure 14 A graph showing the audio signal power 31 calculated by the control unit 120 c is shown. Figure 14 The sections 32a to 32e determined by the control unit 120c in step S35 are shown.

[0133] The control unit 120c selects the interval 32a from the intervals 32a to 32e and sets the power of the interval 32a or higher as the temporary detection threshold. Figure 14 The set temporary detection threshold 33 is shown. The control unit 120c detects a section of the audio signal power 31 that is equal to or greater than the temporary detection threshold 33. For example, the control unit 120c detects sections A1 to A3. That is, the control unit 120c detects three sections.

[0134] Figure 15 A specific example (Part 2) of the process executed by the information processing device according to Embodiment 4 is shown. Next, the control unit 120c selects the interval 32b. The control unit 120c sets the power level in the interval 32b or higher as a temporary detection threshold. Figure 15 The set temporary detection threshold 34 is shown. The control unit 120c detects intervals of the audio signal power 31 that are greater than the temporary detection threshold 34. For example, the control unit 120c detects intervals B1 to B21. That is, the control unit 120c detects 21 intervals.

[0135] The control unit 120c also performs the same process for the sections 32c to 32e.

[0136] The control unit 120c detects the temporary detection threshold value when the number of sections detected in step S43 is the largest. The control unit 120c detects the section detected using the temporary detection threshold value detected in step S45 as the speech section.

[0137] According to Embodiment 4, information processing device 100c uses multiple temporary detection thresholds to detect speech segments. In other words, information processing device 100c varies the temporary detection thresholds to detect speech segments. For example, varying the temporary detection thresholds can improve the accuracy of speech segment detection compared to uniquely determining the detection thresholds as in Embodiment 1.

[0138] The reason for using the noise power with the highest number of detected segments as the final detection result is that, when the noise power (i.e., noise power) is inappropriate, the number of detected segments will be reduced compared to the actual number of speech segments. Specifically, when the noise power is inappropriately low, multiple speech segments will be detected as a single segment, resulting in a reduction in the number of detected segments. On the other hand, when the noise power is inappropriately high, speech segments with lower power will be missed, resulting in a reduction in the number of detected segments.

[0139] Variation of Implementation 4

[0140] Next, a modification of the fourth embodiment will be described.

[0141] Figure 16 This is a flowchart showing a modified example of the fourth embodiment (part 1). Figure 16 In the processing, Figure 12 The difference in the processing is that steps S32a, 34a, 35a, and 36a are executed. Figure 16 In the following, steps S32a, 34a, 35a, and 36a are described. Figure 16 The other steps in Figure 12 The steps are numbered the same as those in the previous step, and thus the description of the processing is omitted.

[0142] (Step S32a) The control unit 120c divides the audio signal into frame units and calculates the voice degree for each frame.

[0143] (Step S34a) The control unit 120c calculates the variation value and the voice degree for each section based on the audio signal.

[0144] (Step S35a) The control unit 120c arranges the voice degrees of the specified intervals in ascending order. The control unit 120c selects the top N intervals in ascending order.

[0145] (Step S36a) The control unit 120c calculates the power of the audio signal in the top N intervals based on the audio signal. Specifically, the control unit 120c calculates the power of the audio signal in the top N intervals using equation (1). The control unit 120c then advances the process to step S41a.

[0146] Figure 17This is a flowchart showing a modified example of the fourth embodiment (part 2). Figure 17 In the processing, Figure 13 The difference in the processing is that steps S41a, 42a, and 44a are executed. Figure 17 In the following, steps S41a, 42a, and 44a are described. Figure 17 The other steps in Figure 13 The steps are numbered the same as those in the previous step, and thus the description of the processing is omitted.

[0147] (Step S41a) The control unit 120c selects one interval from the top N intervals.

[0148] (Step S42a) The control unit 120c sets the power of the audio signal in the selected section to be equal to or higher than the power of the audio signal in the selected section as a temporary detection threshold. In addition, in step S36a, the power of the audio signal in the selected section is calculated.

[0149] (Step S44a) The control unit 120c determines whether the top N intervals are selected. If the top N intervals are selected, the control unit 120c advances the process to step S45. If there are unselected intervals, the control unit 120c advances the process to step S41a.

[0150] In this way, the control unit 120c sets a value based on the power of the audio signal in each of the upper N intervals as a temporary detection threshold, and detects the number of intervals with audio signal power equal to or greater than the temporary detection threshold.

[0151] According to the modification of Embodiment 4, the information processing device 100c can improve the accuracy of detecting a speech section.

[0152] Implementation 5

[0153] Next, the fifth embodiment will be described. In the fifth embodiment, the matters different from the first embodiment will be mainly described. In the fifth embodiment, the matters similar to the first embodiment will be omitted. Figure 1 、 3 .

[0154] In Embodiments 1 to 4, the case of detecting a speech section as a detection target section has been described. In Embodiment 5, the case of detecting a non-stationary noise section as a detection target section will be described.

[0155] Figure 18 : is a block diagram showing the functions of the information processing device of embodiment 5. Figure 3 The same structure as shown Figure 18 Structural annotation and Figure 3 The same reference numerals are shown.

[0156] The information processing device 100d includes a control unit 120d and an output unit 130d. The control unit 120d and the output unit 130d will be described later.

[0157] Figure 19 This is a flowchart (part 1) showing an example of processing executed by the information processing apparatus according to the fifth embodiment.

[0158] (Step S51 ) The acquisition unit 110 acquires a sound signal.

[0159] (Step S52) The control unit 120d divides the audio signal into frames and calculates the power of each frame. In other words, the control unit 120d calculates the audio signal power.

[0160] (Step S53) The control unit 120d divides the audio signal into frames and calculates the speech level for each frame. In other words, the control unit 120d calculates the speech level over time using a predetermined method such as GMM or DNN and the audio signal. Here, the speech level over time can also be expressed as a time-series speech level.

[0161] (Step S54) The control unit 120d determines a section where the voice level is greater than or equal to the voice level threshold. Thus, the control unit 120d determines the voice section. If the voice section is not determined, the control unit 120d may lower the voice level threshold.

[0162] (Step S55) The control unit 120d specifies a section other than the specified section. Thus, the control unit 120d specifies a candidate for a non-stationary noise section.

[0163] Alternatively, the control unit 120d may perform the following process instead of steps S54 and S55. The control unit 120d identifies a section where the voice level is less than the voice level threshold. Thus, the control unit 120d identifies a candidate non-stationary noise section. The control unit 120d then proceeds to step S61.

[0164] Figure 20 This is a flowchart (part 2) showing an example of processing performed by the information processing apparatus of embodiment 5. In the following description, it is assumed that one unstable noise interval candidate is determined. In addition, when multiple unstable noise interval candidates are determined, the process is repeated for the number of unstable noise interval candidates. Figure 20 processing.

[0165] (Step S61) The control unit 120d divides one unstable noise interval candidate into a plurality of intervals. The control unit 120d may divide the unstable noise interval candidate equally or unequally.

[0166] (Step S62 ) The control unit 120 d calculates the variation value of each of the plurality of intervals based on the audio signal.

[0167] (Step S63) The control unit 120d specifies an interval in which the variation value is equal to or smaller than a preset threshold value among the plurality of intervals.

[0168] (Step S64) The control unit 120d calculates the power of the audio signal in the specified interval based on the audio signal. Specifically, the control unit 120d calculates the power of the audio signal in the specified interval using equation (1).

[0169] (Step S65) The control unit 120d determines the maximum value from the power of the audio signal in the determined section, and sets a value equal to or greater than the maximum value as a detection threshold.

[0170] (Step S66) The control unit 120d detects a section within the candidate non-stationary noise section and having a sound signal power not less than the detection threshold as a non-stationary noise section.

[0171] (Step S67) The output unit 130d outputs information indicating the unstable noise interval as the detection target interval. For example, the output unit 130d outputs the start time and end time of the unstable noise interval.

[0172] Figure 21 A specific example (part one) of the processing executed by the information processing apparatus according to the fifth embodiment will be described. Figure 21 Graph showing the audio signal power 41 calculated by the control unit 120d. Figure 21 A graph showing the voice degree 42 is shown. Figure 21 A voice level threshold 43 is shown.

[0173] The control unit 120d determines a section in which the voice level is equal to or greater than the voice level threshold value 43. Figure 21 The determined interval, ie, the speech interval, is shown.

[0174] Figure 22 A specific example (Part 2) of the process executed by the information processing apparatus according to Embodiment 5 is shown. The control unit 120d specifies a section other than the specified section. Figure 22 The determined interval, ie, the non-stationary noise interval candidate is shown.

[0175] For example, the control unit 120d divides the unstable noise interval candidate 1 into multiple intervals. The control unit 120d calculates a variation value for each interval. The control unit 120d identifies intervals where the variation value is below a predetermined threshold. The control unit 120d calculates the power of the sound signal in the identified intervals. The control unit 120d determines the maximum power among the powers calculated for each interval. The control unit 120d sets a value above this maximum value as the detection threshold. The control unit 120d detects intervals within the unstable noise interval candidate 1 where the sound signal power 41 is above the detection threshold as unstable noise intervals.

[0176] Similarly, the information processing device 100d can detect the non-stationary noise section from the non-stationary noise section candidates 2 to 6.

[0177] According to Embodiment 5, information processing device 100d uses speech degrees, enabling stable speech detection. Furthermore, information processing device 100d targets non-speech intervals for non-stationary noise detection and sets a detection threshold for each candidate non-stationary noise interval, enabling high-precision non-stationary noise detection.

[0178] The features of the above-described embodiments can be combined with each other as appropriate.

[0179] Label Description

[0180] 11: Sound signal power; 12: Voice intensity; 13a-13e: Interval; 14: Interval; 15: Detection threshold; 21: Sound signal power; 22: Voice intensity; 23: Voice intensity threshold; 24a-24e: Interval; 25a-25e: Interval; 26: Detection threshold; 31: Sound signal power; 32a-32e: Interval; 33: Temporary detection threshold; 34: Temporary detection threshold; 41: Sound signal power; 42: Voice intensity; 43: Voice intensity threshold; 100, 100a, 100b, 100c, 100d: Information processing device; 101: Processor; 102: Volatile storage device; 103: Non-volatile storage device; 110: Acquisition unit; 120, 120a, 120b, 120c, 120d: Control unit; 130, 130d: Output unit; 900: Range; 901: Threshold; 902: Range.

Claims

1. An information processing device, comprising: an acquisition unit that acquires a sound signal; and A control unit that divides the sound signal into multiple intervals, calculates a change in the power of the sound signal at a time corresponding to each of the multiple intervals based on the sound signal, that is, a change value, calculates a degree of voice similarity, that is, a voice degree, for each of the multiple intervals based on the sound signal, determines an interval in which the change value is below a predetermined threshold and the voice degree is below a predetermined threshold from among the multiple intervals, calculates the power of the sound signal in the determined interval based on the sound signal, determines a maximum value from among the powers of the sound signal in the determined intervals, sets a value based on the maximum value as a detection threshold, and detects an interval above the detection threshold as a detection target interval in the power of the sound signal over time.

2. The information processing device according to claim 1, wherein The control unit detects the detection target section as a speech section.

3. The information processing device according to claim 1 or 2, wherein: The information processing device further includes an output unit that outputs information indicating the detected section.

4. An information processing device, comprising: an acquisition unit that acquires a sound signal; and A control unit that divides the sound signal into multiple intervals, calculates a change in the power of the sound signal at a time corresponding to each of the multiple intervals based on the sound signal, that is, a change value, calculates a degree of voice similarity for each of the multiple intervals based on the sound signal, determines an interval in which the change value is below a predetermined threshold value among the multiple intervals, arranges the voice degrees of the determined intervals in ascending order, selects a predetermined number of intervals in descending order, calculates the power of the sound signal in the selected intervals based on the sound signal, determines a maximum value from the powers of the sound signal in the selected intervals, sets a value based on the maximum value as a detection threshold, and detects an interval above the detection threshold as a detection target interval in the power of the sound signal over time.

5. The information processing apparatus according to claim 4, wherein: The control unit detects the detection target section as a speech section.

6. The information processing device according to claim 4 or 5, wherein: The information processing device further includes an output unit that outputs information indicating the detected section.

7. An information processing device, comprising: an acquisition unit that acquires a sound signal; and A control unit divides the sound signal into a plurality of intervals, calculates a change in the power of the sound signal (i.e., a change value) over time corresponding to each of the plurality of intervals based on the sound signal, calculates a degree of speech similarity (i.e., a voice degree) for each of the plurality of intervals based on the sound signal, determines an interval in which the change value is below a predetermined threshold value from among the plurality of intervals, arranges the voice degrees of the determined intervals in ascending order, selects a predetermined number of intervals in descending order, calculates the power of the sound signal in the selected intervals based on the sound signal, sets a value based on the power of the sound signal in the interval as a temporary detection threshold for each selected interval, detects the number of intervals above the set temporary detection threshold in the power of the sound signal over time, uses the temporary detection threshold at which the maximum number of detected intervals is detected from among the temporary detection thresholds set for each selected interval as the detection threshold, and defines the interval detected when detection is performed using the detection threshold as the detection target interval.

8. The information processing apparatus according to claim 7, wherein: The control unit detects the detection target section as a speech section.

9. The information processing device according to claim 7 or 8, wherein: The information processing device further includes an output unit that outputs information indicating the detected section.

10. A detection method, wherein: An information processing device obtains a sound signal, divides the sound signal into a plurality of intervals, calculates a change in the power of the sound signal at a time corresponding to each of the plurality of intervals based on the sound signal, that is, a change value, calculates a degree of voice similarity, that is, a voice degree, for each of the plurality of intervals based on the sound signal, determines an interval in which the change value is below a predetermined threshold and the voice degree is below a predetermined threshold from among the plurality of intervals, calculates the power of the sound signal in the determined interval based on the sound signal, determines a maximum value from among the powers of the sound signal in the determined intervals, sets a value based on the maximum value as a detection threshold, and detects an interval above the detection threshold as a detection target interval in the power of the sound signal over time.

11. A detection method, wherein: An information processing device obtains a sound signal, divides the sound signal into a plurality of intervals, calculates a change in the power of the sound signal at a time corresponding to each of the plurality of intervals based on the sound signal, that is, a change value, calculates a degree of speech similarity, that is, a voice degree, for each of the plurality of intervals based on the sound signal, determines an interval in which the change value is below a predetermined threshold value among the plurality of intervals, arranges the voice degrees of the determined intervals in ascending order, selects a predetermined number of intervals in descending order, calculates the power of the sound signal in the selected intervals based on the sound signal, determines a maximum value from among the powers of the sound signal in the selected intervals, sets a value based on the maximum value as a detection threshold, and detects, in the power of the sound signal over time, an interval above the detection threshold as a detection target interval.

12. A detection method, wherein: An information processing device obtains a sound signal, divides the sound signal into a plurality of intervals, calculates a change in the power of the sound signal (variation value) corresponding to the time of each of the plurality of intervals based on the sound signal, calculates a degree of speech similarity (speechiness) for each of the plurality of intervals based on the sound signal, identifies intervals in the plurality of intervals in which the change value is less than or equal to a predetermined threshold, arranges the speechiness of the identified intervals in ascending order, selects a predetermined number of intervals in descending order, calculates the power of the sound signal in the selected intervals based on the sound signal, sets a value based on the power of the sound signal in the interval as a temporary detection threshold for each selected interval, detects the number of intervals above the set temporary detection threshold in the power of the sound signal over time, sets the temporary detection threshold at which the maximum number of detected intervals is detected from the temporary detection thresholds set for each selected interval as the detection threshold, and sets the interval detected when detection is performed using the detection threshold as the detection target interval.

13. A recording medium having a detection program recorded thereon, the detection program causing an information processing device to execute the following processing: Acquire a sound signal, divide the sound signal into a plurality of intervals, calculate the amount of change in the power of the sound signal at the time corresponding to each of the plurality of intervals based on the sound signal, that is, the change value, calculate the degree of voice similarity of each of the plurality of intervals based on the sound signal, determine an interval in which the change value is below a predetermined threshold and the voice degree is below a predetermined threshold among the plurality of intervals, calculate the power of the sound signal in the determined interval based on the sound signal, determine a maximum value from the powers of the sound signal in the determined intervals, set a value based on the maximum value as a detection threshold, and detect an interval above the detection threshold as a detection target interval in the power of the sound signal over time.

14. A recording medium having a detection program recorded thereon, the detection program causing an information processing device to execute the following processing: Acquire a sound signal, divide the sound signal into a plurality of intervals, calculate the amount of change in the power of the sound signal corresponding to the time of each of the plurality of intervals based on the sound signal, that is, the change value, calculate the degree of voice similarity of each of the plurality of intervals based on the sound signal, determine the intervals in which the change value is below a predetermined threshold value among the plurality of intervals, arrange the voice degrees of the determined intervals in ascending order, select a predetermined number of intervals in descending order, calculate the power of the sound signal in the selected intervals based on the sound signal, determine a maximum value from the power of the sound signal in the selected intervals, set a value based on the maximum value as a detection threshold, and detect the interval above the detection threshold as the detection target interval in the power of the sound signal as time passes.

15. A recording medium having a detection program recorded thereon, the detection program causing an information processing device to execute the following processing: A sound signal is obtained, the sound signal is divided into a plurality of intervals, a change value (variation value) of the power of the sound signal corresponding to the time of each of the plurality of intervals is calculated based on the sound signal, a degree of speech similarity (speechiness) is calculated based on the sound signal for each of the plurality of intervals, an interval in which the change value is less than or equal to a predetermined threshold is determined from the plurality of intervals, the speechiness of the determined intervals is arranged in ascending order, a predetermined number of intervals are selected in descending order, the power of the sound signal in the selected intervals is calculated based on the sound signal, a value based on the power of the sound signal in the interval is set as a temporary detection threshold for each selected interval, the number of intervals exceeding the set temporary detection threshold is detected in the power of the sound signal over time, the temporary detection threshold at which the maximum number of detected intervals is detected is set as the detection threshold, and the interval detected when detection is performed using the detection threshold is set as the detection target interval.

Citation Information

Patent Citations

  • Noise level estimating method, speech section detecting method, speech recognizing method, speech section detecting device, and speech recognition device

    JP1998288994A

  • Audio processing device, method, program, and integrated circuit

    CN103380457A

  • Voice detecting device

    JP2001067092A