Commentary audio insertion timing learning device and program thereof, and commentary audio insertion timing detection device and program thereof
A neural network system identifies commentary audio insertion timing in program audio, allowing overlap without disrupting the program, enhancing insertion accuracy and speed.
Patent Information
- Application Number
- JP2021165031
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-06
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2041-10-06
AI Technical Summary
Existing methods struggle to insert commentary audio into program audio without overlap, and conventional techniques for detecting utterance ends can cause delays.
A neural network-based system that learns to identify insertion-permitted and prohibited sections in audio frames, using acoustic features to determine commentary audio insertion timing, allowing overlap without impairing sentence meaning.
Enables longer and quicker insertion of commentary audio without disrupting the program audio, improving timing detection accuracy and reducing delays.
Smart Images

Figure 0007730712000002 
Figure 0007730712000003 
Figure 0007730712000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a commentary audio insertion timing learning device and a program therefor that learn the insertion timing of commentary audio to be inserted as a secondary audio into a main audio, and a commentary audio insertion timing detection device and a program therefor that detects the insertion timing. [Background technology]
[0002] Currently, as a broadcasting service for the visually impaired, commentary broadcasting is being implemented, which provides audio explanations about the video in addition to the program audio on the main broadcast channel. Typically, commentary broadcasts are edited so that the commentary audio is placed in sections ("gaps") where there is no program audio, so that the program audio broadcast in the main audio and the commentary audio broadcast in the secondary audio do not overlap and can be heard simultaneously. A conventional method has been disclosed in which the end of an utterance is detected as the beginning of a pause by using short-term and long-term moving averages of pitch frequency, which is an acoustic feature (see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2020-64248 Summary of the Invention [Problem to be solved by the invention]
[0004] Many of the program scenes in which commentary audio needs to be inserted are likely to overlap with the program audio (hereinafter referred to as audio overlap). Although it may be possible to generate a short commentary audio and insert it between program audio segments, it is difficult to insert the commentary audio so that it does not overlap with the program audio at all. Even if the end of an utterance can be detected using the technique described in Patent Document 1, it cannot solve the difficult situation of inserting commentary audio into the "gaps" in program audio. Furthermore, the technique described in Patent Document 1 uses short-term and long-term moving averages, which may cause a delay in detecting the end of an utterance.
[0005] The present invention has been made in view of such problems, and aims to provide a commentary audio insertion timing learning device and a program therefor that can learn the timing of inserting commentary audio into the audio of a program, allowing audio overlap without impairing the meaning of the sentence, and a commentary audio insertion timing detection device and a program therefor that can detect and reduce delays in the insertion timing. [Means for solving the problem]
[0006] In order to solve the above problem, the commentary audio insertion timing learning device of the present invention is a commentary audio insertion timing learning device that learns a neural network type determination model for determining the type of each section of an audio frame in any audio from training data that sets types that distinguish between insertion prohibited sections that prohibit the insertion of commentary audio, overlapping insertion permitted sections that allow the insertion of commentary audio overlappingly just before the end of a speech section, and insertion permitted sections that allow the insertion of commentary audio in a non-speech section, and is configured to include an acoustic feature extraction means, a type determination means, an error calculation means, and a parameter update means.
[0007] In this configuration, the comment audio insertion timing learning device extracts, by the audio feature extracting means, audio features such as pitch frequency from the audio for each audio frame having a predetermined time length. Then, the commentary audio insertion timing learning device inputs the acoustic features extracted by the acoustic feature extraction means into a type determination model in time series using the type determination means, and calculates a probability value of the type of the audio frame as a type determination result. Furthermore, the comment audio insertion timing learning device calculates an error in the type determination by the error calculation means based on the probability value calculated by the type determination means and the teacher data. The comment audio insertion timing learning device then updates the parameters of the category determination model by the parameter update means based on the error calculated by the error calculation means. By updating the parameters in a direction that reduces the error, the category determination model is learned.
[0008] In this way, the commentary audio insertion timing learning device learns the overlapping insertion permission section in which commentary audio is permitted to be inserted overlappingly just before the end of the speech section, thereby making it possible to expand the section in which commentary audio is inserted without compromising the meaning of the program audio. The commentary audio insertion timing learning device can be operated by a commentary audio insertion timing learning program that causes a computer to function as each of the above-mentioned means.
[0009] In addition, in order to solve the above problem, the commentary audio insertion timing detection device of the present invention is a commentary audio insertion timing detection device that detects the timing to insert commentary audio from audio, and is configured to include an acoustic feature extraction means, a type determination means, and a type decision means.
[0010] In this configuration, the comment audio insertion timing detection device extracts an acoustic feature such as a pitch frequency from the audio for each audio frame having a predetermined time length using the audio feature extraction means. The commentary audio insertion timing detection device then uses a type determination means to input acoustic features in a time series into a pre-trained neural network type determination model that determines, for each section of an audio frame, an insertion-prohibited section that prohibits the insertion of commentary audio, an overlapping insertion-permitted section that allows the insertion of overlapping commentary audio immediately before the end of the speech section, and an insertion-permitted section that allows the insertion of commentary audio in a non-speech section, and calculates a probability value of the type of the audio frame as the type determination result. Then, the comment audio insertion timing detection device determines, by the type determination means, the type of the audio frame that has the maximum probability value calculated by the type determination means.
[0011] This allows the comment audio insertion timing detection device to detect an insertion section that extends the section into which comment audio is inserted without impairing the meaning of the program audio. Furthermore, since the commentary audio insertion timing detection device can detect the beginning of the insertion section in audio frame units, it can quickly detect the timing to insert the commentary audio even when inserting the commentary audio in real time. The commentary audio insertion timing detection device can be operated by a commentary audio insertion timing detection program that causes a computer to function as each of the above-mentioned means. [Effects of the Invention]
[0012] The present invention provides the following excellent effects. According to the present invention, it is possible to detect the timing for inserting commentary audio into the audio of a program, allowing audio overlap without impairing the meaning of the sentence. As a result, the present invention can secure a longer time for inserting commentary audio into a program than conventional methods, and can quickly detect the timing for inserting commentary audio when inserting commentary audio in real time. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a block diagram showing the configuration of a comment audio insertion timing learning device according to a first embodiment of the present invention. [Figure 2] 1 is an explanatory diagram for explaining the contents of training data to be input to a commentary audio insertion timing learning device according to a first embodiment of the present invention. FIG. [Figure 3] FIG. 10 is a network diagram showing an example of the configuration of a neural network of a type determination model that determines the type of a speech section from acoustic features. [Figure 4] FIG. 10 is an explanatory diagram for explaining output data of a type determination model. [Figure 5] 3 is a flowchart showing the operation of the comment audio insertion timing learning device according to the first embodiment of the present invention. [Figure 6] FIG. 10 is a block diagram showing the configuration of a comment audio insertion timing detection device according to a second embodiment of the present invention. [Figure 7] 10 is a flowchart showing the operation of the comment audio insertion timing detection device according to the second embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Although the present invention aims to add commentary to television audio, it can also be applied to early determination of the timing of a machine's response to a human when conversing with a robot, smart speaker, or other machine. This makes it possible to realize a conversation similar to that between a human and a machine, without leaving a slight "pause" when responding from the machine.
[0015] [Configuration of the commentary audio insertion timing learning device] First, with reference to FIG. 1, the configuration of a comment audio insertion timing learning device 1 according to a first embodiment of the present invention will be described.
[0016] The commentary audio insertion timing learning device 1 learns a neural network type determination model N for determining the type of each section of an audio frame in any audio from training data that sets types that distinguish between insertion-prohibited sections that prohibit the insertion of commentary audio, overlapping insertion-permitted sections that permit the insertion of overlapping commentary audio immediately before the end of a speech section, and insertion-permitted sections that permit the insertion of commentary audio in a non-speech section.
[0017] The audio V is data (digital data) including the audio of a person's speech section. Since the commentary audio insertion timing learning device 1 learns the timing to insert commentary audio into audio, accuracy improves if the audio V is spoken by the same person as the audio to be actually predicted, or by a similar person, but if the amount of data used is large, unspecified audio or audio from multiple speakers may be used.
[0018] The type label L is a label (numerical value, code, etc.) that indicates the type of audio frame (hereinafter simply referred to as frame) for each predetermined time interval corresponding to the audio V. This type label L is a label if it can identify, for each frame of a predetermined time interval (e.g., 10 ms [milliseconds]), at least an insertion prohibited interval that prohibits the insertion of commentary audio, an overlapping insertion permitted interval that allows the insertion of overlapping commentary audio immediately before the end of the speech interval, and an insertion permitted interval that allows the insertion of commentary audio in a non-speech interval. Note that the type label L is training data that has been generated in advance corresponding to the audio V.
[0019] Here, with reference to FIG. 2, a specific example of the type label L, which is training data, will be described. In FIG. 2, for ease of understanding, the speech V is represented by a speech waveform, and the speech text is provided for reference. Here, five types, type labels L1 to L5, are used to classify the speech V. Using the end position E of one speech section (all speech sections) as a reference, a type label L3 indicating "ending" is assigned to a section that goes back a predetermined time from the end position E. A type label L2 indicating "near the end" is assigned to a section that goes back a further predetermined time from the "ending" (type label L3). For example, the "ending" (type label L3) is a section that goes back 200 ms from the end position E, and the "near the end" (type label L3) is a section that goes back another 200 ms from the beginning of the "ending". This is because "near the end" and "ending" are often two moras (200 to 300 ms) each, but these are not strict times.
[0020] The type label L1 is a label assigned to the entire speech section ALL, excluding sections with type labels L2 and L3. The content of the speech can be roughly understood from the utterances with type label L1.
[0021] The type label L4 is a label assigned to the "pause" section from the end position E to the beginning of the next utterance. The type label L4 is assigned to a non-utterance section (for example, 200 ms or more) into which comment audio can be inserted. A type label L5 is assigned to a short non-speech section (e.g., less than 200 ms) into which commentary audio cannot be inserted. Note that even if there is enough time for commentary audio to be inserted, a non-speech section following a type label L1 is assigned type label L5. This is to prevent the system from mistakenly learning that the speech continuity is low and commentary audio can be inserted, even if the pitch frequency is high and the speech continuity is high, by assigning type label L4.
[0022] The type label L serving as training data is a label string in which types for classifying the audio V are set for each frame of a predetermined time length (for example, 10 ms). For example, the type label L is set to the numeric value "1" consecutively every 10 ms in the time section of type label L1 of audio V. Also, the type label L is set to the numeric value "2" consecutively every 10 ms in the time section of type label L2 of audio V. The same applies to type labels L3 to L5. That is, the meaning of the type label L is as shown in the table below.
[0023] [Table 1]
[0024] However, even if there is no speech (a predetermined time or longer), if the immediately preceding label is type label L1, it is set to type label L5. Sections with type labels L1, L2, and L5 correspond to insertion-prohibited sections that prohibit the insertion of commentary audio. Sections with type label L3 correspond to overlapping insertion-permitted sections that permit the insertion of commentary audio immediately before the end of a speech section. Sections with type label L4 correspond to insertion-permitted sections that permit the insertion of commentary audio into a non-speech section.
[0025] Returning to Figure 1, we continue the explanation. As shown in FIG. 1, the comment audio insertion timing learning device 1 includes an acoustic feature extraction unit 10, a storage unit 11, a type determination unit 12, an error calculation unit 13, and a parameter update unit 14.
[0026] The acoustic feature extraction means 10 extracts acoustic features for each frame from the speech V. Here, the acoustic feature extraction means 10 extracts acoustic features for each frame by sequentially performing acoustic analysis on the speech V at a predetermined frame length (e.g., 10 ms) and a predetermined shift width (e.g., 10 ms). For example, the acoustic feature extraction means 10 extracts a pitch frequency for each frame from the speech V. The acoustic feature extraction means 10 may extract, in addition to the pitch frequency, power, MFCC (Mel-Frequency Cepstrum Coefficients), a filter bank, etc. These acoustic features can be obtained by general acoustic analysis, and therefore detailed explanation of the analysis method will be omitted. The acoustic feature extraction means 10 outputs the extracted acoustic feature for each frame to the type determination means 12 .
[0027] The storage means 11 stores a neural network model. This storage means 11 can be configured with a general storage medium such as a semiconductor memory. Here, the storage means 11 stores a type determination model N. The type determination model N is a neural network (specifically, its parameters) that receives acoustic features in a time series and determines the type of speech for each frame. The type determination model N is a model to be trained, and initial values are set for the parameters in advance.
[0028] A predetermined number of acoustic features are input to the type determination model N. The acoustic features input to this type determination model N are input sequentially in time series, shifted one frame at a time. The output of the type determination model N is the output value of the node for the total number of predetermined type labels. The output of each node indicates a normalized probability value, the sum of which is "1".
[0029] The neural network structure of this classification determination model N may be any as long as it satisfies the above-mentioned data input and output requirements. Specifically, the classification determination model N can be realized by combining an LSTM (Long Short Term Memory), which is a type of RNN (Recurrent Neural Network) that handles time-series data, a fully connected layer, and a softmax function.
[0030] For example, the type determination model N can be realized by a neural network having an input layer IL, a hidden layer ML, and an output layer OL, as shown in FIG. The input layer IL is a layer that inputs acoustic features for each frame. Here, 300 acoustic features (f1 to f300) are input in time series, shifted by one frame at a time. The middle layer ML is a layer composed of two LSTM layers ML1, a forward LSTM and a backward LSTM, and a fully connected layer ML2, a forward propagation neural network (FFNN). The forward LSTM repeats LSTM calculations from the first input acoustic feature f1 to the last input acoustic feature f300. The backward LSTM repeats LSTM calculations from the last input acoustic feature f300 to the first input acoustic feature f1. The middle layer ML concatenates the vectors resulting from the calculations of the two LSTM layers ML1, and outputs the concatenated vector as an output vector via the fully connected layer ML2. The output layer OL is a layer that calculates the ratio (probability value) at the output node by adding weights to the values of each element of the output vector output from the middle layer ML, and then normalizing them. The type corresponding to the node with the largest probability value is the judgment result.
[0031] As shown in FIG. 4, the output layer OL has the number of nodes equal to the number of type labels (here, five) for the output vector (intermediate layer output MLout) output from the intermediate layer ML. This output layer OL calculates probability values P1 to P5 for the type labels L1 to L5 by normalizing the values input to each node (n1 to n5) using a softmax function. Returning to FIG. 1, the description of the configuration of the comment audio insertion timing learning device 1 will be continued.
[0032] The type determination means 12 inputs the acoustic feature extracted by the acoustic feature extraction means 10 in time series to a type determination model N, and calculates a probability value of the type for each frame as a type determination result. Here, the type determination means 12 sequentially shifts the acoustic features for each frame extracted in time series, inputs up to 300 acoustic features into a type determination model N (see FIG. 3), and calculates the type determination model N to calculate probability values P1 to P5 for type labels L1 to L5 (see FIG. 4). The type determination means 12 outputs the calculated probability values P1 to P5 for the type labels L1 to L5 to the error calculation means 13 as type determination results.
[0033] The error calculation means 13 calculates an error in the type determination based on the probability value calculated by the type determination means 12 and the teacher data. Here, the error calculation means 13 calculates the difference between "1" and the probability value corresponding to the type indicated by the type label L, which is the teacher data, in the type determination result calculated by the type determination means 12. For example, if L3 (word ending) is set as the type corresponding to the frame in the type label L, the error calculation means 13 calculates the difference between "1" and the probability value corresponding to the type label L3, among the probability values for each type calculated by the type determination means 12, as the error. The error calculation means 13 outputs the calculated error to the parameter update means 14 .
[0034] The parameter update means 14 updates the parameters of the type determination model N based on the error calculated by the error calculation means 13. That is, the parameter update means 14 updates the parameters in a direction that reduces the error. For updating the parameters in the parameter updating means 14, a general neural network optimization method such as Stochastic Gradient Descent (SGD) or Adaptive moment estimation (Adam) can be used. The parameter update means 14 updates the parameters of the type determination model N stored in the storage means 11 by the stochastic gradient descent method or the like. The parameter updating means 14 updates the parameters and instructs the type determining means 12 to repeatedly perform type determination using the same acoustic feature until the parameter update termination condition is met.
[0035] Furthermore, when the parameter update termination condition is reached, the parameter update means 14 instructs the type determination means 12 to shift the frame and perform type determination on the new frame if there is an acoustic feature of the next frame. The condition for ending this parameter update is, for example, when the amount of change in error becomes less than a predetermined threshold, or when the number of parameter updates for the same acoustic feature exceeds a predetermined number. As a result, the comment audio insertion timing learning device 1 continues learning the category determination model N until the input of the audio V is completed.
[0036] With the above-described configuration, the comment audio insertion timing learning device 1 can learn the type determination model N for determining the type label for each frame in a predetermined time section of any audio. The comment audio insertion timing learning device 1 can be operated by a comment audio insertion timing learning program that causes a computer to function as each of the above-mentioned means.
[0037] [Operation of the commentary audio insertion timing learning device] Next, the operation of the comment audio insertion timing learning device 1 according to the first embodiment of the present invention will be described with reference to Fig. 5 (see Fig. 1 for the configuration as appropriate). It is assumed that the storage means 11 stores a type determination model N with an initial value set.
[0038] In step S1, the acoustic feature extraction means 10 extracts acoustic features for each frame by sequentially performing acoustic analysis on the speech V at a predetermined frame length (for example, 10 ms) and a predetermined shift width (for example, 10 ms). In step S2, the type determination means 12 sequentially inputs the acoustic features extracted in step S1 into the type determination model N stored in the storage means 11, and calculates the type probability value for each frame as the type determination result.
[0039] In step S3, the error calculation means 13 calculates the error of the type determination result calculated in step S2 based on the type for each frame indicated by the training data type label L. Here, the error calculation means 13 calculates the difference between the probability value corresponding to the type indicated by the training data type label L and "1" in the type determination result calculated in step S2 as the error. In step S4, the parameter updating means 14 determines whether or not a parameter update termination condition has been reached, such as when the amount of change in error becomes less than a predetermined threshold or when the number of parameter updates for the same acoustic feature exceeds a predetermined number.
[0040] If the parameter update termination condition has not yet been reached (No in step S4), in step S5, the parameter update means 14 updates the parameters of the type determination model N stored in the storage means 11 by a stochastic gradient descent method or the like. Then, the process returns to step S2, and the type determination means 12 repeats the parameter update by calculating the type determination result using the same acoustic features. On the other hand, if the parameter update termination condition is reached (Yes in step S4), the type determination means 12 determines whether learning is complete in step S6 based on whether a new acoustic feature exists.
[0041] Here, if it is determined that the learning is not complete (No in step S6), the type determination means 12 shifts the frame in step S7 and repeats the learning by calculating the type determination result using the shifted acoustic features in step S2. On the other hand, if it is determined that the learning is completed (Yes in step S6), the comment audio insertion timing learning device 1 ends its operation. Through the above operations, the comment audio insertion timing learning device 1 can learn the category determination model N.
[0042] [Configuration of the commentary audio insertion timing detection device] Next, with reference to FIG. 6, the configuration of a comment audio insertion timing detection device 2 according to a second embodiment of the present invention will be described.
[0043] The commentary audio insertion timing detection device 2 detects the timing of commentary audio insertion from audio. Here, the commentary audio insertion timing detection device 2 detects the timing of commentary audio insertion by detecting the type of frames in each predetermined time interval in the audio. Note that the types identify, for each frame in the predetermined time interval (e.g., 10 ms), at least an insertion prohibited interval that prohibits the insertion of commentary audio, an overlapping insertion permitted interval that allows overlapping insertion of commentary audio immediately before the end of an utterance interval, and an insertion permitted interval that allows the insertion of commentary audio in a non-utterance interval. Here, the types are the type labels L1 to L5 described in FIG. 2. As shown in FIG. 6, the comment audio insertion timing detection device 2 includes an acoustic feature extraction unit 20, a storage unit 21, a type determination unit 22, and a type determination unit .
[0044] The acoustic feature extraction means 20 extracts acoustic features for each frame from the speech V. The acoustic feature extraction means 20 is the same as the acoustic feature extraction means 10 described in Fig. 1, and the frame length for extracting acoustic features, the shift width, and the type of acoustic feature (pitch frequency) to be extracted are the same as those of the acoustic feature extraction means 10. The acoustic feature extraction means 20 outputs the extracted acoustic feature for each frame to the type determination means 22 .
[0045] The storage means 21 stores a neural network model. This storage means 21 can be configured with a general storage medium such as a semiconductor memory. Here, the storage means 21 stores in advance the type determination model N trained by the commentary audio insertion timing training device 1 (see FIG. 1).
[0046] The type determination means 22 inputs the acoustic features extracted by the acoustic feature extraction means 20 in time series to a type determination model N, and calculates a probability value of the type for each frame as a type determination result. The type determination means 22 is the same as the type determination means 12 described in FIG. 1. That is, the type determination means 22 sequentially shifts the acoustic features for each frame extracted in time series, inputs the latest 300 acoustic features to a type determination model N (see FIG. 3), and calculates the type determination model N to calculate probability values P1 to P5 for type labels L1 to L5 (see FIG. 4). The type determination means 22 outputs the calculated probability values P1 to P5 for the type labels L1 to L5 to the type decision means 23 as type determination results.
[0047] The type determination means 23 determines the type of the maximum probability value calculated by the type determination means 22 as the type of the frame. The type determination means 23 sequentially outputs the type label L determined for each frame to the outside. Of course, the type determination means 23 may record and output the type label L for the audio V together as one data file.
[0048] With the above-described configuration, the comment audio insertion timing detection device 2 can detect the type of each frame in any audio. This type for each frame makes it possible to detect insertion-permitted sections that allow commentary audio to be inserted into non-speech sections, as well as overlapping insertion-permitted sections that allow commentary audio to be inserted overlappingly just before the end of a speech section, making it possible to detect the timing for inserting commentary audio that allows overlapping audio without compromising the meaning of the sentence. Furthermore, since the commentary audio insertion timing detection device 2 can detect the beginning of an insertion section in frame units, it can quickly detect the timing to insert commentary audio even when inserting commentary audio in real time. The commentary audio insertion timing detection device 2 can be operated by a commentary audio insertion timing detection program that causes a computer to function as each of the above-mentioned means.
[0049] [Operation of the commentary audio insertion timing detection device] Next, the operation of the comment audio insertion timing detection device 2 according to the second embodiment of the present invention will be described with reference to Fig. 7 (see Fig. 6 for the configuration as appropriate). It is assumed that the storage means 21 stores the type determination model N learned by the comment audio insertion timing learning device 1 (see Fig. 1).
[0050] In step S10, the acoustic feature extraction means 20 extracts acoustic features for each frame by sequentially performing acoustic analysis on the speech V at a predetermined frame length (for example, 10 ms) and a predetermined shift width (for example, 10 ms). In step S11, the type determination means 22 sequentially inputs the acoustic features extracted in step S10 into the type determination model N stored in the storage means 21, and calculates the type probability value for each frame as the type determination result. In step S12, the type determination means 23 determines the type label with the maximum probability value as the type for the frame from the type determination results calculated in step S11, and outputs it to the outside.
[0051] In step S13, the type determination means 22 determines whether or not the analysis of the voice is finished depending on whether or not a new acoustic feature exists. Here, if it is determined that the analysis of the voice has not been completed (No in step S13), the type determination means 22 shifts the frame in step S14 and repeats the determination operation by calculating the type determination result using the shifted acoustic features in step S11. On the other hand, if it is determined that the analysis of the audio has been completed (Yes in step S13), the comment audio insertion timing detection device 2 ends its operation. Through the above operations, the comment audio insertion timing detection device 2 can detect the timing to insert comment audio while allowing audio overlap without impairing the meaning of the sentence.
[0052] The configurations and operations of the commentary audio insertion timing learning device 1 and the commentary audio insertion timing detection device 2 according to the embodiment of the present invention have been described above, but the present invention is not limited to this embodiment.
[0053] Here, the type label L2 is set to indicate that the end of the speech section is near, but this label is not essential and may be included in the "utterance" of the type label L1. Alternatively, the type label L2 may be included in the "ending" of the type label L3, indicating a section where commentary audio may be inserted by overlapping with the audio. In addition, although an LSTM is used in the intermediate layer of the classification determination model N here, any neural network that handles time-series data can be used, and a general RNN may also be used. [Explanation of symbols]
[0054] 1. Audio commentary insertion timing learning device 10 Acoustic feature extraction means 11 Memory means 12. Method of classifying 13 Error calculation means 14 Parameter update method 2. Commentary audio insertion timing detection device 20 Acoustic feature extraction means 21 Memory means 22 Type determination means 23 Classification determination method
Claims
1. A commentary audio insertion timing learning device that learns a neural network type determination model for determining the type of each section of an audio frame in any audio from training data that sets types that distinguish between an insertion prohibited section that prohibits insertion of commentary audio, an overlapping insertion permitted section that allows overlapping insertion of commentary audio immediately before the end of an utterance section, and an insertion permitted section that allows insertion of commentary audio in a non-utterance section, an acoustic feature extraction means for extracting acoustic features from the speech for each of the speech frames; a type determination means for inputting the acoustic feature extracted by the acoustic feature extraction means in time series to the type determination model and calculating a probability value of the type of the speech frame as a type determination result; an error calculation means for calculating an error in the type determination based on the probability value calculated by the type determination means and the teacher data; a parameter update means for updating parameters of the classification determination model based on the error calculated by the error calculation means, thereby learning the classification determination model; A commentary audio insertion timing learning device comprising:
2. 2. The comment audio insertion timing learning device according to claim 1, wherein the overlap insertion permitted section is a time section that is a predetermined time section back from the end position of the speech section.
3. The commentary audio insertion timing training device according to claim 1 or 2, characterized in that the type determination means performs calculations using the type determination model with the LSTM and fully connected layer of a recurrent neural network on the acoustic features input in time series, and normalizes output values of nodes equal to the number of types using a softmax function before outputting the normalized values.
4. 4. The commentary audio insertion timing learning device according to claim 1, wherein the acoustic feature extracting means extracts a pitch frequency as the acoustic feature.
5. A commentary audio insertion timing learning program for causing a computer to function as the commentary audio insertion timing learning device according to any one of claims 1 to 4.
6. A commentary audio insertion timing detection device that detects a commentary audio insertion timing from audio, comprising: an acoustic feature extraction means for extracting acoustic features from the speech for each speech frame; a type determination means for inputting the acoustic features in a time series into a type determination model of a pre-trained neural network that determines, for each section of the audio frame, an insertion prohibited section that prohibits insertion of commentary audio, an overlapping insertion permitted section that allows overlapping insertion of commentary audio immediately before the end of an utterance section, and an insertion permitted section that allows insertion of commentary audio in a non-utterance section, and calculating a probability value of the type of the audio frame as a type determination result; a type determination means for determining the type of the voice frame having the maximum probability value calculated by the type determination means; A commentary audio insertion timing detection device comprising:
7. A commentary audio insertion timing detection program for causing a computer to function as the commentary audio insertion timing detection device according to claim 6.
Citation Information
Patent Citations
Navigation apparatus
JP2006300648A
Detection device, detection method and detection program
JP2019028405A
Device and program for predicting speech end timing
JP2020064248A
Automatic Calculation of Gains for Mixing Narration Into Pre-Recorded Content
US20170092290A1