Speech data provision device, speech estimation system, speech data provision system, speech data provision method and program
The system enhances speech data separation by using AI to estimate and reconstruct missing speech elements in mixed audio data, addressing the challenge of simultaneous speech identification and improving accuracy through contextual analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MITSUBISHI ELECTRIC CORP
- Filing Date
- 2025-07-15
- Publication Date
- 2026-04-17
AI Technical Summary
Existing sound source separation technologies struggle to accurately identify individual utterances in audio data where simultaneous speech occurs, leading to difficulties in separating and reconstructing speech elements with high accuracy.
A system utilizing artificial intelligence to estimate and reconstruct missing speech elements by employing a sound source separation model and a large-scale language model, which processes mixed speech data to identify and fill in missing parts based on contextual information, and evaluates the accuracy of the estimation using an evaluation model.
Improves the accuracy of identifying individual utterances in audio data by reconstructing missing speech elements, ensuring contextual consistency and speaker identification, even in situations with simultaneous speech.
Smart Images

Figure 0007847733000001 
Figure 0007847733000002 
Figure 0007847733000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a speech data providing device, a speech estimation system, a speech data providing system, a speech data providing method, and a program.
Background Art
[0002] In recent years, it has been proposed to apply greatly advanced machine learning techniques to sound source separation (for example, see Patent Document 1). Patent Document 1 describes a technique for generating a voice separation model of a neural network that separates utterances included in a given voice and superimposed on a plurality of different voices. Specifically, this technique avoids the execution of conventional stitching processing by positioning the association between the separated utterance and the speaker as a graph coloring problem. According to the voice separation model of Patent Document 1, a conversation in which utterances overlap can be quickly separated into the utterances of each speaker.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a conversation as assumed in Patent Document 1, it is expected to some extent that the speakers' utterances are made alternately. On the other hand, for example, there is also a desire to obtain the utterance contents of each speaker or conversation group from data recorded in a situation where the timing of simultaneous utterance can occur frequently, such as when recording the voices of different conversation groups simultaneously in a meeting.
[0005] However, separating sound sources at the same time they are spoken is not easy, and this was not particularly considered in the technology described in Patent Document 1. Therefore, there is room for improvement in the accuracy of identifying individual utterances from audio data that includes parts where sound source separation is difficult.
[0006] This disclosure aims to improve the accuracy of identifying individual utterances from audio data that includes portions where sound source separation is difficult. [Means for solving the problem]
[0007] To achieve the above objective, the speech data providing device of this disclosure comprises: acquisition means for acquiring mixed speech data recorded by mixing multiple voices; control means for inputting the mixed speech data into artificial intelligence to estimate each voice or the speech which is the text spoken as each voice; and output means for outputting speech data indicating each estimated speech, wherein at least one estimated speech is a speech element generated based on a first part of the mixed speech data and includes a speech element estimated for a second part different from the first part. Furthermore, the control means obtains an estimation result by having the artificial intelligence estimate the second part without relying on the first part, and has the artificial intelligence or another artificial intelligence evaluation device evaluate the estimation of the utterance elements by comparing the estimated utterance with the estimation result. . [Effects of the Invention]
[0008] According to this disclosure, it is possible to improve the accuracy of identifying individual utterances from audio data that includes parts where sound source separation is difficult. [Brief explanation of the drawing]
[0009] [Figure 1] Diagram showing the configuration of the speech estimation system according to the embodiment. [Figure 2] Diagram showing the hardware configuration of a computer according to the embodiment. [Figure 3] Diagram showing the functional configuration of the provided device according to the embodiment. [Figure 4] A diagram illustrating the separation of audio and the completion of missing portions according to the embodiment. [Figure 5] A diagram showing an example of supplementing missing parts in the embodiment. [Figure 6]Flowchart showing the speech data provision process according to the embodiment [Figure 7] A diagram showing an example of metadata related to the embodiment. [Figure 8] Flowchart showing the evaluation process according to the embodiment [Figure 9] A diagram showing an example of evaluation results related to the embodiment. [Modes for carrying out the invention]
[0010] The speech estimation system according to the embodiment of this disclosure will be described in detail below with reference to the drawings.
[0011] Embodiment. The speech estimation system 1000 according to this embodiment, as shown in Figure 1, is a system that provides speech data 50 obtained by estimating the speeches of each speaker U1 to U4 from mixed speech data 12, which is recorded by mixing the voices of speakers U1 and U2 belonging to conversation group G1 and speakers U3 and U4 belonging to conversation group G2. Because the number of speakers is relatively large, some of the speeches to be restored are lost in conventional sound source separation of the mixed speech data 12, but the speech estimation system 1000 estimates the speeches by using artificial intelligence to fill in the missing parts. The speech estimation system 1000 includes a recording device 10 for recording voice, a providing device 20 for receiving the mixed speech data 12 and providing speech data 50, a generating device 30 for estimating speeches using artificial intelligence, an evaluation device 40 for evaluating the speech estimation using artificial intelligence, and a receiving device 60 for receiving speech data 50.
[0012] The recording device 10, the providing device 20, the generating device 30, the evaluation device 40, and the receiving device 60 communicate with each other by being connected via a network (not shown). This network is, for example, a communication network such as a LAN (Local Area Network) or the Internet. Alternatively, instead of a network, some or all of the above-mentioned devices may communicate via a communication line such as a USB (Universal Serial Bus) cable.
[0013] The recording device 10 may be a PC (Personal Computer), a tablet terminal, or a smartphone, or it may be another device that has the function of creating mixed audio data 12. The recording device 10 is connected by a signal line to a microphone 11 placed on a table in a conference room, and the mixed audio data 12 is created by recording the sound waveform generated in the conference room by this microphone 11. The microphone 11 may be built into the recording device 10. The microphone 11 may be single-channel or have two or more channels. If there are two or more channels, the accuracy of sound source separation described later can be improved by utilizing spatial information such as the phase difference corresponding to the direction of arrival of the sound waves.
[0014] The providing device 20 is, for example, a server device or a PC on a network. The providing device 20, being a server device, may provide a function to estimate a statement as a cloud service upon request from the receiving device 60. The providing device 20 transmits statement data 50, which represents the estimated statement, to the receiving device 60. The statement represented by the statement data 50 may be speech represented in the form of an acoustic waveform, or it may be in the form of text transcribed from the speech. In either form, the estimated statement may include elements estimated by artificial intelligence. The providing device 20 also provides the receiving device 60 with the results of an evaluation of the statement estimation. The providing device 20 is an example of a statement data providing device.
[0015] The generation device 30 and the evaluation device 40 are each, for example, server devices on a network. The generation device 30 and the evaluation device 40 have functions of artificial intelligence as described later, and these functions are used by the providing device 20. Here, artificial intelligence refers to intelligent functions such as inference and judgment realized by a neural network, and a device that provides an operating environment for the functions. The generation device 30 and the evaluation device 40 as artificial intelligence have a multimodal function that can process any of text, which is a natural language, voice, and a combination of text and voice. Further, the generation device 30 and the evaluation device 40 may be able to process data in other forms represented by video that can be converted into text or voice. Although the following description will mainly focus on the example where the generation device 30 and the evaluation device 40 are separate devices, the generation device 30 and the evaluation device 40 may be implemented by a single device, that is, the same artificial intelligence.
[0016] The receiving device 60 is a terminal device equipped with a UI (User Interface) such as a PC, a tablet terminal, or a smartphone. The receiving device 60 receives the speech data 50 from the providing device 20 and presents the speech of each of the speakers U1 to U4 indicated by the speech data 50 to the user of the receiving device 60. The presentation of the speech to the user may be by voice playback or text display. Further, the receiving device 60 receives the evaluation result from the evaluation device 40 from the providing device 20 and presents the evaluation result to the user.
[0017] The recording device 10, the providing device 20, the generation device 30, the evaluation device 40, and the receiving device 60 are each configured as a computer 70 having a hardware configuration as shown in FIG. 2. Specifically, the computer 70 includes a processor 701, a main storage unit 702, an auxiliary storage unit 703, an input unit 704, an output unit 705, and a communication unit 706. The main storage unit 702, the auxiliary storage unit 703, the input unit 704, the output unit 705, and the communication unit 706 are all connected to the processor 701 via an internal bus 707.
[0018] The processor 701 includes a CPU (Central Processing Unit) or MPU (Micro Processing Unit) as a processing circuit. By executing the program P1 stored in the auxiliary storage unit 703, the processor 701 realizes various functions and executes the processes described later. The processor 701 that realizes the operations by the neural network may further include a GPU (Graphical Processing Unit).
[0019] The main storage unit 702 includes a RAM (Random Access Memory). The program P1 is loaded into the main storage unit 702 from the auxiliary storage unit 703. Then, the main storage unit 702 is used as a work area for the processor 701.
[0020] The auxiliary storage unit 703 includes non-volatile memories typified by a semiconductor flash memory and an HDD (Hard Disk Drive). In addition to the program P1, the auxiliary storage unit 703 stores various data used for the processing by the processor 701. The auxiliary storage unit 703 supplies the data used by the processor 701 to the processor 701 according to the instruction of the processor 701. Also, the auxiliary storage unit 703 stores the data supplied from the processor 701.
[0021] The input unit 704 includes input components typified by a hardware switch, an input key, a keyboard, and a pointing device. The input unit 704 acquires the information input by the user of the computer 70 and notifies the acquired information to the processor 701.
[0022] The output unit 705 includes output components typified by an LED (Light Emitting Diode), an LCD (Liquid Crystal Display), and a speaker. The output unit 705 presents various information to the user according to the instruction of the processor 701. The LCD of the output unit 705 and the pointing device of the input unit 704 may be integrally configured as a touch screen.
[0023] The communication unit 706 includes a communication interface circuit for communicating with an external device. The communication unit 706 receives a signal from an external source and outputs the data indicated by this signal to the processor 701. The communication unit 706 also transmits a signal indicating the data output from the processor 701 to the external device.
[0024] The aforementioned hardware configurations work together to enable the providing device 20, the generating device 30, and the evaluation device 40 to perform various functions. In detail, as shown in Figure 3, the providing device 20 includes an acquisition unit 21 that acquires mixed speech data 12, a receiving unit 22 that receives information from the user of the providing device 20, a processing unit 23 that performs various processes, a control unit 24 that controls the generating device 30 and the evaluation device 40 to perform the functions described later, and an output unit 25 that outputs spoken data 50.
[0025] The acquisition unit 21 is primarily implemented by the communication unit 706 of the providing device 20. The acquisition unit 21 acquires mixed voice data 12 from the recording device 10. The source of the mixed voice data 12 to the acquisition unit 21 is not limited to the recording device 10. For example, the acquisition unit 21 may read the mixed voice data 12 from a storage device other than the recording device 10. Alternatively, the acquisition unit 21 may be implemented by the processor 701 of the providing device 20 and read the mixed voice data 12 from a recording medium such as a memory card that is removable from the providing device 20. Furthermore, the acquisition unit 21 may directly acquire the mixed voice data 12 in real time from the microphone 11 instead of the recording device 10. The acquisition unit 21 outputs the acquired mixed voice data 12 to the processing unit 23. The acquisition unit 21 is an example of an acquisition means for acquiring mixed voice data, which is recorded by mixing multiple voices.
[0026] Since the mixed speech data 12 is data that records the speech of different conversation groups G1 and G2, it is thought to contain more portions where each voice is spoken simultaneously than when the speech of a single conversation group is recorded. In such portions, it is difficult to obtain sufficient separation accuracy even when conventional sound source separation is applied, and some of the separated speech may be lost. In addition, due to factors such as microphone 11 malfunction, network failure, and environmental noise, information loss may occur in the mixed speech data 12 itself, resulting in loss of information in the separated speech. Here, loss means the loss of information that was present in the speech before mixing.
[0027] The reception unit 22 is primarily implemented by the input unit 704 of the providing device 20. The reception unit 22 receives filtering conditions from the user of the providing device 20 that the utterance data 50 output from the output unit 25 must satisfy, and notifies the processing unit 23 of the received filtering conditions. The reception unit 22 may also be implemented by the communication unit 706 of the providing device 20 and receive filtering conditions transmitted from an external UI device via a communication line or network. Furthermore, the reception of filtering conditions by the reception unit 22 may be omitted. That is, all utterance data 50 may be output from the output unit 25 without any special conditions being specified by the user. The reception unit 22 is an example of a reception means for receiving filtering conditions.
[0028] The processing unit 23 is primarily implemented by the processor 701 of the supply device 20. The processing unit 23 performs preprocessing on the mixed speech data 12. This preprocessing includes, for example, noise removal and volume normalization. The processing unit 23 outputs the preprocessed mixed speech data 12 to the control unit 24. The processing unit 23 also notifies the control unit 24 of the filtering conditions received by the receiving unit 22. Furthermore, the processing unit 23 acquires the speech data 50 generated based on the mixed speech data 12 and the filtering conditions from the control unit 24 and outputs the acquired speech data 50 to the output unit 25.
[0029] The control unit 24 is primarily realized through the cooperation of the processor 701 and communication unit 706 of the supply device 20. The control unit 24 inputs the mixed speech data 12 to the generation device 30 and causes the generation device 30 to estimate utterances that match the filtering conditions. Details of the process by which the generation device 30 estimates utterances under the control of the control unit 24 will be described later. The control unit 24 is an example of a control means that causes artificial intelligence to estimate utterances, which are each voice or the text spoken as each voice, by inputting the mixed speech data. The control unit 24 also inputs the estimated utterances and the missing data obtained by having the generation device 30 estimate the utterances without filling in the missing parts to the evaluation device 40 and causes the evaluation device 40 to evaluate the estimation of utterance elements corresponding to the missing parts. Details of the process by which the evaluation device 40 evaluates utterances under the control of the control unit 24 will be described later.
[0030] The output unit 25 is primarily realized through the cooperation of the processor 701 and the communication unit 706 of the providing device 20. The output unit 25 may be a web server that provides the utterance data 50 as part of a web page in response to a request from the receiving device 60. The output unit 25 corresponds to an example of an output means that outputs utterance data representing each estimated utterance. The output unit 25 also acquires the evaluation results from the evaluation device 40 via the control unit 24 and the processing unit 23, and outputs information indicating the acquired evaluation results to the receiving device 60.
[0031] The generation device 30 stores a pre-trained sound source separation model 31 and a large-scale language model 32. The sound source separation model 31 is a model that separates the waveforms of mixed sound sources, and it is particularly desirable that it is a model that separates speech. Furthermore, since the sounds produced differ depending on the language, it is desirable that the model be specialized for languages that are expected to be mixed in the mixed speech data 12.
[0032] The large-scale language model 32 is, for example, a model such as GPT (Generative Pretrained Transformer) or BERT (Bideirectional Encoder Representations from Transformer) that can process data in a multimodal manner. However, the large-scale language model 32 is not limited to the above examples, and as will be described later, any model capable of performing natural language processing using the context of what is being said in the speech is acceptable. Furthermore, the large-scale language model 32 may substantially incorporate the sound source separation model 31.
[0033] The generation device 30 separates the voices of each speaker U1 to U4 from the mixed speech data 12 using the sound source separation model 31, in accordance with the instructions of the control unit 24 of the providing device 20. The generation device 30 may separate the voices for each conversation group G1, G2 instead of separating the voices for each speaker U1 to U4. The generation device 30 then identifies the parts where information is missing in the separated voices and fills in the missing parts to make the utterance sound natural based on the surrounding context. For example, the degree to which the consistency of the surrounding context is maintained is determined for each unit period, and if the determined degree is below a threshold, it is determined that the unit period corresponds to the missing part. A unit period is, for example, a frame of a predetermined time length, a period corresponding to one vowel or consonant recognized by speech, or a period corresponding to one character in the transcribed text. The generation device 30 estimates the utterance by generating utterance elements corresponding to the missing parts of the separated voices using the large-scale language model 32, and generates utterance data 50 indicating the estimated utterance and transmits it to the providing device 20.
[0034] Figure 4 schematically illustrates the sound source separation and missing portion completion performed by the generation device 30. At the top of Figure 4, the unmixed speech waveforms of two of the speakers U1 to U4 are shown, and the waveform obtained by mixing these two speech waveforms corresponds to the mixed speech data 12. When this mixed speech data 12 is separated by the sound source separation model 31, the separated speech contains missing portions 80 with low separation accuracy, as indicated by the dashed box. The generation device 30 estimates the speech elements corresponding to such missing portions 80 from the information before and after the missing portion 80.
[0035] Figure 5 shows an example of estimation based on Japanese characters, illustrating the details of estimating the speech element corresponding to the missing portion 80. As shown in Figure 5, the words corresponding to the audio and text "ko,ng,--,chi,wa" have a low probability of occurrence in the large-scale language model 32, or are not registered in the dictionary used by the large-scale language model 32. The same is true even if the word boundaries are changed. For this reason, "--" is identified as the missing portion. Here, if we treat phrase 81 as a single word and assign the character "ni" to the missing portion, it is interpreted that the greeting words "ko,ng,ni,chi,wa" were spoken, and furthermore, no inconsistency arises in the context with other phrases before and after this phrase 81. For this reason, the generation device generates the character "ni" as the speech element for the first missing portion 80 in Figure 5, and if the speech data 50 is in audio format, it synthesizes the audio corresponding to "ni" and inserts it between the separated audio before and after. Figure 5 shows that the speech elements "me" and "shi" are generated for the two missing portions in the same way.
[0036] In the example in Figure 5, utterance elements were estimated to produce natural-sounding speech based on context, but the method for estimating utterance elements is not limited to this. For example, the generator 30 may estimate utterance elements based on speech features such as intonation and vocalization patterns. Also, while Figures 4 and 5 illustrate an example where missing parts are identified and completed according to a predetermined procedure, the completed utterance may be output from a neural network that performs processes that do not necessarily correspond one-to-one with the said procedure.
[0037] Furthermore, the generating device 30, in accordance with the instructions of the control unit 24 of the providing device 20, transmits the missing data representing the separated audio to the providing device 20 for evaluation by the evaluation device 40, without supplementing the missing portions of the audio. The missing data may be audio separated from the mixed audio data 12 by the sound source separation model 31, or text transcribed from said audio, which has not been processed by the large-scale language model 32.
[0038] Furthermore, if filtering conditions are specified by the control unit 24, the generation device 30 transmits utterance data 50 that satisfy those conditions. For utterance data 50 that do not satisfy the conditions, it does not need to transmit them, or it may transmit them with a flag indicating that the conditions are not met. Filtering conditions include, for example, specifying audio data of a single phrase from a specific speaker, and then specifying that the utterance is from that speaker, or from a conversation group that includes that speaker. Also, if the text data "related to semiconductors" is specified as a filtering condition, the generation device 30 outputs utterance data 50 from a single speaker containing a phrase related to the topic of semiconductors, or utterance data 50 from a conversation group that includes that single speaker, using the large-scale language model 32. Furthermore, the generation device 30 may extract conversational portions that follow the specified topic and remove conversational portions that follow other topics to generate utterance data 50.
[0039] The evaluation device 40 stores a pre-trained evaluation model 41. The evaluation model 41 is a model that compares two statements and outputs an index value on a predetermined scale. The evaluation model 41 may be a model specifically trained for comparing statements and outputting index values, or it may be a general-purpose large-scale language model. The scale of the index value may be, for example, the contextual likelihood, which indicates the likelihood of something being natural in context, and the index value may be a numerical score from 1 to 100.
[0040] The evaluation device 40 receives the speech data 50 and the missing data from the control unit 24 and identifies the discrepancies between these data as the completed portion. Alternatively, the generation device 30 may assign a flag or metadata indicating the completed portion to at least one of the speech data 50 and the missing data, and the evaluation device 40 may identify the completed missing portion by referring to this flag or metadata.
[0041] The evaluation device 40 then outputs the time or period in the identified supplementary portion of the speech data 50, the index value obtained by evaluating the supplementary portion, and a low evaluation flag that is assigned when the index value falls below a predetermined threshold. This threshold may be specified by the user of the evaluation device 40 or the user of the providing device 20. When the evaluation device 40 outputs a low evaluation flag, it may generate and output new speech elements to be supplemented for the supplementary portion to which the low evaluation flag has been assigned, in the same manner as the generation device 30. The generation of these new speech elements may be performed by the generation device 30 controlled by the evaluation device 40. The format of the new speech elements may be, for example, speech or text, and may be the same format as the speech data 50, but may be in other formats.
[0042] Next, the processing flow performed by the providing device 20 will be explained using Figures 6 to 9.
[0043] Figure 6 shows the flow of the speech data provision process, which estimates each speech from the mixed speech data 12 and outputs speech data 50. In the speech data provision process, the acquisition unit 21 acquires the mixed speech data 12 (step S1), and the reception unit 22 accepts the filtering conditions (step S2).
[0044] Next, the processing unit 23 performs preprocessing (step S3). By performing preprocessing, the quality of the mixed speech data 12 can be improved, thereby improving the accuracy of subsequent processing.
[0045] Next, the control unit 24 inputs the mixed speech data 12, which has been preprocessed in step S3, and the filtering conditions received in step S2 to the generation device 30, causing it to estimate each utterance with the missing parts completed (step S4). Specifically, the control unit 24 gives the generation device 30 instructions or prompts to complete the missing parts so that the utterance becomes natural in context.
[0046] The generation device 30, in accordance with instructions from the control unit 24, identifies the speaker who uttered each sound based on the meaning of the text in the spoken portion of the sound, as well as speech features including at least one of the fundamental frequency, intonation, and speed of the spoken portion. Furthermore, the generation device 30 identifies conversation groups based on the meaning of the text. Here, the meaning of the text refers to a meaning that makes sense as a natural utterance and does not cause inconsistencies in context. Then, as shown in Figure 7, the generation device 30 adds the identification result as metadata to the speech data 50.
[0047] In the example in Figure 7, speech data 51 and 52, which correspond to an example of speech data 50, are audio from conversation groups G1 and G2, respectively, and one speech data 50 represents the speeches of two speakers. The metadata, as shown in Figure 7, includes a conversation group ID (Identifier) to identify the conversation group and a speaker ID to identify the speaker, as well as a summary of each audio and information indicating the estimated time of the speech element for the missing portion. The metadata may follow the format provided for the audio data format or it may follow a different format from the audio data.
[0048] Returning to Figure 6, following step S4, the control unit 24 acquires each statement that satisfies the filtering conditions, which has been generated by the generation device 30 (step S5). Then, the output unit 25 outputs statement data 50 representing each statement, along with metadata, to the receiving device 60 (step S6). After that, the statement data provision process ends. The statement data provision process is an example of a statement data provision method.
[0049] Figure 8 shows the flow of the evaluation process for evaluating the completion of speech. In the evaluation process, the control unit 24 acquires speech data 50 and missing data from the generation device 30 (step S11). Specifically, the control unit 24 acquires from the generation device 30 the speech data 50 generated in the speech data provision process and missing data indicating speech corresponding to speech separated from the mixed speech data 12 without completing the missing parts.
[0050] Next, the control unit 24 inputs the acquired utterance data 50 and the missing data to the evaluation device 40 and gives instructions to compare the two utterances and evaluate the contextual likelihood of completing the elements contained in the utterances (step S12). Here, a threshold may be set as a parameter for assigning a low evaluation flag, if necessary.
[0051] Next, the control unit 24 acquires the evaluation result for each completed utterance element (step S13), and the output unit 25 outputs information indicating the evaluation result (step S14). The information indicating the evaluation result includes the time of the completed utterance element, the index value as the evaluation result, and a low evaluation flag. If the low evaluation flag is active, other completion candidates for the utterance element are notified from the evaluation device 40 along with the low evaluation flag and output by the output unit 25. Figure 9 shows an example of an evaluation result output from the output unit 25 and displayed by the receiving device 60. In the example in Figure 9, for the utterance element completed at time 00:00:20, the index value is 83 and no low evaluation flag is assigned, but for the utterance element completed at time 00:50:30, the index value is 22 and a low evaluation flag is assigned, and new utterance elements, "I will consider it", "I will research it", and "I will verify it", which are other completion candidates, are indicated. After that, the evaluation process ends.
[0052] As described above, the providing device 20 according to this embodiment provides utterance data 50 that indicates an utterance including complementary elements. This utterance includes utterance elements estimated for parts that are difficult to separate from the mixed speech data 12, and utterance elements generated based on parts that have been separated to some extent. This makes it possible to improve the accuracy of individual utterances identified from speech data that includes parts that are difficult to separate from the sound source. Parts that have been separated to some extent are parts that are estimated to have been recorded when one of multiple speeches, from which an utterance is estimated, is pronounced and the other speeches are not pronounced, and correspond to an example of the first part. Parts that are difficult to separate are parts that are estimated to have been recorded when two or more speeches are pronounced simultaneously, and correspond to an example of the second part, which is different from the first part.
[0053] In detail, the parts that are difficult to separate are the parts where contextual inconsistencies occur when a statement is estimated without considering the context. The control unit 24 compensates for missing utterance elements in the mixed speech data by having the generation device 30 estimate the utterance elements while taking into account the context of the preceding and succeeding utterances. The parts that are difficult to separate correspond to an example of the second part as described above, and correspond to an example of a part where contextual inconsistencies occur between the element corresponding to the second part of a statement and at least one other element before and after that element when the element corresponding to the second part of the statement is estimated without being based on the first part. The control unit 24 corresponds to an example of a control means that compensates for missing utterance elements in the mixed speech data by having artificial intelligence estimate the utterance elements while taking into account the context of other elements. Note that not all estimated utterances necessarily contain elements that have been compensated based on context; it is sufficient that at least one estimated utterance contains such elements.
[0054] Furthermore, the control unit 24 causes the generation device 30 to identify the speaker who produced each sound based on the meaning of the text in the spoken portion of the sound, as well as the speech features including at least one of the fundamental frequency, intonation, and speed of the spoken portion. This allows the speaker to be identified and the utterance to be accurately estimated without having to prepare information about the speaker in advance.
[0055] Furthermore, the output unit 25 outputs each statement data 50 with metadata, and the metadata includes speaker identification information corresponding to the statement indicated by the statement data 50. This makes it easier to use the statement data 50. For example, meeting minutes can be easily created from the statements indicated by the statement data 50.
[0056] Furthermore, each voice may belong to one of two different conversation groups, and the control unit 24 may instruct the generation device 30 to estimate a statement for each conversation group. The output unit 25 then outputs statement data 50 with metadata attached to each conversation group, and the metadata includes identification information for the conversation group corresponding to each statement data 50. This makes it easier to use the statement data 50.
[0057] Furthermore, the output unit 25 outputs utterance data 50 that satisfies the filtering conditions, and the filtering conditions include at least one of the following conditions: having characteristics of a specific speaker and following a specific topic. This makes it easy to obtain the utterances to be extracted from the utterances included in the mixed speech data 12.
[0058] Furthermore, the control unit 24 has the evaluation device 40 evaluate the utterance data 50 and the missing data, thereby evaluating the estimation of the utterance elements. This allows the user to determine whether the completed utterance elements are appropriate. The control unit 24 is an example of a control means that obtains an estimation result by having the artificial intelligence estimate the second part without relying on the first part, and then has the artificial intelligence or another artificial intelligence, which is an evaluation device, evaluate the estimation of the utterance elements by comparing the estimated utterance with the estimation result.
[0059] Furthermore, the output unit 25 adds a low-evaluation flag to the utterance data 50 when the index value indicating the evaluation result by the evaluation device 40 falls below a threshold, and outputs the utterance data 50. This makes it easy to identify utterance elements with low interpolation accuracy.
[0060] Furthermore, the control unit 24 estimates new utterance elements for the portion of the mixed speech data 12 corresponding to the utterance elements indicated by the low-evaluation flag, and the output unit 25 outputs information indicating the estimated new utterance elements. This makes it easy to infer the correct utterance for utterance elements with low completion accuracy. Note that the low-evaluation flag may be any other form of marking information, as long as it is possible to mark utterance elements with low index values.
[0061] While embodiments of this disclosure have been described above, this disclosure is not limited to the embodiments described above.
[0062] For example, the number of conversation groups may be one or three or more. The number of speakers belonging to each conversation group may be one or three or more. For example, the audio of one speaker belonging to one conversation group speaking on the telephone may be recorded on a recording device 10 different from the telephone, and the speech estimation system 1000 may fill in the missing parts caused by environmental noise.
[0063] Furthermore, the processing unit 23 of the providing device 20 may perform conventional speaker separation or speech separation. If the processing unit 23 is responsible for these separation processes, the generating device 30 does not need to store the sound source separation model 31. The control unit 24 may then cause the generating device 30 to identify and fill in any missing parts of the speech separated by the processing unit 23.
[0064] The explanation has focused on examples where filtering conditions are applied to speeches estimated from the mixed speech data 12, but the explanation is not limited to this, and filtering conditions may also be applied to metadata. That is, the generation device 30 may not perform filtering according to the filtering conditions, and the output unit 25 may perform filtering on the metadata attached to the speech data 50 output from the generation device 30, outputting speech data 50 that matches the filtering conditions and excluding speech data 50 that does not match the filtering conditions from the output target.
[0065] While an example where the indicator value is a numerical score has been described, the indicator value may also be a value corresponding to a classification category. Furthermore, the evaluation result may be indicated by the presence or absence of a low-evaluation flag, without assigning an indicator value.
[0066] The recording device 10 and the providing device 20 may be configured as a single, integrated device. The system including the recording device 10 and the providing device 20 is an example of a speech data provision system.
[0067] The functions of the providing device 20 according to the above-described embodiment can be realized by dedicated hardware or by a normal computer system.
[0068] For example, a device that performs the above-mentioned processing can be configured by distributing program P1 on a computer-readable recording medium such as a flexible disk, CD-ROM (Compact Disk Read-Only Memory), DVD (Digital Versatile Disk), or MO (Magneto-Optical disk), and then installing program P1 on a computer.
[0069] Alternatively, program P1 may be stored on a disk drive of a server device on a communication network such as the Internet, and then downloaded to a computer, for example, by superimposing it onto a carrier wave.
[0070] Furthermore, the above-mentioned process can also be achieved by launching and executing program P1 while transferring it over a network such as the Internet.
[0071] Furthermore, the above-described process can also be achieved by having all or part of program P1 run on a server device, and by having a computer send and receive information related to that process via a communication network while executing program P1.
[0072] Furthermore, if the above-mentioned functions are implemented by the OS (Operating System) or through collaboration between the OS and the application, only the parts other than the OS may be stored and distributed on a medium, or they may be downloaded to a computer.
[0073] Furthermore, the means for realizing the functions of the providing device 20 are not limited to software; some or all of them may be realized by dedicated hardware or circuits.
[0074] This disclosure allows for various embodiments and modifications without departing from the broad spirit and scope of this disclosure. Furthermore, the embodiments described above are for illustrative purposes only and do not limit the scope of this disclosure. In other words, the scope of this disclosure is indicated by the claims, not by the embodiments. Various modifications made within the scope of the claims and the equivalent significance of the disclosure are considered to be within the scope of this disclosure. [Industrial applicability]
[0075] This disclosure is suitable for identifying speech from mixed and recorded audio. [Explanation of Symbols]
[0076] 10 Recording device, 11 Microphone, 12 Mixed speech data, 20 Providing device, 21 Acquisition unit, 22 Receiving unit, 23 Processing unit, 24 Control unit, 25 Output unit, 30 Generating device, 31 Sound source separation model, 32 Large-scale language model, 40 Evaluation device, 41 Evaluation model, 50-52 Speech data, 60 Receiving device, 70 Computer, 80 Missing parts, 81 Phrase, 701 Processor, 702 Main memory unit, 703 Auxiliary memory unit, 704 Input unit, 705 Output unit, 706 Communication unit, 707 Internal bus, 1000 Speech estimation system, G1, G2 Conversation groups, P1 Program, U1-U4 Speakers.
Claims
1. An acquisition means for acquiring mixed audio data, which is recorded by mixing multiple audio sources, A control means that inputs the mixed speech data into artificial intelligence to estimate each of the aforementioned speeches or the text spoken as each of the aforementioned speeches, An output means that outputs statement data representing each of the estimated statements, Equipped with, At least one estimated utterance is an utterance element generated based on a first portion of the mixed speech data, and includes an estimated utterance element based on a second portion different from the first portion, The control means obtains an estimation result by having the artificial intelligence estimate the second part without relying on the first part, and has the artificial intelligence or another artificial intelligence evaluation device evaluate the estimation of the utterance elements by comparing the estimated utterance with the estimation result. A device that provides speech data.
2. The second part is presumed to be the part recorded when two or more of the aforementioned sounds were pronounced simultaneously. The first part is the portion that is presumed to have been recorded when one of the plurality of sounds, which is presumed to be the utterance, is spoken, and the other sounds are not spoken. The speech data providing device according to claim 1.
3. The second part is the part of the statement that, when the element corresponding to the second part is inferred without regard to the first part, an inconsistency arises between the context of that element and at least one of the other elements that precede or follow that element, The control means compensates for missing utterance elements in the mixed speech data by having the artificial intelligence estimate the utterance elements, taking into account the context with the other elements. The speech data providing device according to claim 1.
4. The control means causes the artificial intelligence to identify each speaker who produced the speech based on the meaning of the text in the spoken portion of the speech, and on the speech features including at least one of the fundamental frequency, intonation, and speed of the spoken portion. The speech data providing device according to claim 1.
5. The output means outputs each of the statement data with metadata attached, The metadata includes speaker identification information corresponding to the statement indicated by each statement data, The speech data providing device according to claim 4.
6. Each of the aforementioned audio recordings belongs to one of several different conversation groups, The control means causes the statement to be estimated for each conversation group, The output means outputs the utterance data with metadata attached to each conversation group. The metadata includes identification information of the conversation group corresponding to each of the utterance data, The speech data providing device according to claim 4.
7. It further includes a means for accepting filtering conditions, The output means outputs the statement data that satisfies the filtering conditions, The filtering conditions include at least one of the following conditions: having characteristics of a particular speaker and following a particular topic. The speech data providing device according to claim 1.
8. The output means adds mark information indicating the utterance element to the utterance data when the index value indicating the evaluation result by the evaluation device falls below a threshold, and outputs the utterance data. The speech data providing device according to claim 1.
9. The control means causes the artificial intelligence to estimate new utterance elements for the second portion corresponding to the utterance elements indicated by the mark information. The output means outputs information indicating the new statement element. The speech data providing device according to claim 8.
10. A speech data providing device according to any one of claims 1 to 9, The artificial intelligence that estimates the aforementioned statement, A speech estimation system equipped with the following features.
11. A recording device that records multiple audio sources to be mixed and transmits mixed audio data representing the recorded multiple audio sources, A speech data providing device according to any one of claims 1 to 9 that receives the mixed speech data, A speech data provision system equipped with the necessary features.
12. The acquisition method acquires mixed audio data, which is recorded by mixing multiple voices. The control means inputs the mixed speech data to the artificial intelligence, causing it to estimate each of the speeches or the text spoken as each of the speeches. The output means outputs statement data representing each of the estimated statements. This includes, At least one estimated utterance is an utterance element generated based on a first portion of the mixed speech data, and includes an estimated utterance element based on a second portion different from the first portion, The control means obtains an estimation result by having the artificial intelligence estimate the second part without relying on the first part, and has the artificial intelligence or another artificial intelligence evaluation device evaluate the estimation of the utterance elements by comparing the estimated utterance with the estimation result. Method of providing speech data.
13. On the computer, We obtain mixed audio data, which is recorded by mixing multiple voices. By inputting the mixed speech data into artificial intelligence, the AI is made to estimate each of the aforementioned speeches or the text spoken as each of the aforementioned speeches. Output statement data representing each of the estimated statements. Let them do it, At least one estimated utterance is an utterance element generated based on a first portion of the mixed speech data, and includes an estimated utterance element based on a second portion different from the first portion, The computer is further instructed to obtain an estimation result by having the artificial intelligence estimate the second part without relying on the first part, and to have the artificial intelligence or another artificial intelligence evaluation device evaluate the estimation of the utterance elements by comparing the estimated utterance with the estimation result. program.
Citation Information
Patent Citations
Speech authentication device
JP2007017840A
Information processing device, and information processing program
JP2022094027A
Information processing apparatus and program
JP2024078838A
Sound restoring device and sound restoring method
WO2006080149A1
Information processing device, information processing method, and program
WO2019139101A1