A device for determining a sound source separation model, a method for determining a sound source separation model, a system for determining a sound source separation model, and a program for determining a sound source separation model.

The system determines a sound source separation model by selecting and mixing clean audio data similar to the user's environment, enhancing speech recognition accuracy by adapting to environmental factors and user attributes without requiring data from all users.

JP2026121181APending Publication Date: 2026-07-23PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2025-01-10
Publication Date
2026-07-23

Smart Images

  • Figure 2026121181000001_ABST
    Figure 2026121181000001_ABST
Patent Text Reader

Abstract

The system determines the sound source separation model best suited to the user's environment. [Solution] The sound source separation model determination device acquires voice data of one user speaking in a predetermined usage environment, mixes the voice data with at least one clean voice data, separates the mixed voice using each of a plurality of sound source separation models, and determines and outputs a sound source separation model to be used for sound source separation of voice data acquired in the predetermined usage environment based on the speech recognition result of the separated data corresponding to the clean voice data among the plurality of separated data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a sound source separation model determination device, a sound source separation model determination method, a sound source separation model determination system, and a sound source separation model determination program.

Background Art

[0002] Patent Document 1 discloses a speech recognition performance prediction system including a learning model that is machine-learned to output a predicted value of speech recognition performance in a space where reverberant speech is obtained when a value based on the reverberant speech is input.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

Means for Solving the Problems

[0005] This disclosure provides a sound source separation model determination device, comprising: an acquisition unit that acquires audio data in which the speech of one user speaking in a predetermined usage environment is captured; a selection unit that selects at least one clean audio data from among the clean audio data of multiple speakers; a mixing unit that generates a mixed audio by mixing the audio data and the clean audio data; a separation unit that separates the mixed audio into separation data corresponding to the audio data and separation data for each of the clean audio data using each of a plurality of sound source separation models; a speech recognition unit that performs speech recognition on each of the plurality of separated separation data; and a model determination unit that determines and outputs a sound source separation model from each of the sound source separation models that is used for sound source separation of the audio data captured in the predetermined usage environment, based on the speech recognition result of the separation data corresponding to the clean audio data among the plurality of separated data.

[0006] Furthermore, this disclosure provides a method for determining a sound source separation model performed by at least one processor, comprising: acquiring audio data in which the speech of one user speaking in a predetermined usage environment is captured; selecting at least one clean audio data from among the clean audio data of multiple speakers; generating a mixed audio by mixing the audio data and the clean audio data; separating the mixed audio into separation data corresponding to the audio data and separation data for each of the clean audio data using each of the multiple sound source separation models; performing speech recognition on each of the separated separation data; and determining and outputting the sound source separation model from each of the multiple separation data that is used for sound source separation of the audio data captured in the predetermined usage environment, based on the speech recognition result of the separation data corresponding to the clean audio data.

[0007] Furthermore, this disclosure provides a sound source separation model determination system comprising a database storing clean voice data for each of a plurality of speakers, and a processing device capable of communicating with the database, wherein the processing device acquires voice data in which the voice of one user speaking in a predetermined usage environment has been recorded, selects and acquires at least one clean voice data from the clean voice data for each of the plurality of speakers stored in the database, generates a mixed voice by mixing the voice data and the clean voice data, separates the mixed voice into separation data corresponding to the voice data and separation data for each of the plurality of sound source separation models, performs speech recognition on each of the plurality of separated separation data, and determines and outputs a sound source separation model from each of the plurality of sound source separation models to be used for sound source separation of the voice data recorded in the predetermined usage environment based on the speech recognition result of the separation data corresponding to the clean voice data among the plurality of separated data.

[0008] Furthermore, this disclosure provides a sound source separation model determination program that is executed by at least one processor, and which includes the steps of: acquiring audio data in which the speech of one user speaking in a predetermined usage environment has been captured; selecting at least one clean audio data from among the clean audio data of multiple speakers; generating a mixed audio by mixing the audio data and the clean audio data; separating the mixed audio into separation data corresponding to the audio data and separation data for each of the multiple sound source separation models; performing speech recognition on each of the separated multiple separation data; and determining and outputting a sound source separation model from among the multiple separation data that is used for sound source separation of the audio data captured in the predetermined usage environment, based on the speech recognition result of the separation data corresponding to the clean audio data. [Effects of the Invention]

[0009] According to this disclosure, it is possible to determine a sound source separation model that is more suitable for the user's environment. [Brief explanation of the drawing]

[0010] [Figure 1] A diagram showing an example of a use case for the sound source separation model determination system according to the embodiment. [Figure 2] Block diagram showing an example of the internal configuration of the sound source separation model determination system according to the embodiment. [Figure 3] Flowchart illustrating an example of the operation procedure of the processing apparatus in the embodiment. [Figure 4] A flowchart illustrating an example of the environment optimization processing procedure for the sound source separation model of the processing device in the embodiment. [Figure 5] A diagram illustrating an example of a method for selecting clean audio data. [Figure 6] Figures 1-3 show examples of mixing ambient sound data and clean sound data. [Modes for carrying out the invention]

[0011] (Background leading to this disclosure) Traditionally, there has been a demand for technology to separate audio data containing the mixed speech of multiple users into individual user audio data. However, speech is susceptible to external disturbances such as ambient noise, other people's speech, or the microphone device being used. In generating a sound source separation model to separate audio data, it has been difficult to create a single large-scale pre-trained model that can handle all usage scenarios (usage environments).

[0012] Therefore, one approach is to generate numerous sound source separation models, each specializing in sound source separation from pre-generated audio data collected in different scenes, and then evaluate these sound source separation models in parallel to select the model best suited to the usage scenario. However, selecting a sound source separation model requires collecting speech from all users simultaneously speaking in the actual usage environment and evaluating the collected audio data, which is time-consuming. Furthermore, when users are asked to select a sound source separation model suitable for their usage scenario, it is difficult for them to determine which model is best suited to their scenario based on the separation results obtained by separating the audio data.

[0013] Hereinafter, embodiments specifically disclosing the sound source separation model determination apparatus, sound source separation model determination method, sound source separation model determination system, and sound source separation model determination program according to this disclosure will be described in detail with reference to the drawings as appropriate. However, unnecessarily detailed explanations may be omitted. For example, detailed explanations of already well-known matters and redundant explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily verbose and to facilitate understanding by those skilled in the art. The accompanying drawings and the following explanation are provided to enable those skilled in the art to fully understand this disclosure and are not intended to limit the subject matter described in the claims.

[0014] First, the use cases of the sound source separation model determination system 100 according to the embodiment will be described with reference to Figures 1 and 2, respectively. Figure 1 is a diagram showing an example of a use case of the sound source separation model determination system 100 according to the embodiment. Figure 2 is a block diagram showing an example of the internal configuration of the processing device P1 in the embodiment.

[0015] The sound source separation model determination system 100 is a system that determines a sound source separation model used for a sound source separation process of separating voice data in which the voices of at least two or more users US1, US2, US3, US4 collected by one sound collection device 13 are mixed, such as in a group call, a conversation, or a web conference, into voice data for each user. The sound source separation model determination system 100 selects any one of a plurality of pre-registered clean voice data based on voice data in which the voice of at least one user US1 is collected in a usage environment where the voice to be separated is actually collected. The sound source separation model determination system 100 determines and outputs a sound source separation model more suitable for the user's usage environment based on the selected clean voice data and the voice data in which the voice of the user US1 is collected.

[0016] Note that the clean voice data referred to here is voice data without (or with removed) disturbance factors (such as the voices of others, noise, or reverberation).

[0017] The sound source separation model determination system 100 includes at least one sound collection device 13, a processing device P1, a database DB, and a network NW. Note that the sound source separation model determination system 100 shown in FIG. 1 is an example and is not limited thereto. For example, the processing device P1 and the database DB of the sound source separation model determination system 100 may be integrally configured, or may be directly connected without passing through the network NW.

[0018] The processing device P1 is connected to be able to communicate wired or wirelessly with the database DB via the network NW and executes data transmission and reception. Note that the wireless communication referred to here is communication via a wireless Local Area Network (LAN) such as Wi-Fi (registered trademark).

[0019] The processing device P1 is realized by, for example, a Personal Computer (hereinafter referred to as "PC"), a notebook PC, a tablet terminal, or a smartphone, etc. Note that the processing device P1 may be realized by an on-premises server or a cloud server, etc. that can realize the functions of the processor 11 described later, and a separate sound collection device 13 and a display I / F 14. The processing device P1 includes a communication unit 10, a processor 11, a memory 12, a sound collection device 13, a display I / F 14, and a camera 15. Note that each of the sound collection device 13 and the display I / F 14 may be configured separately from the processing device P1. Also, the camera 15 is not essential and may be omitted.

[0020] The communication unit 10 is connected to be able to communicate with the database DB either wired or wirelessly and executes data transmission and reception. The communication unit 10 outputs various data transmitted from the database DB to the processor 11. Also, the communication unit 10 transmits various data (control commands) transmitted from the processor 11 to the database DB.

[0021] The processor 11 is configured using, for example, a Central Processing Unit (CPU) or a Field Programmable Gate Array (FPGA) and performs various processes and controls in cooperation with the memory 12. Specifically, the processor 11 refers to the programs and data held in the memory 12 and executes those programs to realize various functions such as a model determination unit 111, a sound source separation unit 112, a speech recognition unit 113, or an output result generation unit 114, etc.

[0022] The model determination unit 111 determines a sound source separation model suitable for the user's sound collection environment based on the voice data in which at least one user's spoken voice is collected and at least one clean voice data registered in the database DB in advance.

[0023] The sound source separation unit 112 uses the trained model determined by the model determination unit 111 to generate separated data by separating the audio data input from the sound pickup device 13 into sound sources (audio data) for each user. The sound source separation unit 112 outputs each of the separated user-specific separated data to the speech recognition unit 113.

[0024] The speech recognition unit 113 performs speech recognition processing on each of the separated data and obtains speech recognition results for each user. The speech recognition unit 113 outputs the speech recognition results to the output result generation unit 114.

[0025] The output result generation unit 114 generates text data or a transcription screen of the speech of multiple users based on the speech recognition results for each user, and outputs (displays) it on the display interface 14.

[0026] Camera 15, implemented by an optical system including, for example, a lens and an image sensor, captures images of at least one user who is the speaker of the speech data used to determine the sound source separation model. Camera 15 outputs the captured images to the processor 11. The timing of the user's image capture by camera 15 may be before or after the user's speech is captured.

[0027] Memory 12 includes, for example, Random Access Memory (RAM) used as work memory when executing each process of the processor 11, and Read Only Memory (ROM) which stores programs and data that define the operation of the processor 11. Data or information generated or acquired by the processor 11 is temporarily stored in RAM. Programs that define the operation of the processor 11 are written to ROM.

[0028] Memory 12 stores each of the multiple sound source separation models used by the model determination unit 111 for selecting a sound source separation model, and at least one speech recognition model used by the speech recognition unit 113 for speech recognition processing. The multiple sound source separation models are trained models suitable for separating the speech of multiple users recorded in different usage environments. As an example, the memory 12 is described in which each of the multiple sound source separation models is stored, but it may also be stored in an external recording medium or external server that is communicatively connected to the processing unit P1.

[0029] The sound acquisition device 13 is implemented, for example, by a microphone. The sound acquisition device 13 acquires the speech of at least one user during the sound source separation model determination process, or multiple users US1 to US4 during operation, and converts them into electrical signals to generate audio data. The sound acquisition device 13 outputs the generated audio data to the processor 11. The sound acquisition device 13 may be configured separately from the processing unit P1.

[0030] The display interface 14 is configured using, for example, a Liquid Crystal Display (LCD) or an organic electroluminescence (EL). The display interface 14 displays the transcription results generated by the processor 11. The display interface 14 may be configured separately from the processing unit P1.

[0031] The database DB is connected to the processing unit P1 via the network NW, enabling data communication. The database DB stores the clean voice data Dt21, Dt22, Dt23, Dt24 (see Figure 5) for each of multiple speakers, along with information about the speaker and the correct answer data Wd21, Wd22, Wd23, Wd24 (see Figure 5) for the clean voice data, associated with each speaker.

[0032] The speaker information referred to here is speaker attribute information or features that indicate the speaker's individuality. Speaker attribute information includes, for example, the speaker's gender or age group (e.g., child, adult, or elderly). The ground truth data for the clean speech data is text data that shows the content of the utterances captured in the clean speech data.

[0033] <Overall operating procedure> Next, an example of the operation procedure of the processing device P1 will be described with reference to Figure 3. Figure 3 is a flowchart illustrating an example of the operation procedure of the processing device P1 in the embodiment.

[0034] The processor 11 acquires the audio data output from the sound pickup device 13 (St11). The processor 11 determines whether the sound source separation model corresponding to the current user environment has been determined, that is, whether the environment optimization process for the sound source separation model has been performed (St12).

[0035] Furthermore, the process in step St12 may accept user input via a user interface (e.g., touch panel, mouse, or keyboard) regarding the necessity of environment optimization processing for the sound source separation model, and may be executed if the user requests environment optimization processing for the sound source separation model. The process in step St12 may also determine that environment optimization processing for the sound source separation model has not been performed if it is determined that a predetermined period of time has elapsed since the last time environment optimization processing for the sound source separation model was performed.

[0036] If the processor 11 determines that the environment optimization process for the sound source separation model has been completed (St12, YES), it causes the sound source separation unit 112 to perform sound source separation using the sound source separation model determined by the environment optimization process.

[0037] On the other hand, if the processor 11 determines that the environment optimization process for the sound source separation model has not been performed (St12, NO), it instructs the model determination unit 111 to perform the environment optimization process for the sound source separation model (St13). The model determination unit 111 performs the environment optimization process for the sound source separation model shown in Figure 4 and determines a sound source separation model that corresponds to the current user environment. The processor 11 then instructs the sound source separation unit 112 to perform the sound source separation process using the sound source separation model determined by the environment optimization process.

[0038] The sound source separation unit 112 performs sound source separation on the audio data output from the sound acquisition device 13 using a sound source separation model (St14). The sound source separation unit 112 outputs each of the at least one separated data obtained by the sound source separation process to the speech recognition unit 113.

[0039] The speech recognition unit 113 performs speech recognition processing on each of at least one separated data, obtains the speech recognition result for each user, and outputs it to the output result generation unit 114 (St15). The output result generation unit 114 outputs the outputted speech recognition result for each user to the display interface 14 (St16).

[0040] As described above, the processing unit P1 in the embodiment can further improve the sound source separation performance for separating the voice data of multiple users by performing sound source separation processing using a sound source separation model that is more suitable for the user's usage environment, such as external disturbance factors like noise in the usage environment, the attributes or characteristics of the user to be separated, the transmission characteristics of the voice from the user to the sound collection device 13, or the performance of the sound collection device 13. By further improving the sound source separation performance, the processing unit P1 can more effectively improve the accuracy of speech recognition using each separated data.

[0041] <Environmental optimization processing procedure for sound source separation model> Next, the environmental optimization processing procedure for the sound source separation model will be described with reference to Figure 4. Figure 4 is a flowchart illustrating an example of the environmental optimization processing procedure for the sound source separation model of the processing device P1 in the embodiment.

[0042] First, let's explain the sound source separation performance in general sound source separation technologies.

[0043] The first audio data (D1) shown below is audio data obtained by recording the speech of a first user (S1) and the speech of a second user (S2) in a predetermined usage environment. The first audio data (D1) thus obtained includes disturbance factors (N) in the user's usage environment, audio data (=S1×R1) in which the speech transmission characteristics (R1) from the first user to the sound recording device are superimposed on the speech of the first user (S1), and audio data (=S2×R2) in which the speech transmission characteristics from the second user to the sound recording device are superimposed on the speech of the second user (S2).

[0044] Furthermore, the second audio data (D2) is obtained by recording the first user's speech (S1) in a predetermined usage environment, and mixing the recorded first user's speech (S1) with a pre-prepared clean speech of the second user (i.e., clean audio data S3). The second audio data (D2) thus acquired includes disturbance factors (N) in the user's usage environment, audio data (=S1×R1) in which the transmission characteristics of the speech from the first user to the sound recording device (R1) are superimposed on the first user's speech (S1), and the clean audio data of the second user (S3).

[0045] Each of these first and second audio data sets includes disturbance factors in a given usage environment and the transmission characteristics of sound from the user to the sound pickup device, respectively. Therefore, the sound source separation performance achieved using each of the first and second audio data sets will be approximately the same. The first audio data D1 = (S1 × R1) + (S2 × R2) + N The second audio data D2 = (S1 × R1) + S3 + N

[0046] As described above, the processing unit P1 can acquire external disturbances such as noise in the user environment, the attributes or characteristics of the user to be separated, the spatial characteristics of the location where the sound collection device 13 is installed, or the characteristics of the sound collection device 13, by collecting the voice of at least one user in the user's environment. In other words, the processing unit P1 can determine a sound source separation model with sufficient sound source separation performance by collecting the voice of at least one user in the user's environment and mixing it with at least one clean voice data similar to that of the user to be separated, even without collecting the speech of all users to be separated in the user's environment.

[0047] Therefore, the processing device P1 in this disclosure can determine a sound source separation model that is more suitable for the user's environment by performing an optimization process for the sound source separation model. The optimization process for the sound source separation model in this disclosure will be described below.

[0048] The model determination unit 111 acquires the audio data obtained in step St11 as audio data to be used for optimizing the sound source separation model (hereinafter referred to as "environmental audio data") (St131). The model determination unit 111 performs audio analysis of the acquired environmental audio data Dt11, or image analysis of the user's face image whose speech is captured in the environmental audio data Dt11 (St132).

[0049] The speech analysis processing of the environmental audio data Dt11 referred to here is a process that analyzes user attribute information using fundamental frequency analysis, or a process that extracts features indicating the individuality of the user's speech. The image analysis of the user's face image is a process that analyzes the user's face image captured by the camera 15 or captured in advance, and is a process that analyzes the user's attribute information. The user attribute information referred to here is the user's gender or age group (e.g., child, adult, or elderly). Note that the user attribute information may include multiple attribute information, such as "female, child" or "male, elderly".

[0050] The model determination unit 111 selects at least one clean audio data to be used in the optimization process of the sound source separation model based on speaker attribute information Cd21 to Cd24 or feature quantities (not shown) corresponding to each clean audio data Dt21 to Dt24 stored in the database DB, and user attribute information Cd11 or feature quantities (not shown) obtained by analyzing the environmental audio data Dt11 (St133). Specifically, the model determination unit 111 selects one or more clean audio data that are most similar to the user attribute information Cd21 to Cd24 or feature quantities (not shown), or multiple clean audio data with the highest similarity.

[0051] The model determination unit 111 generates a mixed voice Mx11 by mixing the environmental voice data Dt11 with at least one selected clean voice data Dt22 (St134). The characteristics of the sound pickup device 13 referred to here are the characteristics of the voice transmission from the user to the sound pickup device 13, or characteristics derived from the quality (performance) of the sound pickup device 13 itself.

[0052] The model determination unit 111 uses each of the multiple sound source separation models stored in the memory 12 to perform sound source separation (St135) by separating the generated mixed voice Mx11 into user speech (environmental voice data Dt11) and clean voice data.

[0053] The model determination unit 111 acquires each of the multiple separated data sets separated using each of the multiple sound source separation models. The model determination unit 111 then performs speech recognition on each of the multiple separated data sets for each sound source separation model. The speech recognition model used here is a pre-trained model usable by the speech recognition unit 113. This allows the processing unit P1 to determine a sound source separation model suitable not only for the user's environment but also for the pre-trained model used by the speech recognition unit 113 of the processing unit P1.

[0054] The model determination unit 111 acquires the speech recognition results for each of the multiple separated data sets for each sound source separation model performed by the speech recognition unit 113. The model determination unit 111 calculates the degree of agreement between the speech recognition result corresponding to the mixed clean audio data Dt22 from among the speech recognition results for each sound source separation model and the correct data Wd22 of the clean audio data Dt22, and performs an evaluation of the sound source separation model based on the calculated degree of agreement (St136).

[0055] The model determination unit 111 determines the sound source separation model with the highest degree of agreement as the sound source separation model best suited to the user's environment (St137). If there are multiple clean audio data sets to be mixed, the processing unit P1 calculates the degree of agreement between the speech recognition result corresponding to the clean audio data and the correct data for the clean audio data for each sound source separation model. The processing unit P1 may determine the sound source separation model with the largest overall value, average value, or median of the calculated degree of agreement for each sound source separation model as the sound source separation model best suited to the user's environment.

[0056] Hereinafter, we have described an example in which the processing unit P1 performs the process of determining a sound source separation model using environmental speech data Dt11, which is obtained by capturing the user's free speech, but it is not limited to this example. The environmental speech data Dt11 may also be speech data in which at least one user reads a predefined sentence. In such a case, the processing unit P1 may determine the sound source separation model that is optimal for the user's environment based on the degree of agreement between the speech recognition result corresponding to the environmental speech data Dt11 and the predefined sentence, and the degree of agreement between the speech recognition result corresponding to the clean speech data Dt22 and the correct answer data.

[0057] Furthermore, the processing unit P1 may accept input of correct data corresponding to the environmental sound data Dt11 via a user interface. In such a case, the processing unit P1 may determine the optimal sound source separation model for the user's environment based on the degree of agreement between the speech recognition result corresponding to the environmental sound data Dt11 and the input correct data, and the degree of agreement between the speech recognition result corresponding to the clean sound data Dt22 and the correct data.

[0058] As described above, the processing unit P1 in the embodiment can determine a sound source separation model that can more accurately separate the voice of the user being separated from audio data containing the speech of multiple users by selecting one or more clean audio data similar to attribute information or features similar to at least one user being separated from. In particular, when the attribute information of multiple users being separated from is similar or identical, or when the features of multiple users are similar, it becomes difficult for the processing unit P1 to separate the voice of each user. Therefore, the processing unit P1 can determine (select) a speech separation model suitable for separating voices that are more difficult to separate by selecting one or more clean audio data similar to attribute information or features similar to at least one user being separated from.

[0059] Furthermore, the sound source separation model determination system 100 in this embodiment manages the clean audio data Dt21 to Dt24 used to determine the sound source separation model by associating them with the correct answer data Wd21 to Wd24. As a result, the processing unit P1 can evaluate which of the multiple sound source separation models stored in the database DB is more suitable for the user's environment, without needing the correct answer data for the user's utterance content Wd11 included in the environmental audio data Dt11.

[0060] Furthermore, in this embodiment, the processing unit P1 evaluates the sound source separation model based on the speech recognition results of the trained model used by the speech recognition unit 113 for speech recognition. This allows the processing unit P1 to determine a sound source separation model that is suitable not only for the user's environment but also for the trained model used by the speech recognition unit 113 of the processing unit P1.

[0061] <How to select clean audio data> Next, we will explain the method for selecting clean audio data, referring to Figure 5. Figure 5 is a diagram showing an example of a method for selecting clean audio data.

[0062] As shown in Figure 5, the database DB stores, in association with at least one speaker attribute information Cd21-Cd24 and correct answer data Wd21-Wd24 for each of the speaker's clean voice data Dt21-Dt24. Note that the utterances (correct answer data Wd21-Wd24) for each of the clean voice data Dt21-Dt24 may be the same or different.

[0063] For example, clean voice data Dt21 is associated with the speaker attribute information Cd21 "adult female" and the correct utterance data Wd21 "It will be sunny tomorrow". Clean voice data Dt22 is associated with the speaker attribute information Cd22 "adult male" and the correct utterance data Wd22 "It's a nice day today". Clean voice data Dt23 is associated with the speaker attribute information Cd23 "elderly" and the correct utterance data Wd23 "It's very cold today". Clean voice data Dt24 is associated with the speaker attribute information Cd24 "child" and the correct utterance data Wd24 "It's a nice day today".

[0064] In the example shown in Figure 5, the model determination unit 111 selects one clean voice data Dt22 that has attribute information most similar to the user attribute information Cd11 "adult male" obtained by analyzing the environmental voice data Dt11. Although not shown in Figure 5, the model determination unit 111 may also select one or more clean voice data that have feature quantities similar to the user feature quantities obtained by analyzing the environmental voice data Dt11, or it may select one or more clean voice data that have attribute information and feature quantities similar to the user attribute information Cd11 and feature quantities obtained by analyzing the environmental voice data Dt11.

[0065] Furthermore, if the number of users to be separated is N (N: an integer of 3 or more), the model determination unit 111 may perform environmental optimization processing of the sound source separation model using environmental audio data that includes the speech of multiple users, but fewer than N users. In such a case, the model determination unit 111 may analyze the speech of each of the fewer than N users and select one or more clean audio data by comparing each of the attribute information or features of the multiple users with each of the speaker-specific attribute information or features stored in the database DB.

[0066] Furthermore, if the model determination unit 111 finds that the attribute information of fewer than N users obtained from the environmental voice data is "adult female" and "adult male", it may select one or more clean voice data Dt21, Dt22 that have attribute information identical or similar to each attribute information, or it may select one or more clean voice data that have attribute information identical or similar to the most frequent attribute information among the attribute information of fewer than N users obtained.

[0067] Similarly, the model determination unit 111 may select one or more clean voice data sets having features similar to the features of each of the fewer than N users obtained from the environmental voice data, or it may cluster the acquired user features and select one or more clean voice data sets having features identical or similar to the features of the cluster with the most features.

[0068] <Method for mixing ambient sound data and clean sound data> Next, referring to Figure 6, we will explain how to mix the ambient audio data Dt11 and the clean audio data Dt22. Figure 6 shows examples 1 to 3 of mixing the ambient audio data Dt11 and the clean audio data Dt22. Note that the utterance Wd11 "XXXXXXXXXXXXXXX" indicates that the correct utterance is unknown.

[0069] The processing unit P1 mixes one or more ambient audio data with one or more clean audio data using a pre-configured or user-selected mixing method. While this disclosure shows the processing unit P1 as an example of mixing one ambient audio data Dt11 with one clean audio data Dt22, it is not limited to this example. Furthermore, when mixing a total of three or more audio data, such as when mixing one ambient audio data with multiple clean audio data, the processing unit P1 may perform a mixing process by arbitrarily combining mixing examples 1 to 3 described later.

[0070] In Mix Example 1, the processing unit P1 mixes the environmental audio data Dt11 and the clean audio data Dt22 so that their total lengths L11 and L22 overlap. In this case, the mixed audio Mx11 is audio data in which the total length L11 of the utterance Wd11 "XXXXXXXXXXXXXXX" from the environmental audio data Dt11 and the total length L22 of the utterance (correct data Wd22 "It's a nice day today") from the clean audio data Dt22 overlap. In other words, the mixed audio Mx11 is audio data that includes utterance Wd31, which is a mixture of utterance Wd11 "XXXXXXXXXXXXXXX" and correct data Wd22 "It's a nice day today".

[0071] In Mix Example 2, the processing unit P1 mixes the environmental audio data Dt11 and the clean audio data Dt22 so that they partially overlap. In this case, the mixed audio Mx12 is audio data in which the utterance Wd11 "XXXXXXXXXXXXXXX" from the environmental audio data Dt11 and the utterance (correct data Wd22 "It's a nice day today") from the clean audio data Dt22 partially overlap. In other words, the mixed audio Mx12 is audio data that includes utterance Wd32, which is a partial mixture of utterance Wd11 "XXXXXXXXXXXXXXX" and correct data Wd22 "It's a nice day today".

[0072] In Mix Example 3, the processing unit P1 mixes the ambient audio data Dt11 and the clean audio data Dt22 in a concatenated manner. In this case, the mixed audio Mx13 is audio data in which the utterance Wd11 "XXXXXXXXXXXXXXX" from the ambient audio data Dt11 and the utterance (correct data Wd22 "It's a nice day today") from the clean audio data Dt22 are concatenated without overlap. In other words, the mixed audio Mx13 is audio data that includes utterance Wd33, which is a partial mixture of utterance Wd11 "XXXXXXXXXXXXXXX" and correct data Wd22 "It's a nice day today".

[0073] As described above, the processing device P1 in the embodiment can obtain audio data (mixed audio) equivalent to audio data spoken by the user corresponding to the environmental audio data Dt11 and the speaker corresponding to the clean audio data in the same usage environment by mixing one or more clean audio data with the environmental audio data Dt11 of at least one user that is the target of sound source separation.

[0074] Furthermore, the processing unit P1 in this embodiment can acquire audio data equivalent to audio data spoken by a user corresponding to the ambient audio data Dt11 and a speaker corresponding to the clean audio data in the same usage environment (mixed audio). In addition, the processing unit P1 can increase the difficulty of sound source separation as the length of the overlap section L22 of the clean audio data Dt22 with the total length L11 of the ambient audio data Dt11 increases. In other words, the processing unit P1 can adjust the sound source separation performance determined by the length of the overlap section between the ambient audio data Dt11 and the clean audio data Dt22.

[0075] (Note) The following technologies are disclosed based on the above description of embodiments.

[0076] (Technology 1) An acquisition unit (processor 11) acquires audio data (environmental audio data Dt11) which is the voice of one user speaking in a predetermined usage environment, A selection unit (model determination unit 111) selects at least one clean voice data from among the clean voice data of multiple speakers, A mixing unit (model determination unit 111) mixes the aforementioned audio data (environmental audio data Dt11) and the aforementioned clean audio data, A separation unit (model determination unit 111) separates the mixed audio (for example, the mixed audio Mx11~Mx13 shown in Figure 6) into separation data corresponding to the audio data and separation data for each of the clean audio data, using each of multiple sound source separation models. A speech recognition unit 113 performs speech recognition on each of the separated data, The system includes a model determination unit 111 that, based on the speech recognition results of the separated data corresponding to the clean audio data among the plurality of separated data, determines and outputs a sound source separation model from each of the sound source separation models to be used for sound source separation of the audio data (environmental audio data Dt11) acquired in the predetermined usage environment. Sound source separation model determination device (processing device P1). As a result, the processing unit P1 can obtain audio data (mixed audio) equivalent to audio data spoken by the user corresponding to the environmental audio data Dt11 and the speaker corresponding to the clean audio data in the same usage environment by mixing one or more clean audio data with the environmental audio data Dt11 of one user that is the target of sound source separation. Therefore, the processing unit P1 can evaluate and determine a sound source separation model that is more suitable for the user's usage environment using only the speech of one user, without having to collect the speech of all of the multiple users that are actually the target of sound source separation.

[0077] (Technology 2) The clean voice data Dt21 to Dt24 are associated with the correct speech content Wd21 to Wd24 recorded in the clean voice data Dt21 to Dt24. The model determination unit 111 determines and outputs the sound source separation model based on the degree of agreement between the speech recognition result of the separated data corresponding to the clean audio data and the correct answer data. A device (processing device P1) for determining the sound source separation model described in (Technology 1). As a result, even if the correct speech of the user that has been recorded is unknown, the processing unit P1 can evaluate the source separation performance for each source separation model based on the ground truth data corresponding to the clean speech data, since the ground truth data corresponding to the clean speech data is known.

[0078] (Technology 3) The selection unit (model determination unit 111) selects the clean voice data which is similar to the voice data (environmental voice data Dt11). A device (processing device P1) for determining the sound source separation model described in (Technology 1) or (Technology 2). As a result, the processing unit P1 can select clean audio data, which is more difficult to separate from the sound source, as the mixing target, thereby determining (selecting) an audio separation model that is suitable for separating audio that is more difficult to separate, that is, an audio separation model with higher sound source separation performance.

[0079] (Technology 4) The clean voice data Dt21~Dt24 is associated with the speaker attribute information Cd21~Cd24. The selection unit (model determination unit 111) analyzes the user's attribute information based on the voice data (environmental voice data Dt11) and selects the clean voice data having speaker attribute information that is the same as or similar to the analyzed user's attribute information. A device (processing device P1) for determining the sound source separation model described in any one of (Technology 1) to (Technology 3). This allows the processing unit P1 to select clean audio data, which is more difficult to separate from the sound source, as the target for mixing, based on the user's attribute information. Therefore, by selecting clean audio data, which is more difficult to separate from the sound source, as the target for mixing, the processing unit P1 can determine (select) an audio separation model that is suitable for separating audio that is more difficult to separate, that is, an audio separation model with higher sound source separation performance.

[0080] (Technology 5) The system further includes an imaging unit (camera 15) for imaging the user, The clean voice data has the speaker's attribute information, The selection unit (model determination unit 111) analyzes the user's attribute information based on the user's captured image (face image) captured by the imaging unit (camera 15), and selects the clean voice data having speaker attribute information that is the same as or similar to the analyzed user attribute information. A device (processing device P1) for determining the sound source separation model described in any one of (Technology 1) to (Technology 3). This allows the processing unit P1 to select clean audio data, which is more difficult to separate from the sound source, as the target for mixing, based on the user's attribute information. Therefore, by selecting clean audio data, which is more difficult to separate from the sound source, as the target for mixing, the processing unit P1 can determine (select) an audio separation model that is suitable for separating audio that is more difficult to separate, that is, an audio separation model with higher sound source separation performance.

[0081] (Technology 6) The selection unit (model determination unit 111) analyzes the user's features based on the audio data (environmental audio data Dt11) and selects the clean audio data having features similar to the analyzed user's features. A device (processing device P1) for determining the sound source separation model described in any one of (Technology 1) to (Technology 3). This allows the processing unit P1 to select clean speech data, which is more difficult to separate from the source, as the target for mixing, based on the feature quantities of the user's speech. Therefore, by selecting clean speech data, which is more difficult to separate from the source, as the target for mixing, the processing unit P1 can determine (select) a speech separation model that is suitable for separating speech that is more difficult to separate, that is, a speech separation model with higher sound source separation performance.

[0082] (Technology 7) The mixing unit (model determination unit 111) generates a mixed audio (for example, the mixed audio Mx11, Mx12 shown in Figure 6) in which at least a portion of the audio data (environmental audio data Dt11) and the clean audio data overlap. A device (processing device P1) for determining the sound source separation model described in any one of (Technology 1) to (Technology 6). As a result, the processing unit P1 can acquire audio data (mixed audio) equivalent to audio data spoken by a user corresponding to the environmental audio data Dt11 and a speaker corresponding to the clean audio data in the same usage environment. Therefore, the processing unit P1 can determine a sound source separation model suitable for the user's usage environment.

[0083] (Technology 8) The mixing unit (model determination unit 111) generates the mixed audio (for example, the mixed audio Mx11 to Mx13 shown in Figure 6) by concatenating the audio data (environmental audio data Dt11) and the clean audio data. A device (processing device P1) for determining the sound source separation model described in any one of (Technology 1) to (Technology 6). As a result, the processing unit P1 can acquire audio data (mixed audio) equivalent to audio data spoken by a user corresponding to the environmental audio data Dt11 and a speaker corresponding to the clean audio data in the same usage environment. Therefore, the processing unit P1 can determine a sound source separation model suitable for the user's usage environment.

[0084] (Technology 9) A method for determining a sound source separation model performed by at least one processor 11, The system acquires audio data (environmental audio data Dt11) containing the speech of one user speaking in a specified environment. Select at least one clean audio data from the clean audio data of each of the multiple speakers, The aforementioned audio data (environmental audio data Dt11) and the clean audio data are mixed together. The mixed audio (for example, the mixed audio Mx11~Mx13 shown in Figure 6) is separated into separate data corresponding to the audio data and separate data for each of the clean audio data using each of the multiple sound source separation models, and each of the separated sets of separate data is subjected to speech recognition. Based on the speech recognition results of the separated data corresponding to the clean audio data among the plurality of separated data, the system determines and outputs the sound source separation model used for sound source separation of the audio data (environmental audio data Dt11) acquired in the predetermined usage environment, from among the sound source separation models. Method for determining the sound source separation model. As a result, the processing unit P1 can obtain audio data (mixed audio) equivalent to audio data spoken by the user corresponding to the environmental audio data Dt11 and the speaker corresponding to the clean audio data in the same usage environment by mixing one or more clean audio data with the environmental audio data Dt11 of one user that is the target of sound source separation. Therefore, the processing unit P1 can evaluate and determine a sound source separation model that is more suitable for the user's usage environment using only the speech of one user, without having to collect the speech of all of the multiple users that are actually the target of sound source separation.

[0085] (Technology 10) A database DB that stores clean audio data for each of multiple speakers, A sound source separation model determination system 100 comprising a processing device P1 capable of communicating with the aforementioned database DB, The processing unit P1 is Audio data (environmental audio data Dt11) is obtained from the voices of multiple users speaking in a specified environment. Select and acquire at least one clean voice data from among the clean voice data of each of the multiple speakers stored in the database DB. The aforementioned audio data (environmental audio data Dt11) and the clean audio data are mixed together. The mixed audio (for example, the mixed audio Mx11~Mx13 shown in Figure 6) is separated into separate data corresponding to the audio data and separate data for each of the clean audio data using each of the multiple sound source separation models, and each of the separated sets of separate data is subjected to speech recognition. Based on the speech recognition results of the separated data corresponding to the clean audio data among the plurality of separated data, the system determines and outputs the sound source separation model used for sound source separation of the audio data (environmental audio data Dt11) acquired in the predetermined usage environment, from among the sound source separation models. Sound source separation model determination system 100. As a result, the sound source separation model determination system 100 can obtain audio data (mixed audio) equivalent to audio data spoken by the user corresponding to the environmental audio data Dt11 and the speaker corresponding to the clean audio data in the same usage environment, by mixing one or more clean audio data with the environmental audio data Dt11 of one user who is the target of sound source separation. Therefore, the sound source separation model determination system 100 can evaluate and determine a sound source separation model that is more suitable for the user's usage environment using only the audio of one user, without having to collect the speech of all of the multiple users who are the actual target of sound source separation.

[0086] (Technology 11) A program for determining a sound source separation model, which is executed by at least one processor 11, The steps include: acquiring audio data (environmental audio data Dt11) in which the speech of one user speaking in a predetermined usage environment is captured; The steps include selecting at least one clean audio data from among the clean audio data of multiple speakers, The steps include mixing the aforementioned audio data (environmental audio data Dt11) and the clean audio data, and separating the mixed audio (for example, the mixed audio Mx11~Mx13 shown in Figure 6) into separate data corresponding to the aforementioned audio data and separate data for each of the clean audio data, using each of the multiple sound source separation models, The steps include: performing speech recognition on each of the separated data sets, To realize the step of determining and outputting a sound source separation model from each of the sound source separation models that is used for sound source separation of the audio data (environmental audio data Dt11) acquired in the predetermined usage environment, based on the speech recognition result of the separated data corresponding to the clean audio data among the plurality of separated data, A program for determining the sound source separation model. As a result, the processing unit P1 can obtain audio data (mixed audio) equivalent to audio data spoken by the user corresponding to the environmental audio data Dt11 and the speaker corresponding to the clean audio data in the same usage environment by mixing one or more clean audio data with the environmental audio data Dt11 of one user that is the target of sound source separation. Therefore, the processing unit P1 can evaluate and determine a sound source separation model that is more suitable for the user's usage environment using only the speech of one user, without having to collect the speech of all of the multiple users that are actually the target of sound source separation.

[0087] Although various embodiments have been described above with reference to the drawings, it goes without saying that this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the various embodiments described above can be combined arbitrarily without departing from the spirit of the invention. [Industrial applicability]

[0088] This disclosure is useful as a sound source separation model determination device, a sound source separation model determination method, a sound source separation model determination system, and a sound source separation model determination program for determining a sound source separation model more suitable for the user's environment. [Explanation of symbols]

[0089] 10 Communications Department 11 processors 12 memory 13. Sound collection device 14 Display I / F 15 Cameras 100 Sound Source Separation Model Determination System 111 Model Determination Section 112 Sound source separation section 113 Voice Recognition Unit 114 Output result generation section Cd11, Cd21, Cd22, Cd23, Cd24 Attribute Information DB Database Dt11 Environmental Audio Data Dt21, Dt22, Dt23, Dt24 Clean Audio Data Mx11, Mx12, Mx13 mixed audio NW Network P1 Processing Unit US1, US2, US3, US4 users Wd21, Wd22, Wd23, Wd24 Correct Data

Claims

1. An acquisition unit that acquires audio data in which the speech of one user speaking in a predetermined usage environment is captured, A selection unit that selects at least one clean audio data from among the clean audio data of multiple speakers, A mixing unit that mixes the aforementioned audio data and the aforementioned clean audio data, A separation unit that separates the mixed audio into separate data corresponding to the audio data and separate data for each of the clean audio data, using each of multiple sound source separation models, A speech recognition unit that performs speech recognition on each of the separated data sets, The system includes a model determination unit that, based on the speech recognition results of the separated data corresponding to the clean audio data among the plurality of separated data, determines and outputs a sound source separation model from each of the sound source separation models to be used for sound source separation of the audio data acquired in the predetermined usage environment. A device for determining the sound source separation model.

2. The clean audio data is associated with the correct data of the spoken content captured in the clean audio data. The model determination unit determines and outputs the sound source separation model based on the degree of agreement between the speech recognition result of the separated data corresponding to the clean audio data and the ground truth data. A device for determining the sound source separation model according to claim 1.

3. The selection unit selects clean audio data that is similar to the audio data. A device for determining the sound source separation model according to claim 1.

4. The clean audio data is associated with the speaker's attribute information. The selection unit analyzes the user's attribute information based on the audio data and selects the clean audio data having speaker attribute information that is identical or similar to the analyzed user's attribute information. A device for determining the sound source separation model according to claim 1.

5. The system further comprises an imaging unit for imaging the user, The clean voice data has the speaker's attribute information, The selection unit analyzes the user's attribute information based on the user's image captured by the imaging unit, and selects the clean voice data having speaker attribute information that is identical or similar to the analyzed user attribute information. A device for determining the sound source separation model according to claim 1.

6. The selection unit analyzes the user's features based on the audio data and selects the clean audio data having features similar to the analyzed user features. A device for determining the sound source separation model according to claim 1.

7. The mixing unit generates the mixed audio in which at least a portion of the audio data and the clean audio data overlap. A device for determining the sound source separation model according to claim 1.

8. The mixing unit generates the mixed audio by concatenating the audio data and the clean audio data. A device for determining the sound source separation model according to claim 1.

9. A method for determining a sound source separation model performed by at least one processor, The system acquires audio data containing the speech of one user speaking in a specified environment. Select at least one clean audio data from the clean audio data of multiple speakers, The aforementioned audio data and the clean audio data are mixed together. The mixed audio is separated into separate data corresponding to the audio data and separate data for each of the clean audio data using each of the multiple sound source separation models, and each of the separated sets of separate data is subjected to speech recognition. Based on the speech recognition results of the separated data corresponding to the clean audio data among the plurality of separated data, the system determines and outputs the sound source separation model used for sound source separation of the audio data acquired in the predetermined usage environment from among the sound source separation models. Method for determining the sound source separation model.

10. A database that stores clean voice data from multiple speakers, A sound source separation model determination system comprising a processing device capable of communicating with the aforementioned database, The aforementioned processing apparatus is The system acquires audio data containing the speech of one user speaking in a specified environment. Select and acquire at least one clean voice data from among the clean voice data of each of the multiple speakers stored in the database. The aforementioned audio data and the clean audio data are mixed together. The mixed audio is separated into separate data corresponding to the audio data and separate data for each of the clean audio data using each of the multiple sound source separation models, and each of the separated sets of separate data is subjected to speech recognition. Based on the speech recognition results of the separated data corresponding to the clean audio data among the plurality of separated data, the system determines and outputs the sound source separation model used for sound source separation of the audio data acquired in the predetermined usage environment from among the sound source separation models. Sound source separation model determination system.

11. A program for determining a sound source separation model, which is executed by at least one processor, The steps include: acquiring audio data in which the speech of one user speaking in a predetermined usage environment is captured; The steps include selecting at least one clean audio data from among the clean audio data of multiple speakers, The steps include mixing the aforementioned audio data and the clean audio data, and separating the mixed audio into separate data corresponding to the aforementioned audio data and separate data for each of the clean audio data using each of a plurality of sound source separation models, The steps include: performing speech recognition on each of the separated data sets, To realize the step of determining and outputting, based on the speech recognition result of the separated data corresponding to the clean audio data among the plurality of separated data, which of the sound source separation models is used for sound source separation of the audio data acquired in the predetermined usage environment, A program for determining the sound source separation model.