Information processing method and information processing system

The method addresses format and localization issues in target sound extraction by using reference-based beamforming and super-resolution processing to maintain the quality of extracted sounds.

WO2026100237A1PCT designated stage Publication Date: 2026-05-15SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2025-09-26
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Conventional target sound extraction methods often change the format of the extracted sound information, such as the number of channels, frequency, and localization, when extracting target sounds from mixed sounds, which can degrade the quality of the output.

Method used

An information processing method that includes a restoration step to maintain the format of the extracted sound information by using reference-based beamforming and super-resolution processing to adjust the format of the query-based target sound extraction result.

Benefits of technology

The method effectively extracts target sounds while preserving the original format and localization of the retrieved sound information, ensuring high-quality output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025034040_15052026_PF_FP_ABST
    Figure JP2025034040_15052026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing method is executed by a computer, and comprises: a reception step for receiving an input of a query that is information for searching sound information; a search step for searching the sound information on the basis of the received query; a target sound extraction step for extracting, from the searched sound information, a target sound that is sound information corresponding to the received query; and a restoration step for restoring, on the basis of the searched sound information, a format of a query-based target sound extraction result that is an output in the target sound extraction step.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Method and Information Processing System

[0001] The present disclosure relates to an information processing method and an information processing system.

[0002] Techniques for searching for sound information such as audio data like waveform files are known. Further, there are techniques for extracting or emphasizing a target sound, which is the sound information to be extracted, from mixed sound that is sound information in which the target sound and interfering sound, which is the sound information to be removed, are mixed. In other words, there are techniques for removing or attenuating interfering sound, which is sound other than the target sound (see, for example, Patent Documents 1 to 3 below). "Extracting or emphasizing the target sound" and "removing or attenuating interfering sound, which is sound other than the target sound" are synonymous. Hereinafter, it will basically be expressed as "extracting the target sound".

[0003] Further, as a different type of target sound extraction from these patent documents, there is also known a technique called query-based target sound extraction in which various modal data such as text, images, and audio are given as a query (hint information regarding the target sound), and the sound information corresponding to the query, that is, the target sound, can be extracted from the mixed sound (see, for example, Non-Patent Documents 1 to 4 below).

[0004] Japanese Patent Application Laid-Open No. 2021-152623 International Publication No. 2022 / 190615 Japanese Patent Application Laid-Open No. 2010-233173

[0005] Hao-Wen Dong and Naoya Takahashi and Yuki Mitsufuji and Julian McAuley and Taylor Berg-Kirkpatrick, “CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos,” Proceedings of International Conference on Learning Representations (ICLR), 2023Liu, Xubo and Liu, Haohe and Kong, Qiuqiang and Mei, Xinhao and Zhao, Jinzheng and Huang, Qiushi, and Plumbley, Mark D and Wang, Wenwu, “Separate What You Describe: Language-Queried Audio Source Separation,” Proc. Interspeech, 1411-1805, 2022Ohishi Y, Delcroix M, Ochiai T, Araki S, Takeuchi D, Niizumi D, Kimura A et al., “ConceptBeam: Concept Driven Target Speech Extraction,” Proceedings of the 30th ACM International Conference on Multimedia, 2022M. Delcroix, J. B. Vazquez, T. Ochiai, K. Kinoshita, Y. Ohishi, S. Araki, "SoundBeam: Target Sound Extraction Conditioned on Sound-Class Labels and Enrollment Clues for Increased Performance and Continuous Learning," IEEE / ACM Trans. on Audio, Speech, and Language Processing, Vol. 31, pp.121-136, 2023.

[0006] Conventional technology can extract a target sound from a mixed sound. Furthermore, by combining it with sound information retrieval technology, target sound extraction can be applied to the search results, which are the retrieved sound information. However, a side effect is that the format of the extracted target sound (the target sound extraction result) may change between the extracted target sound and the aforementioned search results. For example, the format of the target sound, such as the number of channels, frequency, and localization, may be converted.

[0007] Therefore, this disclosure aims to propose an information processing method and information processing system that can extract target sounds while maintaining the format of the retrieved sound information.

[0008] The information processing method relating to this disclosure is an information processing method performed by a computer and includes: a reception step of receiving input which is a query which is information for searching for sound information; a search step of searching for sound information based on the received query; a target sound extraction step of extracting a target sound which is sound information corresponding to the received query from the searched sound information; and a restoration step of restoring the format of the query-based target sound extraction result which is the output of the target sound extraction step, based on the searched sound information.

[0009] This is a diagram illustrating the outline of an information processing system according to an embodiment. This is a block diagram illustrating an example of the configuration of an information processing device according to an embodiment. This is a diagram illustrating the control unit. This is a diagram illustrating the search unit. This is a diagram illustrating the query-based target sound extraction unit. This is a diagram illustrating the format and localization restoration unit. This is a diagram (1) illustrating an example of reference processing. This is a diagram (2) illustrating an example of reference processing. This is a flowchart (1) showing an example of the information processing flow according to an embodiment. This is a flowchart (2) showing an example of the information processing flow according to an embodiment. This is a flowchart (3) showing an example of the information processing flow according to an embodiment. This is a flowchart (4) showing an example of the information processing flow according to an embodiment. This is a block diagram illustrating an example of the configuration of an information processing device according to a first modification. This is a diagram illustrating an example of waveform division. This is a diagram illustrating an example of conversion of multiple partial waveforms. This is a flowchart illustrating an example of the search processing flow according to a first modification. This is a flowchart illustrating an example of the detailed matching flow. This is a block diagram illustrating an example of the configuration of an information processing device according to a second modification. This is a diagram illustrating an example of search results when detailed matching is performed. This is a block diagram illustrating an example of the configuration of an information processing device according to a third modification. This is a diagram illustrating a first example of searching for sound information based on multiple queries. This is a diagram illustrating a second example of searching for sound information based on multiple queries. This is a hardware configuration diagram showing an example of a computer that implements the functions of an information processing device.

[0010] The embodiments of this disclosure will be described in detail below with reference to the drawings. In the following embodiments, the same parts will be denoted by the same reference numerals, and redundant descriptions will be omitted.

[0011] The embodiments of this disclosure will be described below in the following order: 1. Embodiments 1-1. Overview of the information processing system according to the embodiment 1-2. Configuration of the information processing device according to the embodiment 1-3. Flow of information processing according to the embodiment 2. Modifications 2-1. First modification 2-2. Second modification 2-3. Third modification 3. Other embodiments 4. Effects of the information processing method according to this disclosure 5. Hardware configuration 6. Supplementary information

[0012] (1. Embodiments) (1-1. Overview of the Information Processing System According to the Embodiment) An overview of the information processing system 1 according to the embodiment will be described using Figure 1. Figure 1 is a diagram for explaining the overview of the information processing system according to the embodiment.

[0013] Information processing system 1 includes an information processing device 100. The information processing device 100 is, for example, a server. Specifically, the information processing device 100 is an audio data retrieval device that searches for sound information, such as audio data, based on queries. As an example, the information processing device 100 is a cross-modal audio retrieval device that searches for audio data corresponding to various modal queries (for example, audio data such as waveform files, text, images, etc.).

[0014] Sound information can be any audio data, such as human voices, sounds emitted by animals or insects, sound effects, music, or ambient sounds. The sound information can be digital audio signals, such as waveform files, or analog audio. A cross-modal audio search device is a system that searches for audio data, such as waveform files, that corresponds to various modal queries. It is a device that can use non-audio modals such as text and images, as well as audio data, as queries.

[0015] The following describes an example of the process by which the information processing device 100 searches for sound information based on a query, specifically the process performed in steps S1 to S4 of Figure 1.

[0016] The information processing device 100 accepts the input of a query 200 (step S1). For example, the information processing device 100 accepts the input of text from the user as a query 200, which is "a dog barking in a park," and is intended to search for audio data related to "a dog barking in a park."

[0017] Next, the information processing device 100 searches for sound information based on the received query 200 (step S2). For example, the information processing device 100 searches for the audio data 400 to be searched as sound information from the audio database 300, which is a database that stores various waveform files in various formats.

[0018] Specifically, the information processing device 100 searches the audio database 300 for audio data 400, which includes audio data 401 most similar to the target sound corresponding to the query 200, up to the Nth most similar audio data 40N. Hereinafter, N is an integer greater than or equal to 1.

[0019] At least one of the audio data 401 to 40N may contain interfering sounds, which are sound information that does not correspond to the query. For example, if audio data 401 is ambient sound recorded in a park, it may include not only the target sound, which is sound information corresponding to query 200, such as a dog barking, but also interfering sounds such as the voices of people playing in the park or birdsong.

[0020] The audio data 401 to 40N mentioned above contain the target sound, so the search results are not incorrect. On the other hand, if a user intends to use such audio data as sound effects in movies or sound dramas, it would be preferable to have the interfering sounds removed. However, since what constitutes the target sound and what constitutes the interfering sound depends on the search query, it is not possible to apply a process to remove interfering sounds in advance to the data contained in the audio database 300.

[0021] Therefore, the information processing device 100 performs a process to extract sound information corresponding to the received query 200 from the retrieved sound information (step S3). For example, after receiving the input of the query 200, the information processing device 100 generates a query-based target sound extraction result 450 by extracting the target sound, which is the sound information corresponding to the query 200, from the audio data 400, which is the retrieved sound information.

[0022] Specifically, as a preprocessing step for extraction, the information processing device 100 performs format conversion as needed, such as reducing the number of channels or frequency of the audio data 401 that is most similar to the target sound among the audio data 401 to audio data 40N. Subsequently, the information processing device 100 inputs the format-converted audio data 401 and the query 200 into a machine learning model for target sound extraction, thereby generating sound information in which the target sound has been extracted or emphasized from the audio data 401, i.e., a query-based target sound extraction result 450.

[0023] The information processing device 100 can extract the target sound from the mixed sound, which is sound information in which the target sound and interfering sound are mixed, by processing up to step S3. However, the technique of extracting the target sound from the mixed sound alone may not be able to extract the target sound while maintaining the format of the retrieved sound information, in which case the format may change as a side effect. For example, the query-based target sound extraction result 450 may have a format that has been converted compared to the audio data 400, such as the number of channels, frequency, and localization of the target sound.

[0024] Specifically, the following problems exist regarding the number of channels and frequency. When extracting a target sound by inputting the format-converted audio data 401 and query 200 into a machine learning model for target sound extraction, as in step S3, the input and output formats of the machine learning model for target sound extraction depend on the training data. Therefore, in order to handle the diverse formats of the retrieved sound information, one of the following can be considered: a) Train a machine learning model for target sound extraction for each format. During inference, i.e., when extracting the target sound, switch the machine learning model for target sound extraction according to the format of the search results. b) Fix the input and output formats of the machine learning model for target sound extraction. Before inference, match the format of the search results to the input format for target sound extraction.

[0025] However, achieving a) becomes difficult as the number of search result formats increases. On the other hand, b) is easily achieved if the search results are converted in a way that reduces the number of channels or sampling frequency. For example, if the input and output format of the machine learning model for extracting target sounds is fixed to mono and 16 kHz, and the number of channels or sampling frequency of the search results is high, it is easy to perform processing such as mixdown or downsampling. However, when such conversion processing is performed, the information deteriorates compared to the original search results.

[0026] Furthermore, there are the following issues regarding localization. Assuming that multiple microphones are used, it is possible to record audio data that expresses the direction and position of each sound source. For example, it is possible to record data where a dog barking sounds from the left and a bird chirp sounds from the right. In addition, audio data containing a mixture of such dog and bird sounds is presented as a search result for the query "dog barking" (200).

[0027] In this example, the user might want to remove bird sounds but maintain the sense of localization, such as the dog barking coming from the left. However, if the machine learning model for target sound extraction is trained on monaural data, maintaining localization is not considered in the resulting target sound extraction output.

[0028] Three methods can be considered to adapt a machine learning model for mono-compatible target sound extraction to handle multi-channel search results. However, maintaining a sense of localization is difficult in any of these cases. i) Mix down the search results to mono before inputting them into the machine learning model for target sound extraction. ii) Treat the search results as C independent mono audio data (where C is the number of channels in the search results), and input each of the C audio data into the machine learning model for target sound extraction. From the C outputs, C-channel audio data is reconstructed. iii) If training data recorded with the same microphone array as the search results is prepared and the machine learning model for target sound extraction is trained, it may be possible to generate target sound extraction results that maintain a sense of localization. However, since there are countless patterns for microphone array arrangements, it is difficult to prepare such training data.

[0029] In order to solve the problem of extracting the target sound while maintaining the format of the retrieved sound information, the information processing device 100 restores the format of the query-based target sound extraction result 450 based on the retrieved sound information after step S3 (step S4). As a result, even if a decrease in the number of channels, frequency, or loss of localization occurs during the extraction of the target sound, the information processing device 100 can restore these formats, and therefore, as a whole, extract the target sound while maintaining the format of the retrieved sound information.

[0030] In step S4, for example, the information processing device 100 restores the sampling frequency, number of channels, direction, and position of the sound source of the target sound extracted from the audio data 401 as a format.

[0031] The key feature of this process is that, instead of directly restoring the format of the query-based target sound extraction result 450, it generates a reference for reference-based beamforming (BF) from the query-based target sound extraction result 450. Another feature of this process is that it performs target sound extraction again by applying beamforming using that reference to the retrieved sound information. Reference-based BF can handle any sampling frequency and number of channels as long as a reference is available.

[0032] Although the output signal is monaural, it is easy to adjust the phase of the output to match each channel of the input signal, making it possible to maintain the number of channels and restore the sense of localization.

[0033] Specifically, the information processing device 100 performs super-resolution processing as needed, which is a process that estimates frequency components higher than the sampling frequency of the query-based target sound extraction result 450. Subsequently, the information processing device 100 performs reference generation processing to generate a reference for use in the reference base BF.

[0034] A reference is an estimated value of the amplitude spectrogram of the target sound. Only one channel of reference needs to be generated. Furthermore, the reference can differ to some extent from the true amplitude of the target sound. Phase is also unnecessary. Therefore, generating the reference is easier than directly reconstructing the format of the query-based target sound extraction result 450.

[0035] However, the sampling frequency of the reference must match that of the audio data 401. Therefore, if the sampling frequency of the sound information that has undergone super-resolution processing differs from that of the audio data 401, the information processing device 100 performs processing to compensate for the difference between the two sampling frequencies.

[0036] For example, if the sampling frequency of the sound information that has undergone super-resolution processing is greater than or equal to the sampling frequency of the retrieved sound information, the information processing device 100 copies the frequency components corresponding to half or less of the sampling frequency of the retrieved sound information as a reference generation process.

[0037] Furthermore, if the sampling frequency of the sound information that has undergone super-resolution processing is less than the sampling frequency of the retrieved sound information, the information processing device 100 estimates, as a reference generation process, a frequency component that is greater than half the sampling frequency of the sound information that has undergone super-resolution processing and less than or equal to half the sampling frequency of the retrieved sound information, using the method described later.

[0038] As a result, the information processing device 100 restores the sampling frequency in the reference.

[0039] Next, the information processing device 100 performs beamforming using the reference and audio data 401, and then estimates the phase for each channel of the beamformed sound information. This restores the number of channels, direction and position of the sound source that have changed in the query-based target sound extraction result 450.

[0040] As described above, the information processing device 100 restores the format of the query-based target sound extraction result 450 based on the retrieved sound information. Therefore, even if a decrease in the number of channels, frequency, or loss of localization occurs during the target sound extraction process, these formats can be restored. For this reason, the information processing device 100 can generate a restored target sound extraction result 500 as the final search result, in which the target sound is extracted while maintaining the format of the retrieved sound information.

[0041] For example, when the information processing device 100 receives a query 200 in the form of text meaning the sound of a cow mooing, it generates a restored target sound extraction result 500 from a waveform file of search results that includes cow mooing and bird calls, etc., with the bird calls removed and the formatting maintained.

[0042] In addition, when the information processing apparatus 100 receives an input of a query 200 of text that means music characterized by a bass, the information processing apparatus 100 generates a restored target sound extraction result 500 in which parts other than the bass are removed from the music in the search result including parts other than the bass and the format is maintained.

[0043] (1-2. Configuration of Information Processing Apparatus According to Embodiment) An example of the configuration of the information processing apparatus 100 will be described using FIG. 2. FIG. 2 is a block diagram showing an example of the configuration of the information processing apparatus according to the embodiment. The information processing apparatus 100 includes a communication unit 110, a storage unit 120, a reception unit 130, a presentation unit 140, and a control unit 150.

[0044] (Communication Unit) The communication unit 110 is realized by, for example, a network interface controller (Network Interface Controller) or a NIC (Network Interface Card). The communication unit 110 may be a USB (Universal Serial Bus) interface configured by a USB host controller, a USB port, etc. Further, the communication unit 110 may be a wired interface or a wireless interface. For example, the communication unit 110 may be a wireless communication interface using a wireless LAN method or a cellular communication method.

[0045] The communication unit 110 functions as a communication means or a transmission means of the information processing apparatus 100. For example, the communication unit 110 is connected to the network N by wire or wirelessly, and transmits and receives information to and from an external device such as a cloud server or another information processing terminal via the network N. The network N is realized by, for example, a wireless communication standard or method such as Bluetooth (registered trademark), the Internet, Wi-Fi (registered trademark), UWB (Ultra-Wide Band), LPWA (Low Power Wide Area), ELTRES (registered trademark).

[0046] For example, the communication unit 110 receives the audio database 300 from an external terminal device. Also, the communication unit 110 can receive a query from an external terminal device.

[0047] (Memory Unit) The memory unit 120 is realized by, for example, a semiconductor memory device such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. For example, the memory unit 120 stores various information such as an audio database 300.

[0048] (Reception Unit) The reception unit 130 is realized by, for example, a keyboard, a camera with a character recognition function, a microphone with a voice recognition function, etc. when the query 200 is text. Also, when the query 200 is an image, the reception unit 130 is realized by a keyboard that receives an input of a file in which an image is recorded or the file name of the file via a GUI (Graphical User Interface) such as a button on the screen, or a camera that directly acquires the image.

[0049] When the query 200 is audio, the reception unit 130 is realized by a keyboard that receives an input of a file in which audio data is recorded or the file name of the file via a GUI, or a microphone that directly acquires the audio. When the query 200 is in a modality other than the above-mentioned modalities, the reception unit 130 is realized by a keyboard that receives an input of a file in which data of the modality is recorded or the file name of the file via a GUI, or a sensor that directly acquires the data of the modality.

[0050] (Presentation Unit) The presentation unit 140 is realized by, for example, a display unit such as a touch panel or a display, or an audio output unit such as a speaker. For example, the presentation unit 140 displays a waveform file of the restored target sound extraction result 500 or plays the sound of the restored target sound extraction result 500.

[0051] (Control Unit) The control unit 150 is implemented, for example, by a CPU (Central Processing Unit) or MPU (Micro Processing Unit) executing a program stored inside the information processing device 100 (for example, the information processing program according to this disclosure) using RAM or the like as a working area. The control unit 150 is also a controller and may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).

[0052] The control unit 150 comprises a query input unit 151, a search unit 152, a query-based target sound extraction unit 153, and a format / localization restoration unit 154. Furthermore, as described above, some of the processing in the search unit 152 is performed in advance (before query input), and the results are saved as an index, thereby shortening the search time. In such an embodiment, the control unit 150 also includes an index creation unit 156. The control unit 150 will be described below with reference to Figure 3. Figure 3 is a diagram illustrating the control unit.

[0053] (Query Input Unit) The query input unit 151 receives input of a query 200, which is information for searching for sound information to be searched. For example, if the search target is audio data containing dog barking, the query input unit 151 receives the following data as query 200 from the user via the reception unit 130: (1) The text "dog barking" (2) An image of a barking dog (3) A video containing a scene of a dog barking (4) A waveform file containing the sound of a dog barking

[0054] Furthermore, TB-AR (text-based audio retrieval), a technology that accepts text input as query 200, includes, for example, the technology described in Reference 1 below.

[0055] (Reference 1) Huang Xie, Samuel Lipping, and Tuomas Virtanen, “Language-based audio retrieval task in dcase 2022 challenge,” In Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop (DCASE), 216-220. 2022

[0056] The search format has variations in each of the following categories (I) to (IV). Note that when there are two or more channels, as in (IV) below, and there are also phase differences or volume differences between channels, a sense of localization is created, which expresses the spread of sound and the direction of the sound source. Localization depends not only on the number of channels but also on the microphone placement during recording. (I) Duration (II) Sampling frequency (16kHz, 44.1kHz, 48kHz, 96kHz, etc.) (III) Number of channels (mono, stereo, surround, etc.) (IV) Microphone placement during recording (when there are two or more channels)

[0057] (Search Unit) The search unit 152 searches for sound information based on the received query 200. For example, the search unit 152 searches the audio database 300 for audio data 400 to be searched as sound information. Specifically, the search unit 152 searches the audio database 300 for audio data 400, from the audio data 401 that is most similar to the target sound corresponding to the query 200 to the Nth similar audio data 40N.

[0058] At least one of the audio data 401 to 40N may contain interference sounds.

[0059] For example, if audio data 401 is ambient sound recorded in a park, it may include not only the target sound, which is the sound information corresponding to query 200, such as a dog barking, but also interfering sounds such as the voices of people playing in the park or birdsong. As another example, if the query 200 is "playing a guitar" and a guitar performance with vocals is searched, audio data 401 may include not only the target sound, which is the sound information corresponding to query 200, such as a guitar performance, but also interfering sounds such as vocals.

[0060] The details of the search unit 152 will be explained below using Figure 4. Figure 4 is a diagram illustrating the search unit. In Figure 4, the search unit 152 comprises a format conversion unit 1521, an audio encoder 1522, a second encoder 1523, and a matching unit 1524.

[0061] Figure 4 illustrates the case where the audio data stored in the audio database 300 is a waveform file. Even when the query 200 is a modal other than the waveform file, such as text, images, or videos, the search unit 152 can compare the similarity between different modals by converting the data from different modals into vectors belonging to a common space. As a result, the search unit 152 searches the audio database 300, which is the search target, for the audio data 400, which is the search result.

[0062] (Format Conversion Unit) If the format of the waveform file in the audio database 300 is different from the format corresponding to the audio encoder 1522, the format conversion unit 1521 retrieves the waveform file from the audio database 300 and converts its format. For example, the format conversion unit 1521 converts the format of the waveform file retrieved from the audio database 300 to the format corresponding to the audio encoder 1522.

[0063] (Audio Encoder) The audio encoder 1522 converts each waveform file, which has been converted to a format compatible with the audio encoder 1522, into a feature vector. The audio encoder 1522 generates an index 350 that contains the feature vector.

[0064] The index 350 includes a feature vector buffer 351, which is a buffer containing the feature vector, and a feature-file correspondence table 352, which indicates which waveform file in the audio database 300 corresponds to the feature vector. The index 350 may be generated in advance. In that case, in Figure 2, the control unit 150 includes an index creation unit 156, and the index creation unit 156 includes a format conversion unit 1521 and an audio encoder 1522. The pre-created index 350 is then used by the search unit 152 during a search.

[0065] (Second Encoder) The second encoder 1523 converts query 200 into a feature vector by encoding it. Hereafter, the feature vector converted from query 200 will be referred to as query vector 250.

[0066] (Matching Unit) The matching unit 1524 searches the audio data 400 based on the similarity or distance between each feature vector in the feature vector buffer 351 and the query vector 250. For example, the matching unit 1524 calculates the similarity between each feature vector in the feature vector buffer 351 and the query vector 250. As an example, the matching unit 1524 calculates the cosine similarity shown in the following equation (1) as the similarity.

[0067]

[0068] In equation (1), e_k are the feature vectors in the feature vector buffer 351. In the text of this specification, the underscore "_" indicates that the following character is a subscript in a mathematical expression. q is the query vector 250. In mathematical expressions, vectors are represented as column vectors.

[0069] The matching unit 1524 may calculate the distance between each feature vector in the feature vector buffer 351 and the query vector 250 as the similarity measure, instead of using cosine similarity. In this case, the smaller the value, the more similar the two vectors are. As an example, the matching unit 1524 calculates the Euclidean distance shown in the following equation (2) as the distance.

[0070]

[0071] The matching unit 1524, once the similarity between all feature vectors in the feature vector buffer 351 and the query vector 250 has been calculated, selects the first to Nth feature vectors in descending order of similarity. For example, the matching unit 1524 sorts the feature vectors in descending order of cosine similarity or in ascending order of Euclidean distance. The matching unit 1524 selects the first to Nth feature vectors in descending order of similarity.

[0072] Next, the matching unit 1524 retrieves waveform files corresponding to each of the selected N feature vectors from the audio database 300 by referring to the feature vector-file correspondence table 352. The matching unit 1524 retrieves pairs of the selected N feature vectors and the corresponding waveform files as the search result audio data 400.

[0073] In this way, the matching unit 1524, for example, converts both the text query 200 and the audio data into vectors and then calculates the similarity of the vectors, so it can perform searches even if the waveform file does not have a tag that represents the contents of the waveform file.

[0074] In addition, the matching unit 1524 can also select feature vectors whose similarity is equal to or greater than a threshold, rather than selecting the first to Nth feature vectors in order of decreasing similarity.

[0075] Furthermore, the matching unit 1524 can also have the waveform file of the audio data 400 displayed on the presentation unit 140 once it has acquired the search result audio data 400. For example, the matching unit 1524 can display the waveform file of the audio data 400 on the presentation unit 140, which functions as a display unit, or play it back from the presentation unit 140, which functions as an audio output unit.

[0076] (Query-based target sound extraction unit) Returning to the explanation of Figure 3, the query-based target sound extraction unit 153 extracts the target sound, which is the sound information corresponding to the received query 200, from the retrieved sound information. For example, the query-based target sound extraction unit 153 receives the input of query 200 and then extracts the target sound, which is the sound information corresponding to query 200, to perform target sound extraction on the audio data 400. In a configuration that simply combines search and query-based target sound extraction, both processes require a query, so even if both queries are the same, the user would have to input it twice. In contrast, in this disclosure, the query is shared between the two processes, so the user only needs to input the query once before the search by the search unit 152.

[0077] Specifically, the query-based target sound extraction unit 153 selects one of the audio data 401 to audio data 40N as the target for target sound extraction based on the user input received by the reception unit 130.

[0078] The following describes the case where the query-based target sound extraction unit 153 selects the audio data 401 that is most similar to the target sound as the target for target sound extraction. In this case, the query-based target sound extraction unit 153 performs format conversions such as reducing the number of channels or frequency of the audio data 401 as necessary.

[0079] Next, the query-based target sound extraction unit 153 inputs the format-converted audio data 401 and the query 200 into a machine learning model for target sound extraction, thereby generating a query-based target sound extraction result 450 from the audio data 401.

[0080] The query-based target sound extraction unit 153 can extract the target sound even if the query 200 is data of various modals. For example, if the query 200 is text, the query-based target sound extraction unit 153 extracts the target sound based on the TQ-TSE (Text-queried target sound extraction) technology described in Non-Patent Documents 1 and 2 above.

[0081] For example, if the query-based target sound extraction unit 153 searches for a guitar playing sound with vocals in response to a query 200 with the text "playing a guitar," it retains the target sound, which is the sound information corresponding to the query 200, and removes the vocals, which are a distracting sound that does not correspond to the query. Similarly, if the query-based target sound extraction unit 153 searches for a guitar playing sound with vocals in response to a query 200 with the text "singing voice," it retains the target sound, which is the sound information corresponding to the query 200, and removes the guitar playing sound, which is a distracting sound.

[0082] As another example, if the query-based target sound extraction unit 153 searches for dog barkings that include the sounds of people playing in a park and birds chirping in response to the query 200 of the text "dog barking," it will remove the sounds of people playing in a park and birds chirping that are not dog barking. Similarly, if the query-based target sound extraction unit 153 searches for child voices that include the sounds of dogs and birds in response to the query 200 of the text "child voice," it will remove the sounds of dogs and birds that are not child voices.

[0083] Furthermore, if the query 200 is an image or audio, the query-based target sound extraction unit 153 extracts the target sound based on the techniques described in Non-Patent Documents 3 and 4 above.

[0084] The details of the query-based target sound extraction unit 153 will be explained below using Figure 5. Figure 5 is a diagram illustrating the query-based target sound extraction unit. In Figure 5, the query-based target sound extraction unit 153 comprises a format conversion unit 1531 and a target sound extraction model 1532, which is a machine learning model used for query-based target sound extraction.

[0085] Furthermore, since the input / output format of the target sound extraction model 1532 is fixed, the query-based target sound extraction unit 153 does not need to prepare a separate target sound extraction model 1532 for each format; it only needs to have one.

[0086] Figure 5 illustrates the case where the query-based target sound extraction unit 153 selects audio data 401 from among the N audio data 401 to audio data 40N that constitute the audio data 400 as the target for target sound extraction.

[0087] Furthermore, as mentioned above, the input and output formats of the target sound extraction model 1532 are fixed. For example, the number of channels is fixed at 1 (monaural). Also, the sampling frequency of the target sound extraction model 1532 is fixed to F_enh, which is the lowest value among the audio data 401 to 40N to be searched. Note that the sampling frequency may be a fixed value such as 16 kHz instead of the lowest value among these audio data.

[0088] The reason this format is fixed is that it depends on the data used to train the target sound extraction model 1532. Furthermore, there are many different formats for the search results, making it difficult to train the target sound extraction model 1532 for each format.

[0089] (Format conversion unit) If the format of the audio data 401 is different from the input format of the target sound extraction model 1532, the format conversion unit 1531 adjusts the format of the audio data 401 to match the input format of the target sound extraction model 1532.

[0090] For example, if the number of channels in the audio data 401 is C and the sampling frequency is F_ret, the format conversion unit 1531 converts the number of channels in the audio data 401 to 1 and the sampling frequency to F_enh.

[0091] C is a value of 1 or more representing the number of channels in the search results. F_ret is the sampling frequency of the search results. F_enh is the input / output sampling frequency of the target sound extraction model 1532. Note that F_enh is a fixed value, while C and F_ret can take arbitrary values ​​for each search result.

[0092] Specifically, the format conversion unit 1531 converts the number of channels in the audio data 401 to 1 by one of the following methods: - Selecting one channel from multiple channels - Performing a mixdown by combining the data of one or more selected channels

[0093] Furthermore, if the sampling frequency of the audio data 401 is higher than the sampling frequency of the target sound extraction model 1532 (F_ret > F_enh), the format conversion unit 1531 performs downsampling to match the sampling frequency to F_enh, resulting in the loss of high-frequency components in the search results.

[0094] On the other hand, if the sampling frequency of the audio data 401 is lower than the sampling frequency of the target sound extraction model 1532 (F_ret < F_enh), the format conversion unit 1531 performs upsampling to match the sampling frequency to F_enh.

[0095] (Target Sound Extraction Model) The target sound extraction model 1532 outputs a query-based target sound extraction result 450, which is sound information from which the target sound has been extracted, in response to the input of a query 200 and the search results whose format has been converted. For example, the target sound extraction model 1532 outputs a query-based target sound extraction result 450 in response to the input of a query 200 and audio data 401 which has been converted to have 1 channel and a sampling frequency of F_enh.

[0096] (Format and localization restoration unit) Returning to the explanation of Figure 3, the format and localization restoration unit 154 performs processing equivalent to restoring the format of the query-based target sound extraction result 450 based on the retrieved sound information.

[0097] For example, the format and localization restoration unit 154 restores the sampling frequency, number of channels, and direction and position of the sound source of the query-based target sound extraction result 450 as the format. As described above, in this embodiment, instead of directly performing format restoration on the query-based target sound extraction result 450, a signal in which the format has been restored is generated by a method using a reference-based BF.

[0098] Specifically, the format and localization restoration unit 154 performs a process to generate a reference for use in the reference base BF. To this end, it performs a super-resolution process on the query-based target sound extraction result 450, which estimates frequency components higher than the sampling frequency of the extraction result. The super-resolution process can estimate high-frequency components with respect to the target sound when high-frequency components are lost as a side effect of downsampling in the format conversion unit 1531.

[0099] However, it is not possible to restore other lost information (such as a decrease in the number of channels and the resulting loss of localization). Therefore, in this embodiment, the format and localization restoration unit 154 uses super-resolution processing to generate the data necessary for the reference generation process, and the restoration of format and localization is performed by the reference base BF and subsequent processing.

[0100] Next, the format and localization restoration unit 154 generates a reference based on the super-resolution processing result, which is the sound information after applying super-resolution processing. In reference generation, a process is performed to compensate for the difference between the sampling frequency of the super-resolution processing result and the sampling frequency of the retrieved sound information.

[0101] In the resulting reference, the sampling frequency is restored from F_enh (the sampling frequency of the target sound extraction result) to F_ret (the sampling frequency of audio data 401). Furthermore, frequency components between F_enh / 2 and F_ret / 2 are estimated.

[0102] Next, the format and localization restoration unit 154 performs beamforming (BF) using the reference and audio data 401, and then estimates the phase for each channel of the audio data 401 based on the BF application result, which is the output of the beamforming.

[0103] Since beamforming and phase estimation processes do not use training data, they can be applied to input signals with any sampling frequency and number of channels, provided the reference is compatible. Although the BF application result is monaural, it is easy to adjust the output phase to match each channel of the input signal.

[0104] Therefore, the format and localization restoration unit 154 can restore the number of channels in the audio data 401, as well as the direction and position of the sound source. In this case, the format and localization restoration unit 154 restores the number of channels from 1 to C.

[0105] In this way, the format and localization restoration unit 154 outputs a restored target sound extraction result 500 in which the target sound is extracted while maintaining the format of the searched sound information, in response to the input of the query-based target sound extraction result 450 and the audio data 401.

[0106] As a result, the format and localization restoration unit 154 can generate a restored target sound extraction result 500 in which the number of channels and sampling frequency are the same as the search result audio data 401, and the localization of the target sound is maintained. Furthermore, the format and localization restoration unit 154 can convert the audio data 401, in which the target sound is heard as being to the right and the interfering sound as being to the left, into a restored target sound extraction result 500 in which the target sound is heard as being to the right, while the interfering sound is removed.

[0107] The details of the format and localization restoration unit 154 will be explained below using Figure 6. Figure 6 is a diagram illustrating the format and localization restoration unit. In Figure 6, the format and localization restoration unit 154 comprises a super-resolution unit 1541, an STFT (Short-Time Fourier Transform) unit 1542, an STFT unit 1543, a reference generation unit 1544, a reference base BF unit 1545, a phase estimation unit 1546, and an ISTFT unit 1547.

[0108] Furthermore, since F_enh and F_sr, which are inputs and outputs of the super-resolution model used for super-resolution processing by the super-resolution unit 1541, are both fixed values, it is not required to prepare a super-resolution model for each sampling frequency. One super-resolution model is sufficient. Note that F_sr is a fixed value greater than F_enh.

[0109] (Super-resolution unit) The super-resolution unit 1541 performs super-resolution processing on the query-based target sound extraction result 450, which is the extracted target sound, and estimates the frequency components higher than the sampling frequency of the query-based target sound extraction result 450. If the sampling frequency of a signal is F, then the upper limit of the range in which the frequency components of that signal exist is F / 2. Therefore, as a super-resolution process, the super-resolution unit 1541 estimates the frequency components higher than the sampling frequency of the query-based target sound extraction result 450, which is F_enh / 2.

[0110] Specifically, the super-resolution unit 1541 performs super-resolution processing by inputting the query-based target sound extraction result 450 into a super-resolution model, which is a machine learning model used for super-resolution processing. This converts the sampling frequency of the query-based target sound extraction result 450 from F_enh to F_sr, and estimates the corresponding frequency components. As an example, the super-resolution unit 1541 applies the means described in the following references 2 and 3 as super-resolution processing.

[0111] (Reference 2) Han, Seungu and Lee, Junhyeok, “NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates,” arXiv preprint arXiv:2206.08545, 2022 (Reference 3) Liu, Haohe and Chen, Ke and Tian, ​​Qiao and Wang, Wenwu and Plumbley, Mark D, “AudioSR: Versatile Audio Super-resolution at Scale,” arXiv preprint arXiv:2309.07314, 2023

[0112] (STFT section) The STFT section 1542 converts the super-resolution processing result, which is sound information that has undergone super-resolution processing, from waveform data, which is time-domain data, to a spectrogram (complex spectrogram), which is time-frequency domain data.

[0113] For example, the STFT unit 1542 converts waveform data into a complex spectrogram by applying a short-time Fourier transform to the waveform data. Note that when the sampling frequency of time-frequency domain data is F, it means that the data consists of frequency components from 0 Hz to F / 2.

[0114] Furthermore, the STFT unit 1543 converts the audio data 401, which is one of the search results, from waveform data to a complex spectrogram, similar to the STFT unit 1542. However, unlike the STFT unit 1542, the number of channels in the STFT result in the STFT unit 1543 is C, the same as the audio data 401.

[0115] (Reference Generation Unit) The reference generation unit 1544 generates a reference used in the reference base BF. The sampling frequency of the reference must be the same as that of the audio data 401, and since this value may differ from the sampling frequency of the super-resolution processing, a process is performed to compensate for the difference between the two frequencies.

[0116] The reference generation unit 1544 converts the super-resolution processing result, which has been converted into a complex spectrogram by the STFT unit 1542, into an amplitude spectrogram. Let the amplitude spectrogram converted from the super-resolution processing result be |Y_sr|.

[0117] Similarly, one channel of the audio data 401, which has been converted into a complex spectrogram by the STFT unit 1543, is selected and converted into an amplitude spectrogram. The amplitude spectrogram converted from the selected channel of the audio data 401 is denoted as |X_k|, where k is the number of the selected channel.

[0118] Next, the reference generation unit 1544 performs a process to match F_sr, which is the sampling frequency of the super-resolution processing result converted into an amplitude spectrogram, with F_ret, which is the sampling frequency of the audio data 401 converted into an amplitude spectrogram.

[0119] The following describes an example of reference processing by the reference generation unit 1544. First, the reference generation unit 1544 compares the relative magnitudes of F_sr, which is the sampling frequency of the super-resolution processing result converted into an amplitude spectrogram, and F_ret, which is the sampling frequency of the audio data 401.

[0120] Using Figure 7, we will explain an example of reference processing when the sampling frequency of the super-resolution processing result is greater than or equal to the sampling frequency of the audio data 401 (F_sr ≥ F_ret). Figure 7 is Figure (1) for explaining an example of reference processing.

[0121] The left side of Figure 7 shows the amplitude spectrogram 600 (|Y_sr|) of the super-resolution processing result, where the vertical axis represents frequency and the horizontal axis represents time. The component at frequency f and time t is represented as |y_sr(f,t)|. On the other hand, the right side shows the reference 700, where the vertical and horizontal axes are the same as in the amplitude spectrogram. The entire reference is represented as R, and the component at frequency f and time t is represented as r(f,t).

[0122] In this case, the reference generation unit 1544 generates a reference by copying the frequency components of the amplitude spectrogram 600 that correspond to frequencies between 0 and F_ret / 2, which correspond to less than or equal to half the sampling frequency of the reference 700.

[0123] Next, using Figure 8, we will explain an example of reference processing when the sampling frequency of the super-resolution processing result is less than the sampling frequency of the audio data 401 (F_sr < F_ret). Figure 8 is Figure (2) for illustrating an example of reference processing.

[0124] In this case, the reference generation unit 1544 copies all of the frequency components of the amplitude spectrogram 600 as part of the reference generation process. The reference generation unit 1544 also estimates the frequency components that are greater than half the sampling frequency of the super-resolution processing result (F_sr / 2) and less than or equal to half the sampling frequency of the audio data 401 (F_ret / 2) using the following method.

[0125] For example, when F_sr = 48 kHz and F_ret = 96 kHz, the reference generation unit 1544 estimates high-frequency components corresponding to frequencies greater than 24 kHz and less than or equal to 48 kHz. Specifically, the reference generation unit 1544 estimates high-frequency components using the following equations (3) to (5) and uses them as a reference at that frequency.

[0126]

[0127] In equations (3) to (5), t is the time index. f is the frequency index. r(f, t) is the reference. y_sr(f, t) is the super-resolution result after STFT. F_min is a value that satisfies 0 < F_min < F_sr. <F_min / 2 ≤ f ≤ F_sr / 2 means that the average value is calculated within the specified frequency range.

[0128] The reference generation unit 1544 calculates r(f,t) from frequency components near F_sr / 2 by using one of equations (3) and (4) and equation (5). Equations (3) and (4) represent the property that the time-direction pattern of the high-frequency amplitude spectrogram does not change significantly with frequency and is almost identical. Equation (5) is an equation that applies the tendency that the amplitude decreases as the frequency increases to r_est(t). In equation (5), |x_k(f,t)| is the frequency f and time t element in the aforementioned |X_k|, and X_k is the STFT search result corresponding to the k-th channel (microphone).

[0129] Since x_k(f,t) is thought to have a tendency for its amplitude to decrease as the frequency increases, it is thought that r(f,t) will have a similar tendency if multiplied by the power ratio of x_k(f,t) and r_est(t).

[0130] (Reference-based BF section) The reference-based BF section 1545 applies beamforming to the search results after STFT application, based on a reference generated by the process to set the sampling frequency to be the same as that of the search results, as described above. For example, the reference-based BF section 1545 takes the reference generated by the reference generation section 1544 and the complex spectrogram derived from the audio data 401, which is the search result, as input, performs beamforming on the latter, and outputs a signal corresponding to the former.

[0131] For example, the reference-based beamforming unit 1545 performs RB-TSE (Reference-Based Target Sound Extraction) as beamforming. That is, the reference-based beamforming unit 1545 uses a roughly estimated target sound amplitude spectrogram as a reference to extract the sound corresponding to the reference from a multi-channel signal in which multiple sounds are mixed.

[0132] Specifically, the reference-based BF unit 1545 applies a reference-based BF, called SIBF (Similarity-and-Independence-aware Beam Former), to the complex spectrogram derived from the audio data 401, which is the search result, based on a reference generated from the query-based target sound extraction result 450 on which the above-described transformation has been performed.

[0133] More specifically, the reference base BF section 1545 applies SIBF as described in the above-mentioned Patent Documents 2 and 3, and the following Reference Documents 4 and 5, as beamforming.

[0134] (Reference 4) Atsuo Hiroe, “Similarity-and-Independence-Aware Beamformer With Iterative Casting and Boost Start for Target Source Extraction Using Reference,” IEEE Open Journal of Signal Processing, vol. 3, pp. 1-20, 2022 (Reference 5) Atsuo Hiroe, “Online Similarity-and-Independence-Aware Beamformer for Low-latency Target Sound Extraction,” arXiv preprint arXiv: 2312.16449, 2023

[0135] SIBF is characterized by using a rough estimate of the amplitude spectrogram of the target sound as a reference, and then extracting sound information with higher accuracy (less distortion and interference) than that estimate. Furthermore, because SIBF is a linear process, it is less likely to generate sound distortion compared to nonlinear processes such as DNN (Deep Neural Network). In addition, if a reference is available, SIBF itself can handle any sampling frequency and number of channels. However, since it is beamforming, monaural input is not supported.

[0136] Furthermore, the sampling frequency of the BF application result, which is the output from the reference base BF unit 1545, is F_ret, the same as the audio data 401. On the other hand, the number of channels in this output is 1 (monaural), unlike the audio data 401 where C is the number of channels. In other words, although the sampling frequency is restored in the BF application result, the number of channels and the sense of localization are not restored.

[0137] The following describes an example of beamforming when the reference base BF unit 1545 uses SIBF. In this case, the reference base BF unit 1545 generates the BF application result y_bf(f,t) by using the following equations (6) to (9).

[0138]

[0139] < and > in equations (6) and (7) t This represents the average operation over all time points. In equation (7), max(a, b) represents the operation of selecting the larger argument. In equation (7), ε and β are small positive values ​​to avoid division by zero and positive values ​​to control the influence of the reference, respectively. Reference 9 states that ε = 10 -9 β = 1 / 2 is used.

[0140] In equations (6) and (7), x(f,t) is a quantity that represents the combined STFT results for all channels as a column vector. x(f,t) can be expressed, for example, as in equation (10) below. In equation (8), GEV_min(A,B) represents the operation of finding the eigenvector corresponding to the minimum eigenvalue λ_min in the generalized eigenvalue problem represented by equation (11) below. In equation (9), w(f) is a beamforming filter for extracting y_bf(f,t).

[0141]

[0142] (Phase Estimation Unit) The phase estimation unit 1546 takes the monaural BF application result and the search result, which is multi-channel audio data 401, as input and performs phase estimation for each channel of the search result (audio data 401) for the beamforming sound information.

[0143] The phase of a mixed sound is influenced by the phases of the sounds that make it up. Therefore, the phase of the extracted target sound is different from the phase of the mixed sound. The phase estimation unit 1546 then estimates the phase using the method described below.

[0144] For example, the phase estimation unit 1546 converts the beamforming filter used in SIBF into a steering vector while referring to the search result, such as the multi-channel audio data 401. The phase estimation unit 1546 also restores the number of channels in the BF application result from 1 to C, the same as the audio data 401, and restores the direction and position of the sound source based on the steering vector.

[0145] Furthermore, the phase estimation unit 1546 estimates the phase by projection processing, which projects the BF application results onto the projection space, based on the signals of each channel of the audio data 401, which is the search result.

[0146] As an example, the phase estimation unit 1546 performs phase estimation using the Minimal Distortion Principle (MDP), which is used as a post-processing step to adjust the amplitude and phase of the separation result in a sound source separation means that separates a sound mixed with multiple sounds into individual sound sources, as part of the projection processing. Specifically, the phase estimation unit 1546 adjusts both the amplitude and phase of the BF application result using the method described in Patent Document 3 and the following Reference 10 as the MDP. As a result, the sense of localization of the target sound is also restored.

[0147] (Reference 6) Kiyotoshi. Matsuoka, “Minimal distortion principle for blind source separation,” in Proceedings of the 41st SICE Annual Conference. SICE 2002, Soc. Instrument & Control Eng. (SICE), 2003

[0148] More specifically, the phase estimation unit 1546 estimates the phase by using the following equations (12) and (13). In equation (12), y_bf(f,t)* represents the complex conjugate of y_bf(f,t). In equation (13), z_k(f,t) is the restored target sound extraction result 500 corresponding to the k-th channel.

[0149]

[0150] The phase estimation unit 1546 calculates γ_k(f) from equation (12), which is the coefficient that minimizes the squared error between x_k(f,t), the search result after STFFT used as a reference signal for phase estimation, and y_bf(f,t), the output of the reference base BF unit 1545.

[0151] Furthermore, the phase estimation unit 1546 obtains z_k(f,t), which is the product of γ_k(f), the coefficient that minimizes the squared error, and y_bf(f,t), the output of the reference base BF unit 1545, from equation (13). The phase estimation unit 1546 generates the time-frequency domain restored target sound extraction result 500 by performing calculations corresponding to equations (12) and (13) for each channel where 1 ≤ k ≤ C.

[0152] Furthermore, the phase estimation unit 1546 can also perform phase estimation using the SWF (Single-Channel Wiener Filter) described in the above-mentioned reference 9, instead of the MDP, as a projection process. In this case, equation (14) below is used instead of equation (12).

[0153]

[0154] In SWF, the phase estimation unit 1546 generates a phase estimation reference signal by multiplying the reference r(f,t) by x_k(f,t) / |x_k(f,t)|, which is the phase component of the search result after STFT. The phase estimation unit 1546 then finds γ_k(f), which is the coefficient that minimizes the squared error between the generated reference signal and y_bf(f,t), which is the output of the reference base BF unit 1545.

[0155] The phase estimation unit 1546 calculates z_k(f,t), which is the product of γ_k(f) and y_bf(f,t), as shown in equation (13). The phase estimation unit 1546 generates the time-frequency domain restored target sound extraction result 500 by performing calculations corresponding to equations (13) and (14) for each channel where 1 ≤ k ≤ C.

[0156] Furthermore, if the search result is monaural, the phase estimation unit 1546 needs to estimate the phase using a method that does not depend on the BF. For example, it uses the phase of the search result as is, as shown in equation (15) below.

[0157]

[0158] (ISTFT section) The ISTFT section 1547 applies an I (Inverse) STFT, which is the inverse transform of STFT, to the output of the phase estimation section 1546. As a result, the ISTFT section 1547 generates a restored target sound extraction result 500, which is a time-domain signal. For example, the ISTFT section 1547 generates a restored target sound extraction result 500 in which the target sound is extracted while maintaining the sampling frequency, number of channels, localization, etc., of the search result audio data 401.

[0159] (Subsequent Processing Unit) Returning to the explanation of Figure 2, the subsequent processing unit 155 performs subsequent processing on the restored target sound extraction result 500. For example, the subsequent processing unit 155 saves the restored target sound extraction result 500 as a file to a storage device separate from the audio database 300 (for example, a storage device such as an HDD (Hard Disk Drive)). The subsequent processing unit 155 causes the presentation unit 140 to play (output) the restored target sound extraction result 500 as sound. The subsequent processing unit 155 also performs various editing operations on the restored target sound extraction result 500 using an application for editing audio waveforms.

[0160] (1-3. Flow of Information Processing According to the Embodiment) (First Information Processing Example) A first information processing example by the information processing device 100 according to the embodiment will be explained using Figure 9. Figure 9 is a flowchart (1) showing an example of the flow of information processing according to the embodiment.

[0161] The query input unit 151 receives the query 200 (step S11). For example, if the search target is audio data containing dog barking sounds, the query input unit 151 receives the query 200 from the user via the reception unit 130, including text, images, audio, and other queries.

[0162] Next, the search unit 152 searches for sound information based on the received query 200 (step S12). For example, the search unit 152 searches the audio database 300 for audio data 400, which is the search result for N Best, as sound information.

[0163] Next, the reception unit 130 receives a selection from the user via a GUI or the like, for one of the NBEST search results (step S13).

[0164] Next, the query-based target sound extraction unit 153 extracts the target sound, which is the sound information corresponding to the received query 200, from the retrieved sound information (step S14).

[0165] For example, if the format of the audio data 401 differs from the input format of the target sound extraction model 1532, the query-based target sound extraction unit 153 adjusts the format of the audio data 401 to match the input format of the target sound extraction model 1532. Then, the query-based target sound extraction unit 153 performs the target sound extraction process.

[0166] Next, the format and localization restoration unit 154 restores the format of the query-based target sound extraction result 450 based on the retrieved sound information (step S15). For example, the format and localization restoration unit 154 restores the sampling frequency, number of channels, and localization, which are the format, of the query-based target sound extraction result 450.

[0167] The subsequent processing unit 155 performs subsequent processing on the restored target sound extraction result 500 (step S16). For example, the subsequent processing unit 155 saves the restored target sound extraction result 500 as a file to a storage device. The subsequent processing unit 155 plays the restored target sound extraction result 500 as sound on the presentation unit 140. The subsequent processing unit 155 also performs various editing operations on the restored target sound extraction result 500 using an application for editing audio waveforms.

[0168] (Second Information Processing Example) A second information processing example by the information processing device 100 will be explained using Figure 10. Figure 10 is a flowchart (2) showing an example of the information processing flow according to the embodiment. Figure 10 is also a diagram for explaining in detail the search process of step S12 in Figure 9.

[0169] The second encoder 1523 converts the query 200 into a feature vector called the query vector 250 (step S21).

[0170] Next, the matching unit 1524 calculates the similarity between one of the feature vectors and the query vector (step S22). If the index creation unit 156 is included in Figure 2, or if the search process is the second or subsequent time, the feature vector buffer 351 is pre-constructed, so one feature vector is taken from the buffer and the similarity between it and the query vector 250 is calculated.

[0171] On the other hand, if the feature vector buffer 351 has not been constructed, one audio data file, i.e., a waveform file, is taken from the audio database 300, converted into a feature vector corresponding to the waveform by the format conversion unit 1521 and the audio encoder 1522, and then the similarity with the query vector 250 is calculated. The matching unit 1524 then saves the feature vector to the feature vector buffer 351 and adds information indicating which waveform file in the audio database 300 the feature vector corresponds to to the feature vector to the feature-file correspondence table 352. In other words, even if the index creation unit 156 is not provided, the index 350 is constructed when the search process is executed once.

[0172] Next, the matching unit 1524 determines whether or not similarity has been calculated for all audio data in the audio database 300 (step S23). If the index 350 has already been constructed, this determination is the same as determining whether or not similarity has been calculated for all feature vectors contained in the feature vector buffer 351. If there are waveform files or feature vectors for which similarity has not been calculated (step S23; No), the matching unit 1524 returns to step S22.

[0173] If the matching unit 1524 determines that the similarity calculation is complete (step S23; Yes), it sorts the feature vectors in descending order of similarity (step S24). The matching unit 1524 selects the first to Nth feature vectors in descending order of similarity. Then, for each of the N feature vectors, it obtains the corresponding waveform file name from the feature-file correspondence table 352, and then obtains the sound information (waveform data) corresponding to that file name from the audio database 300.

[0174] In Figure 10, the similarity between all feature vectors in the feature vector buffer 351 and the query vector 250 is calculated by forming a loop in steps S22 and S23. The matching unit 1524 can also be replaced with a mechanism that processes multiple waveforms and feature vectors simultaneously, such as parallel processing or matrix operations, instead of a loop.

[0175] (Third Information Processing Example) A third information processing example using the information processing device 100 according to the embodiment will be explained using Figure 11. Figure 11 is a flowchart (3) showing an example of the information processing flow according to the embodiment. Figure 11 is also a diagram for explaining in detail the format and localization restoration processing of step S15 in Figure 9.

[0176] The super-resolution unit 1541 performs super-resolution processing on the query-based target sound extraction result 450, which is the extracted target sound (step S31). If the number of channels in the audio data 401 is 1 and F_sr ≥ F_ret, the super-resolution unit 1541 can also consider the result of the super-resolution processing as the restored target sound extraction result 500. In that case, the super-resolution unit 1541 can omit steps S32 to S37.

[0177] The STFT unit 1542 converts the sound information, which has undergone super-resolution processing on the query-based target sound extraction result 450, from waveform data, which is time-domain data, to a complex spectrogram, which is time-frequency domain data (step S32).

[0178] The reference generation unit 1544 performs a reference generation process (step S33). For example, as a prerequisite for beamforming by the reference base BF unit, the reference generation unit 1544 generates a reference in which the sampling frequency of the sound information after the reference generation process is matched to the sampling frequency of the audio data 401.

[0179] If the reference-based beamforming unit 1545 determines that the number of channels in the search result audio data 401 is greater than 1 (step S34; Yes), it performs beamforming based on the reference (step S35). For example, the reference-based beamforming unit 1545 uses the query-based target sound extraction result 450, which has undergone the above-described transformation, as a reference, performs beamforming on the audio data 401 after STFT application, and generates a beamforming application result.

[0180] The phase estimation unit 1546 performs phase estimation for each channel of the search results based on the BF application result (step S36). For example, the phase estimation unit 1546 performs phase estimation by projection processing, which projects the BF application result onto the projection space, based on the signal of each channel of the audio data 401, which is the search result.

[0181] If the reference base BF unit 1545 determines that the number of channels in the search result audio data 401 is 1 (monaural) (step S34; No), beamforming cannot be applied, so instead, phase estimation for a monaural signal is performed (step S38).

[0182] The ISTFT unit 1547 generates a restored target sound extraction result 500, which is a time-domain signal, by applying ISTFT to the output of the phase estimation unit 1546 (step S37).

[0183] (Fourth Information Processing Example) A fourth information processing example using the information processing device 100 according to the embodiment will be explained using Figure 12. Figure 12 is a flowchart (4) showing an example of the flow of information processing according to the embodiment. Figure 12 is also a diagram for explaining in detail the reference generation process of step S33 in Figure 11.

[0184] The reference generation unit 1544 secures an area for storing the reference (step S41).

[0185] The reference generation unit 1544 generates a reference (step S43) by copying the frequency components of |Y_sr| corresponding to 0 ≤ f ≤ F_ret / 2 if F_sr ≥ F_ret (step S42; Yes). Note that the notation (f, *) in step S43 represents the total time at a specific frequency f.

[0186] The reference generation unit 1544 copies all of the frequency components of |Y_sr| if F_sr < F_ret (step S42; No) (step S44). The reference generation unit 1544 also estimates the frequency components that are greater than F_sr / 2 and less than or equal to F_ret / 2 (step S45).

[0187] (2. Modification) (2-1. First Modification) The information processing device 100 can also identify a time range corresponding to the received query 200 as sound information to be searched.

[0188] The time range refers to information about approximately how many seconds from the beginning of the audio data in the search results the section corresponding to the query is located in. For example, if a query like "dog barking" returns a file about 10 minutes long, and the dog only barks a few times within that file, if the time from the beginning of the file is identified, the user can easily play the audio corresponding to the query immediately or use it in other applications.

[0189] The following describes a modified example in which a time range is determined by performing detailed matching after the search process described in the embodiment. Figure 13 is a block diagram showing an example of the configuration of an information processing device according to the first modified example. The information processing device 100A includes a control unit 150A instead of a control unit 150. The control unit 150A includes a search unit 152A instead of a search unit 152. Also, as in Figure 2, index creation may be performed in advance, and in such an embodiment, the control unit 150A also includes an index creation unit 156A.

[0190] (Search Unit) In addition to the processing performed by the search unit 152, the search unit 152A identifies a time range for the audio data obtained as a search result by the search unit 152. To do this, the audio data is divided into partial waveforms and converted into feature vectors for each partial waveform. The waveform division is explained below in Figure 14, and the conversion to feature vectors is explained in Figure 15. Figure 14 shows an example of waveform division.

[0191] In Figure 14, waveform file 401A is a waveform approximately 2 minutes and 35 seconds long, and will be treated as an example of a long waveform hereafter. The basic idea for identifying the time waveform corresponding to the query for this waveform is to divide the long waveform into predetermined lengths to generate multiple sub-waveforms, and convert each of these into a different feature vector. However, considering the possibility that a continuous sound may be divided, the sub-waveforms will overlap in time.

[0192] In the example in Figure 14, a waveform of approximately 2 minutes and 35 seconds is divided into 10-second sub-waveforms, but each sub-waveform has a 5-second overlap. Specifically, the first 10 seconds are divided into the first sub-waveform 411, and the period from 5 seconds to 15 seconds is divided into the second sub-waveform 412. Similarly, the period from 2 minutes and 20 seconds to 2 minutes and 30 seconds is divided into the 29th sub-waveform 413, and the period from 2 minutes and 25 seconds to the end is divided into the 30th sub-waveform 414.

[0193] Next, we will explain how multiple partial waveforms are transformed using Figure 15. Figure 15 is a diagram illustrating an example of the transformation of multiple partial waveforms. In Figure 15, the search unit 152A includes an audio encoder 1522A instead of an audio encoder 1522. The search unit 152A also includes a vector integration unit 1525. Furthermore, the search unit 152A includes a matching unit 1524A (not shown in Figure 15) instead of a matching unit 1524. The matching unit 1524A will be explained in Figures 16 and 17.

[0194] (Audio Encoder) The audio encoder 1522A converts the first partial waveform 411 to the 30th partial waveform 414 into the first feature vector 421 to the 30th feature vector 424. In this way, unlike the audio encoder 1522, the audio encoder 1522A generates multiple feature vectors from a single waveform file depending on the length of the waveform file.

[0195] The audio encoder 1522A may perform the division into partial waveforms and conversion into feature vectors during the search, but it may also perform these operations in advance and save the results as index 350. In the latter embodiment, the search unit 152A is further equipped with an index creation unit 156A. The index creation unit 156A is equipped with the audio encoder 1522A and a vector integration unit 1525.

[0196] (Vector Integration Unit) The vector integration unit 1525 integrates the first feature vectors 421 to the 30th feature vectors 424, which are multiple feature vectors originating from the same waveform file 401A, into a single representative feature vector 430. The purpose of using the representative feature vector 430 will be described later. The vector integration unit 1525 generates the representative feature vector 430 by averaging the first feature vectors 421 to the 30th feature vectors 424. At that time, a process to normalize the norm of the feature vectors to 1 may be added before or after the averaging operation, or to one or both.

[0197] As mentioned above, the vector integration unit 1525 can perform the division into partial waveforms and conversion into feature vectors either during the search or beforehand. The former embodiment will be explained in the search processing example described later, and the latter will be explained below.

[0198] (Index Creation Unit) Returning to the explanation of Figure 13, the index creation unit 156A divides one of the audio data, i.e., waveform files, contained in the audio database 300 into partial waveforms of a predetermined length, as shown in Figure 14.

[0199] Next, the audio encoder 1522A is used to convert each partial waveform into a corresponding feature vector. Then, the index creation unit 156A stores the feature vector in the feature vector buffer 351 and adds time range information, indicating which waveform file the feature vector corresponds to and from which second, to which second, to the feature-file correspondence table.

[0200] Furthermore, the index creation unit 156A uses the vector integration unit 1525 to generate a representative feature vector corresponding to the waveform file. Then, it stores this representative feature vector in the feature vector buffer 351 and adds information indicating which waveform file the representative feature vector corresponds to to the feature-file correspondence table.

[0201] The index creation unit 156A performs the above processing for all audio data included in the audio database 300, thereby generating an index 350 that corresponds to the specified time range.

[0202] (Example of search processing) An example of search processing related to the first modified example by the matching unit 1524A of the search unit 152A of the information processing device 100A will be explained using Figure 16. Figure 16 is a flowchart showing an example of the flow of search processing related to the first modified example.

[0203] Steps S51 to S54 may be the same as steps S21 to S1 in Figure 10. However, in step S52, "calculating similarity," a representative feature vector may be used as the feature vector corresponding to each waveform file. When the "sorting" in step S54 is completed, the search results from 1st to Nth place, along with their respective file names and audio data (waveform data), have been obtained, but the time range corresponding to the query in each audio data has not yet been determined. After step S54, the matching unit 1524A performs "detailed matching" (step S55).

[0204] Next, we will explain the "detailed matching" in step S55 using Figure 17. Figure 17 is a flowchart showing an example of the detailed matching process.

[0205] The matching unit 1524A obtains a feature vector corresponding to one of the search results (step S61).

[0206] If the index 350 has already been constructed, the matching unit 1524A obtains a feature vector for each partial waveform corresponding to the file name from the feature vector buffer 351. If the index 350 has not yet been constructed, the matching unit 1524A converts the audio data into feature vectors for each partial waveform using the method shown in Figures 14 and 15.

[0207] The matching unit 1524A then stores the feature vector in the feature vector buffer 351 and adds time range information indicating which waveform file the feature vector corresponds to to which second, and to which second, into the feature-file correspondence table.

[0208] Furthermore, the matching unit 1524A uses the vector integration unit 1525 to generate a representative feature vector corresponding to the waveform file. Then, it stores this representative feature vector in the feature vector buffer 351 and adds information indicating which waveform file the representative feature vector corresponds to to the feature-file correspondence table.

[0209] The matching unit 1524A calculates the similarity (step S62). For example, if one of the search results is waveform 401 file A shown in Figure 14, the matching unit 1524A calculates the similarity between one of the first feature vectors 421 to the 30th feature vector 424 and the query vector 250.

[0210] The matching unit 1524A determines whether the similarity calculation is complete (step S63). For example, the matching unit 1524A determines that the similarity calculation is complete when it has calculated the similarity between all of the first feature vectors 421 to the 30th feature vectors 424 and the query vector 250.

[0211] If the matching unit 1524A determines that the similarity calculation is not yet complete (step S63; No), it returns to step S62. If the matching unit 1524A determines that the similarity calculation is complete (step S63; Yes), it rearranges the feature vectors in descending order of similarity to the query vector 250 (step S64). For example, the matching unit 1524A retains the top M feature vectors with the highest similarity to the query vector 250. M is a natural number set independently of N.

[0212] Since the feature vectors and partial waveforms correspond in a one-to-one manner, the matching unit 1524A can find the portion corresponding to the query 200 within a single waveform file 401A by processing in step S64.

[0213] The matching unit 1524A determines whether the processing of the N best search results, which are the top N audio data with the greatest similarity to the target sound corresponding to the query 200, has been completed (step S65). For example, the matching unit 1524A determines whether the processing in steps S61 to S64 has been completed for each of the audio data from 401 to 40N.

[0214] If the matching unit 1524A determines that it has finished processing the search results for NBEST (step S65; Yes), it terminates the process. If the matching unit 1524A determines that it has not finished processing the search results for NBEST (step S65; No), it returns to step S62.

[0215] (2-2. Second Modification) The first modification concerned the process of identifying the time range corresponding to the query for each search result. In contrast, the second modification concerns how to present the time range obtained in this way to the user.

[0216] The information processing device 100 can also display the time range identified as sound information corresponding to the received query 200. The configuration of the information processing device 100B according to the second modification will be described below with reference to Figure 18. Figure 18 is a block diagram showing an example of the configuration of the information processing device according to the second modification. The information processing device 100B includes a control unit 150B instead of a control unit 150. The control unit 150B includes a search unit 152B instead of a search unit 152.

[0217] (Search Unit) The search unit 152B identifies the time range corresponding to the received query 200 and displays it superimposed on the display of the entire waveform. For example, the search unit 152B displays the search results on the display unit 140 when detailed matching is performed, which is the process in steps S51 to S55. An example of the display of search results when detailed matching is performed by the search unit 152A will be explained using Figure 19. Figure 19 is a diagram showing an example of search results when detailed matching is performed.

[0218] In Figure 19, the search unit 152B displays a screen 900 on the presentation unit 140 that includes a waveform file 401B with a long recording period as one of the search results for audio data. If the audio data consists of N waveform files, the search unit 152B can also display a total of N screens similar to screen 900, which can be switched between. Furthermore, the search unit 152B can also allow the user to select an arbitrary time range for the waveform file 401B on screen 900.

[0219] Furthermore, the search unit 152B displays the results of the detailed matching and the first section 411B to the sixth section 416B, which correspond to each partial waveform into which the waveform file 401B has been divided, on the display unit 140. In addition, the search unit 152B displays ranks 4111 to 4161, which represent the ranking of the degree of similarity between the query vector 250 and the feature vectors of each section, on the display unit 140.

[0220] For example, the search unit 152B displays the first interval 411B, from 30 seconds to 40 seconds, on the display unit 140 as the interval most similar to the query vector 250. The search unit 152B displays the second interval 412B, from 35 seconds to 45 seconds, on the display unit 140 as the interval next most similar to the query vector 250. The search unit 152B displays the third interval 413B, from 55 seconds to 65 seconds, on the display unit 140 as the interval next most similar to the query vector 250.

[0221] When the reception unit 130 receives a selection of a section via a click or other means on the screen 900, the search unit 152B can also display the selected section on the presentation unit 140 in a different display format from the other sections.

[0222] Furthermore, the search unit 152B displays GUI elements such as buttons 930 to 960 on the display unit 140. When button 930 is clicked, the search unit 152B plays the waveform file 401B from the audio output unit of the display unit 140, such as a speaker. Also, when a time range of the waveform file 401B, such as a section, is selected, the search unit 152B plays only the sound corresponding to that time range on the display unit 140.

[0223] Furthermore, when button 940 is clicked or otherwise, the search unit 152B stores the waveform data of the entire waveform file 401B or a selected time range in a storage device such as an HDD. Also, when button 950 is clicked or otherwise, the search unit 152B copies the same waveform data as the entire waveform file 401B or a selected time range to a mechanism for sharing data between applications, such as a clipboard.

[0224] Furthermore, when button 960 is clicked or otherwise, the search unit 152B applies both a query-based target sound extraction process, which extracts the target sound corresponding to the query, and a format and localization restoration process, which restores the format, to all or part of the waveform file 401B. Subsequently, when buttons 930 to 950 are clicked or otherwise, the search unit 152B plays, saves, and copies the results of both processes applied.

[0225] Furthermore, the search unit 152B may display buttons 930 to 950 on the display unit 140 as buttons for playback, saving, and copying the results of both processes, and omit button 960. In that case, the search unit 152B will perform one of the following three processes depending on when both processes are performed.

[0226] As a first example, the search unit 152B applies both processes to each of the NBEST search results. In this case, the search unit 152B omits step S13 in Figure 9. The search unit 152B also displays the results of applying both processes as a waveform file 401B on the display unit 140.

[0227] As a second example, the search unit 152B performs both processes when buttons 930 to 950 are clicked or otherwise affected. The search unit 152B also performs playback, saving, and copying of the results of applying both processes. If a time range for the waveform file 401B is selected, the search unit 152B applies both processes to that time range. If no time range is selected, the search unit 152B applies both processes to the entire waveform file 401B.

[0228] As a third example, the search unit 152B performs query-based target sound extraction processing for each search result of NBEST, and then performs format and localization restoration processing when buttons 930 to 950 are clicked or otherwise activated. In this case, the search unit 152B displays the result of the query-based target sound extraction as a waveform file 401B on the display unit 140.

[0229] Even if the format of the query-based target sound extraction process is converted and the result of the query-based target sound extraction process is displayed on the display unit 140, the final generated sound information is the restored target sound extraction result 500. Therefore, the format is maintained regardless of which of the three timings described above is adopted.

[0230] Furthermore, instead of displaying a screen 900 on the display unit 140 for each waveform data, the search unit 152B can also display the following screen 900 on the display unit 140.

[0231] For example, when the sorting in step S24 in Figure 10 is completed, the search unit 152B displays the path name of the waveform file corresponding to the N-best search result on the display unit 140. Also, when the user selects one of the waveform files corresponding to the N-best search result, the search unit 152B displays the screen corresponding to the selected waveform file on the display unit 140. The search unit 152B may apply both of the above processes to only the selected waveform file, or to each of the N-best search results.

[0232] The search unit 152B may display only the file path name or, as an intermediate between a simple display and a detailed display like the screen 900, it may display the time range of the intervals with particularly high similarity to the query 200 (for example, within the top 5) along with the file path name. For example, if the search unit 152B displays the top 5 time ranges with particularly high similarity to the query 200, the display unit 140 will further display the following strings on the screen 900: 1. 30 seconds to 40 seconds 2. 35 seconds to 45 seconds 3. 55 seconds to 1 minute 5 seconds 4. 1 minute 20 seconds to 1 minute 30 seconds 5. 50 seconds to 1 minute

[0233] (2-3. Third Modification) The above example described the case in which only one query 200 is used in a single search. On the other hand, the information processing device 100 can also accept multiple queries as queries 200 and search for sound information based on the multiple queries that have been accepted. The configuration of the information processing device 100C according to the third modification will be explained below using Figure 20. Figure 20 is a block diagram showing an example of the configuration of the information processing device according to the third modification.

[0234] The information processing device 100C includes a control unit 150C instead of a control unit 150. The control unit 150C includes a query input unit 151C and a search unit 152C instead of a query input unit 151 and a search unit 152.

[0235] (Query Input Section) The query input section 151C accepts multiple queries as queries 200. The query input section 151C may accept queries as queries 200 in which the modals of each query are the same, or it may accept queries in which the modals of each query are different.

[0236] Furthermore, the query input unit 151C may accept duplicate modal queries as query 200. For example, the query input unit 151C may accept two text queries and one audio data query as query 200 from among three queries.

[0237] (Search Unit) The search unit 152C searches for sound information based on multiple received queries. Below, Figure 21 will be used to explain a first example of searching for sound information based on multiple queries. Figure 21 is a diagram illustrating a first example of searching for sound information based on multiple queries.

[0238] In Figure 21, the search unit 152C performs a search after integrating the query vectors corresponding to each query into one. Also in Figure 21, the search unit 152C searches for sound information based on two queries. However, the search unit 152C can search for sound information based on any number of queries. Furthermore, in Figure 21, the search unit 152C is further equipped with a third encoder 1523C and a query integration unit 1526 in addition to the search unit 152.

[0239] (Third Encoder) The third encoder 1523C converts the second query 200C into the query vector 250C.

[0240] (Query Integration Unit) The query integration unit 1526 generates a single integrated vector 325 from the query vectors 250 and 250C, by integrating these multiple query vectors. For example, the query integration unit 1526 generates the integrated vector 325 by multiplying each of these multiple vectors by a weight coefficient 275 prepared for each modal or query, and then performing a weighted average.

[0241] The query integration unit 1526 obtains audio data 400 by inputting the integration vector 325 to the matching unit 1524 instead of the query vector 250.

[0242] Furthermore, if the search unit 152C is to support three or more queries, it only needs to be equipped with an encoder corresponding to the number of queries. However, if modals overlap between queries, the search unit 152C can share an encoder. For example, even if there are three queries, if two of them are text and the remaining one is audio, the search unit 152C can share an encoder (text encoder) for converting text queries.

[0243] Furthermore, the search unit 152C can also set the weight coefficient 275 in response to input received by the reception unit 130 via a GUI such as a slider bar. In this case, even if the query itself is not changed, the search results may change if the weight coefficient 275 changes. Therefore, when the search unit 152C detects a change in the weight coefficient 275, it may perform the search process without the user having to initiate a search, and cause the presentation unit 140 to redisplay the search results.

[0244] When multiple queries are received by the query input unit 151C, there are several possible embodiments for how these multiple queries are handled in the query-based target sound extraction process, as described below. In other words, in Figure 3, the single input query, query 200, was supplied to both the search unit 152 and the query-based target sound extraction unit 153, whereas in the third modified example, the search unit 152C uses all the input queries, while the query-based target sound extraction unit basically uses only one of the multiple queries, and there are several ways to select it.

[0245] As an example, when the reception unit 130 receives a selection of a query from the user for which query-based target sound extraction processing will be performed, it selects the query to be input to the query-based target sound extraction unit 153. As another example, the reception unit 130 selects the first query entered into the query input unit 151C as the query to be input to the query-based target sound extraction unit 153.

[0246] Furthermore, if the same encoder is used in the query-based target sound extraction unit 153 as in the search unit 152C, the search unit 152C can also use the integrated vector 325 as the encoder's output vector. In other words, a common query vector can be used in both the query-based target sound extraction unit 153 and the search unit 152C only when the same encoder is used. That is, in that case, the query-based target sound extraction unit 153 can also reflect multiple queries.

[0247] In the example above, all queries were used to retrieve N search results. In another embodiment, the search unit 152C can retrieve the N search results using the first query, and the second query can be used to change the ranking of the search results.

[0248] The following describes a second example of searching for sound information based on multiple queries, using Figure 22. Figure 22 is a diagram illustrating a second example of searching for sound information based on multiple queries. In Figure 22, the search unit 152C obtains the N-best search result, which is audio data 400, using one query 200, and then performs rematching to calculate the similarity between the N-best search result and the second query 200C. In Figure 22, the search unit 152C is further equipped with a third encoder 1523C and a rematching unit 1524C in addition to the search unit 152.

[0249] (Rematching Unit) The rematching unit 1524C calculates the similarity between the query vector 250C converted from the second query 200C by the third encoder 1523C and each search result of the N best that constitute the audio data 400 obtained based on the first query 200. The rematching unit 1524C sorts the N best based on the calculated similarity to obtain the audio data 400C, which is the search result of the rematched N best.

[0250] As a result, although the search results that make up NBEST themselves do not change, the rematching unit 1524C changes the order of each audio data that makes up audio data 400 in the rematching process to audio data 401C to 401N that make up audio data 400C.

[0251] Furthermore, the rematching unit 1524C may perform detailed matching using a second query after rematching, similar to the first modified example. In this case, the time range particularly similar to the query within a single search result also changes. Therefore, the user is more likely to easily discover the time range compared to searching for a time range similar to the query based on only one query.

[0252] (3. Other Embodiments) Of the processes described in each of the above embodiments, all or part of the processes described as being performed automatically may be performed manually, or all or part of the processes described as being performed manually may be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings may be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.

[0253] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads, usage conditions, etc.

[0254] Furthermore, the embodiments and modifications described above can be combined as appropriate, provided that the processing content is not inconsistent. Also, the steps shown in the sequence diagram or flowchart of this embodiment can be changed in order as appropriate. For example, each step may be processed chronologically, repeatedly, or partially in parallel. Moreover, the effects described herein are merely illustrative and not limiting, and other effects may exist.

[0255] (4. Effects of the Information Processing Method relating to the Disclosure) As described above, the information processing method relating to the Disclosure is an information processing method performed by a computer. The information processing method includes a reception step (step S11 in the embodiment), a search step (step S12 in the embodiment), a target sound extraction step (step S14 in the embodiment), and a restoration step (step S15 in the embodiment).

[0256] The reception step accepts a query, which is information for searching for sound information. The search step searches for sound information based on the accepted query. The target sound extraction step extracts the target sound, which is the sound information corresponding to the accepted query, from the searched sound information. The restoration step restores the format of the query-based target sound extraction result, which is the output of the target sound extraction step, based on the searched sound information.

[0257] As a result, depending on the information processing method, even if the number of channels, frequency, or sense of localization decreases during target sound extraction compared to the search results, these formats can be restored. Therefore, overall, the target sound can be extracted while maintaining the format of the retrieved sound information.

[0258] The restoration step restores the number of channels of the extracted target sound as a format.

[0259] This allows the number of channels in the extracted target sound to be restored according to the information processing method, thus better preserving the format of the retrieved sound information.

[0260] The restoration step restores the sampling frequency of the extracted target sound as a format.

[0261] This allows the sampling frequency of the extracted target sound to be restored according to the information processing method, thus better preserving the format of the retrieved sound information.

[0262] The restoration step involves performing super-resolution processing on the extracted target sound, which estimates frequency components higher than the sampling frequency of the extracted target sound.

[0263] This allows the information processing method to estimate higher frequency components than the sampling frequency of the extracted target sound, thereby bringing it closer to the sampling frequency of the search results. Therefore, the information processing method makes it possible to maintain the sampling frequency of the retrieved sound information or to perform beamforming while matching the sampling frequency of the extracted target sound with the relevant frequency.

[0264] The restoration step involves a reference generation process, which compensates for the difference between the sampling frequency of the super-resolution processed audio information and the sampling frequency of the retrieved audio information.

[0265] This allows the information processing method to compensate for the difference between the sampling frequency of the super-resolution processed sound information and the sampling frequency of the retrieved sound information, thereby maintaining the sampling frequency of the retrieved sound information. Furthermore, the information processing method enables beamforming while matching the sampling frequency of the super-resolution processed sound information with the sampling frequency of the retrieved sound information.

[0266] The restoration step, if the sampling frequency of the super-resolution processed audio information is greater than or equal to the sampling frequency of the retrieved audio information, copies the frequency components corresponding to half or less of the sampling frequency of the retrieved audio information as a reference generation process.

[0267] Therefore, according to the information processing method, if the sampling frequency of the sound information that has undergone super-resolution processing is greater than or equal to the sampling frequency of the retrieved sound information, the above-mentioned difference can be compensated for simply by copying the frequency components corresponding to half or less of that sampling frequency.

[0268] In the restoration step, if the sampling frequency of the super-resolution processed audio information is less than the sampling frequency of the retrieved audio information, the reference generation process estimates frequency components that are greater than half the sampling frequency of the super-resolution processed audio information and less than or equal to half the sampling frequency of the retrieved audio information.

[0269] Therefore, the reference used in beamforming such as SIBF can be a rough estimate of the amplitude spectrogram of the target sound, and according to the information processing method, the aforementioned difference can be compensated for by performing the estimation described above. Furthermore, the time-direction pattern in the high-frequency amplitude spectrogram does not change significantly with frequency and is almost identical, so according to the information processing method, the aforementioned difference can be compensated for.

[0270] The restoration step restores, as a format, at least one of the direction and location of the sound source of the extracted target sound.

[0271] This allows the localization of the extracted target sound to be restored according to the information processing method, thus better preserving the format of the retrieved sound information.

[0272] The restoration step involves beamforming.

[0273] Therefore, according to the information processing method, beamforming itself is unrelated to machine learning, and can be applied to search results with any sampling frequency and number of channels, thereby restoring the sense of localization lost during target sound extraction.

[0274] The restoration step involves generating a reference from the extracted target sound, which has been processed to have the same sampling frequency as the retrieved sound information. Based on this reference, beamforming is then performed on the retrieved sound information.

[0275] This allows the sampling frequency of the extracted target sound to be avoided depending on the information processing method. At the time of beamforming, the number of channels in the output beamformed sound information is monaural. However, the output sampling frequency maintains the same value as the searched sound information.

[0276] The restoration step involves estimating the phase of each channel of the retrieved audio information for the beamforming audio information.

[0277] As a result, the information processing method estimates the phase relative to the beamforming output, eliminating the need to train a machine learning model. Furthermore, since the information processing method estimates the phase for each channel, it can estimate the phase for any number of channels or microphone arrangement. Consequently, the number of channels and localization of the beamforming sound information are restored.

[0278] The restoration step involves estimating the phase by projection processing, which projects the beamformed sound information into the projection space.

[0279] As a result, according to the information processing method, the appropriate phase can be estimated for each channel, allowing for the proper restoration of the number of channels and localization.

[0280] The reconstruction step involves estimating the phase using MDP as a projection process.

[0281] As a result, according to the information processing method, a more appropriate phase can be estimated for each channel as part of the projection process, allowing for a more accurate restoration of the number of channels and localization.

[0282] The reconstruction step involves estimating the phase using SWF as a projection process.

[0283] As a result, according to the information processing method, a more appropriate phase can be estimated for each channel as part of the projection process, allowing for a more accurate restoration of the number of channels and localization.

[0284] The target sound extraction step accepts query inputs received before the search step.

[0285] In a configuration that simply links the search step and the target sound extraction step, the target sound extraction step is also required to accept query input from the user. In contrast, according to the information processing method, the query received in the search step is also accepted in the target sound extraction step, so query input from the user can be omitted in the target sound extraction step.

[0286] The search step identifies the time range corresponding to the received query as sound information.

[0287] As a result, depending on the information processing method, even if a query such as "dog barking" yields a waveform file that is 10 minutes or longer in duration, the system can automatically identify the parts in the waveform file where a dog barks.

[0288] The search step displays the specified time range.

[0289] As a result, depending on the information processing method, even if a long waveform file is searched, the time range in which the query corresponds to the relevant section within the waveform file is displayed, allowing the user to be efficiently presented with that section.

[0290] The reception step accepts multiple queries as input queries, and the search step searches for sound information based on the accepted queries.

[0291] According to the information processing method, even when there are two or more queries, sound information can be retrieved in the same way as when there is only one query, for example, by integrating feature vectors.

[0292] (5. Hardware Configuration) The information processing device 100 etc. according to the present disclosure described above is realized by a computer 1000 having a configuration such as that shown in Figure 23. Figure 23 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device. Hereinafter, as an example of the computer 1000, the information processing device 100 according to the embodiment will be described as an example. The computer 1000 has a CPU 1100, RAM 1200, ROM (ReadOnly Memory) 1300, HDD 1400, communication interface 1500, and input / output interface 1600. The parts of the computer 1000 are connected by a bus 1050.

[0293] The CPU 1100 operates based on programs stored in the ROM 1300 or HDD 1400 and controls each part. For example, the CPU 1100 loads the programs stored in the ROM 1300 or HDD 1400 into the RAM 1200 and executes processing corresponding to various programs.

[0294] ROM 1300 stores boot programs such as the BIOS (Basic Input Output System) that are executed by the CPU 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.

[0295] The HDD 1400 is a computer-readable recording medium that non-temporarily stores programs executed by the CPU 1100 and data used by said programs. Specifically, the HDD 1400 is a recording medium that stores an information processing program related to this disclosure, which is an example of program data 1450.

[0296] The communication interface 1500 is an interface for the computer 1000 to connect to an external network 1550 (e.g., the Internet). For example, the CPU 1100 can receive data from other devices or transmit data it has generated to other devices via the communication interface 1500.

[0297] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from input devices such as a keyboard or mouse via the input / output interface 1600. The CPU 1100 also transmits data to output devices such as a display, speaker, or printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, or semiconductor memory.

[0298] For example, when the computer 1000 functions as an information processing device 100 according to the embodiment, the CPU 1100 of the computer 1000 realizes functions such as the control unit 150 by executing an information processing program loaded on the RAM 1200. The HDD 1400 stores the information processing program according to this disclosure and data in the storage unit 120. The CPU 1100 reads and executes the program data 1450 from the HDD 1400, but as another example, these programs may be obtained from other devices via an external network 1550.

[0299] (6. Supplement) This technology can also be configured as follows: (1) An information processing method performed by a computer, comprising: a reception step of receiving input of a query which is information for searching for sound information; a search step of searching for sound information based on the received query; a target sound extraction step of extracting a target sound which is sound information corresponding to the received query from the searched sound information; and a restoration step of restoring the format of the query-based target sound extraction result which is the output of the target sound extraction step based on the searched sound information. (2) The information processing method according to (1) wherein the restoration step restores the number of channels of the extracted target sound as the format. (3) The information processing method according to (1) or (2) wherein the restoration step restores the sampling frequency of the extracted target sound as the format. (4) The information processing method according to (3) wherein the restoration step performs super-resolution processing on the extracted target sound, which is a process of estimating frequency components higher than the sampling frequency of the extracted target sound. (5) The information processing method according to (4), wherein the restoration step comprises performing a reference generation process which compensates for the difference between the sampling frequency of the sound information subjected to the super-resolution processing and the sampling frequency of the retrieved sound information. (6) The information processing method according to (5), wherein the restoration step comprises copying frequency components corresponding to half or less of the sampling frequency of the retrieved sound information as the reference generation process when the sampling frequency of the sound information subjected to the super-resolution processing is greater than or equal to the sampling frequency of the retrieved sound information. (7) The information processing method according to (5), wherein the restoration step comprises estimating frequency components corresponding to more than half of the sampling frequency of the sound information subjected to the super-resolution processing and half or less of the sampling frequency of the retrieved sound information as the reference generation process when the sampling frequency of the sound information subjected to the super-resolution processing is less than the sampling frequency of the retrieved sound information.(8) The information processing method according to (1) wherein the restoration step restores at least one of the direction and position of the sound source of the extracted target sound as the format. (9) The information processing method according to (8) wherein the restoration step performs beamforming. (10) The information processing method according to (9) wherein the restoration step generates a reference from the extracted target sound that has been processed to have the same sampling frequency as the searched sound information, and then performs beamforming on the searched sound information based on the generated reference. (11) The information processing method according to (10) wherein the restoration step performs phase estimation for each channel of the searched sound information on the beamformed sound information. (12) The information processing method according to (11) wherein the restoration step performs phase estimation by projection processing that projects the beamformed sound information into projection space. (13) The information processing method according to (12), wherein the restoration step is the projection process, in which the phase is estimated by MDP (Minimal Distortion Principle). (14) The information processing method according to (12), wherein the restoration step is the projection process, in which the phase is estimated by SWF (Single-Channel Wiener Filter). (15) The information processing method according to any one of (1) to (14), wherein the target sound extraction step is the input of a query received before the search step. (16) The information processing method according to any one of (1) to (15), wherein the search step is the time range corresponding to the received query as sound information. (17) The information processing method according to (16), wherein the search step is the time range specified is displayed. (18) The information processing method according to any one of (1) to (17), wherein the reception step receives a plurality of queries as the query, and the search step searches for the sound information based on the plurality of queries received.(19) An information processing system comprising: a receiving unit that receives input of a query which is information for searching for sound information; a searching unit that searches for sound information based on the received query; a target sound extraction unit that extracts a target sound which is sound information corresponding to the received query from the searched sound information; and a restoration unit that restores the format of the query-based target sound extraction result which is the output of the target sound extraction unit based on the searched sound information. (20) An information processing method performed by a computer, comprising: a receiving step of receiving input of a query which is information for searching for sound information; and a presentation step of presenting a target sound extraction result which has the same format as the sound information searched based on the received query and in which a target sound which is sound information corresponding to the received query has been extracted.

[0300] 100 Information Processing Unit 110 Communication Unit 120 Storage Unit 130 Reception Unit 140 Presentation Unit 150 Control Unit 151 Query Input Unit 152 Search Unit 153 Query-Based Target Sound Extraction Unit 154 Format / Localization Restoration Unit 155 Post-Processing Unit N Network

Claims

1. An information processing method performed by a computer, comprising: a reception step of receiving input which is a query that is information for searching for sound information; a search step of searching for sound information based on the received query; a target sound extraction step of extracting a target sound that is sound information corresponding to the received query from the searched sound information; and a restoration step of restoring the format of the query-based target sound extraction result, which is the output of the target sound extraction step, based on the searched sound information.

2. The information processing method according to claim 1, wherein the restoration step restores the number of channels of the extracted target sound as the format.

3. The information processing method according to claim 1, wherein the restoration step restores the sampling frequency of the extracted target sound as the format.

4. The information processing method according to claim 3, wherein the restoration step involves performing a super-resolution process on the extracted target sound, which is a process of estimating frequency components higher than the sampling frequency of the extracted target sound.

5. The information processing method according to claim 4, wherein the restoration step includes a reference generation process which compensates for the difference between the sampling frequency of the sound information subjected to the super-resolution processing and the sampling frequency of the retrieved sound information.

6. The information processing method according to claim 5, wherein, if the sampling frequency of the sound information on which the super-resolution processing has been performed is equal to or greater than the sampling frequency of the retrieved sound information, the restoration step is to copy a frequency component corresponding to half or less of the sampling frequency of the retrieved sound information as the reference generation process.

7. The information processing method according to claim 5, wherein, if the sampling frequency of the sound information subjected to the super-resolution processing is less than the sampling frequency of the retrieved sound information, the restoration step is to estimate, as a reference generation process, a frequency component that is greater than half the sampling frequency of the sound information subjected to the super-resolution processing and less than or equal to half the sampling frequency of the retrieved sound information.

8. The information processing method according to claim 1, wherein the restoration step restores at least one of the direction and position of the sound source of the extracted target sound as the format.

9. The information processing method according to claim 8, wherein the restoration step is to perform beamforming.

10. The information processing method according to claim 9, wherein the restoration step involves generating a reference from the extracted target sound that has been processed to have the same sampling frequency as the retrieved sound information, and then performing beamforming on the retrieved sound information based on the generated reference.

11. The information processing method according to claim 10, wherein the restoration step involves estimating the phase for each channel of the retrieved sound information with respect to the sound information on which beamforming has been performed.

12. The information processing method according to claim 11, wherein the restoration step is performed by estimating the phase by projection processing, which projects the beamforming sound information into the projection space.

13. The information processing method according to claim 12, wherein the restoration step involves estimating the phase by MDP (Minimal Distortion Principle) as the projection process.

14. The information processing method according to claim 12, wherein the restoration step involves estimating the phase using a Single-Channel Wiener Filter (SWF) as the projection process.

15. The information processing method according to claim 1, wherein the target sound extraction step receives input of a query received prior to the search step.

16. The information processing method according to claim 1, wherein the search step specifies a time range corresponding to the received query as sound information.

17. The information processing method according to claim 16, wherein the search step is to display the identified time range.

18. The information processing method according to claim 1, wherein the reception step receives a plurality of queries as the query, and the search step searches for the sound information based on the plurality of queries received.

19. An information processing system comprising: a reception unit that receives input of a query, which is information for searching for sound information; a search unit that searches for sound information based on the received query; a target sound extraction unit that extracts a target sound, which is sound information corresponding to the received query, from the searched sound information; and a restoration unit that restores the format of the query-based target sound extraction result, which is the output of the target sound extraction unit, based on the searched sound information.

20. An information processing method performed by a computer, comprising: a reception step of receiving input of a query which is information for searching for sound information; and a presentation step of presenting a target sound extraction result in which a target sound which is sound information corresponding to the received query has the same format as the sound information searched based on the received query.