Sound estimation model acquisition device, sound estimation device, sound estimation model acquisition method, sound estimation method and program

The sound estimation model acquisition device improves sound estimation accuracy by using a mathematical model to refine time series representation through machine learning, addressing the limitations of text and acoustic signal queries.

JP7727219B2Active Publication Date: 2025-08-21NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023573825
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-01-17
Filing Date
2022-07-04
Publication Date
2025-08-21
Estimated Expiration
2042-07-04

AI Technical Summary

Technical Problem

Existing sound estimation methods using text queries lack precision due to the difficulty in perfectly describing acoustic signals with text, while acoustic signal queries are inflexible and less accurate due to the challenge of adjusting search ranges.

Method used

A sound estimation model acquisition device that utilizes a mathematical model to estimate a time series based on sound time series and difference information, improving accuracy by machine learning and vector representation techniques.

Benefits of technology

Enhances the precision of sound estimation by reducing the difference between input and estimated time series, allowing for more accurate acoustic signal retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007727219000013
    Figure 0007727219000013
  • Figure 0007727219000014
    Figure 0007727219000014
  • Figure 0007727219000015
    Figure 0007727219000015
Patent Text Reader

Abstract

One embodiment of the present invention is a sound estimation model acquisition device comprising a model acquisition unit for obtaining a mathematical model that estimates an estimated time series satisfying a predetermined estimation condition from among one or a plurality of time series, the mathematical model being obtained on the basis of a first sound time series that indicates a first sound, a second sound time series that indicates a second sound, and first difference information that indicates at least some of the difference between the first and the second sounds. The mathematical model estimates the estimated time series on the basis of an input time series that indicates an inputted sound and second difference information that indicates at least some of the difference between the input time series and the estimated time series. The estimation condition represents whether the difference between the difference from the input time series and the difference indicated by the second difference information is smaller than a prescribed difference.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound estimation model acquisition device, a sound estimation device, a sound estimation model acquisition method, a sound estimation method, and a program. [Background technology]

[0002] The number of sound time series available on the Internet is increasing every day, making it essential to have a method for searching for target acoustic signals from a large database of acoustic signals. Most of the methods proposed so far can be broadly divided into two types: methods that use text as a query and methods that use acoustic signals as a query.

[0003] In the text-query method, data pairs consisting of an acoustic signal and a text containing explanatory sentences are used. The text-query method then uses these data to learn a common latent space between the acoustic signal and the text, and searches for acoustic signals with embeddings similar to those of the input text query.

[0004] The acoustic signal query approach is also called content-based approach, which searches for acoustic signals with similar embeddings to the input acoustic signal query. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] P. Manocha, R. Badlani, A. Kumar, A. Shah, B. Elizalde, and B. Raj, "Content-based representations of audio using siamese neural networks", in Proc. IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP). IEEE, 2018, pp. 3136-3140. Summary of the Invention [Problem to be solved by the invention]

[0006] Since text queries can be easily edited manually, text-based query methods allow for flexible adjustment of the search scope. However, since it is difficult to perfectly describe the content of an acoustic signal using text alone, the search may be too general. As a result, the accuracy of acoustic signal estimation may be low.

[0007] In the acoustic signal query method, information that cannot be expressed in text can be used as a query in the form of an acoustic signal. This allows for more detailed search range settings than in the text query method. However, since the center of the search range is embedded in the acoustic signal query and editing the acoustic signal is not easy, flexible adjustment of the search range can be difficult. As a result, the accuracy of acoustic signal estimation can be low.

[0008] This situation is not limited to acoustic signals, but is common to time series of sounds such as voice signals.

[0009] In view of the above circumstances, an object of the present invention is to provide a technique for improving the accuracy of estimating a time series representing a sound. [Means for solving the problem]

[0010] One aspect of the present invention is a sound estimation model acquisition device including: a model acquisition unit that acquires a mathematical model that estimates an estimated time series, which is a time series that satisfies a predetermined estimation condition from one or more time series, based on a first sound time series that is a time series indicating a first sound, a second sound time series that is a time series that is indicating a second sound, and first difference information that indicates at least a portion of the differences between the first sound and the second sound, wherein the mathematical model estimates the estimated time series based on an input time series that is a time series that indicates an input sound, and second difference information that is information that indicates at least a portion of the differences between the input time series and the estimated time series, and the estimation condition is a condition that the difference between the input time series and the difference indicated by the second difference information is smaller than a predetermined difference.

[0011] One aspect of the present invention includes a query acquisition unit that acquires query information including a sound time series that is a time series indicating a sound and second difference information that indicates at least a part of a difference between an estimated time series that is a time series that satisfies a predetermined estimation condition and the sound time series, and a model acquisition unit that acquires a mathematical model that estimates a time series that satisfies the estimation condition from one or more time series based on a first sound time series that is a time series indicating a first sound, a second sound time series that is a time series that indicates a second sound, and first difference information that indicates at least a part of a difference between the first sound and the second sound, wherein the mathematical model is based on the first sound time series that indicates an input sound. and an estimation unit that estimates a time series that satisfies the estimation condition from one or more time series, using the mathematical model acquired by a sound estimation model acquisition device, based on the query information acquired by the query acquisition unit, wherein the mathematical model is a mathematical model that estimates a time series that satisfies the estimation condition, based on an input time series that is a time series, and second difference information that is information that indicates at least a part of a difference between the input time series and a time series that satisfies the estimation condition, and the estimation condition is a condition that a difference between the input time series and a difference indicated by the second difference information is smaller than a predetermined difference.

[0012] One aspect of the present invention is a sound estimation model acquisition method including: a sound estimation model acquisition step of acquiring a mathematical model that estimates an estimated time series, which is a time series that satisfies a predetermined estimation condition from one or more time series, based on a first sound time series that is a time series indicating a first sound, a second sound time series that is a time series that is indicating a second sound, and first difference information that indicates at least a portion of the differences between the first sound and the second sound, wherein the mathematical model is a mathematical model that estimates the estimated time series based on an input time series that is a time series that indicates an input sound, and second difference information that is information that indicates at least a portion of the differences between the input time series and the estimated time series, and the estimation condition is a condition that the difference between the input time series and the difference indicated by the second difference information is smaller than a predetermined difference.

[0013] One aspect of the present invention includes a query acquisition step of acquiring query information including a sound time series that is a time series indicating a sound and second difference information that indicates at least a part of a difference between an estimated time series that is a time series that satisfies a predetermined estimation condition, and a model acquisition unit that acquires a mathematical model for estimating a time series that satisfies the estimation condition from one or more time series, based on a first sound time series that is a time series indicating a first sound, a second sound time series that is a time series indicating a second sound, and first difference information that indicates at least a part of a difference between the first sound and the second sound, wherein the mathematical model is a time series indicating an input sound. an estimation step of estimating a time series that satisfies the estimation condition from one or more time series, based on the query information acquired in the query acquisition step, using the mathematical model acquired by a sound estimation model acquisition device, the mathematical model being configured to estimate a time series that satisfies the estimation condition based on an input time series and second difference information that is information indicating at least a part of a difference between the input time series and a time series that satisfies the estimation condition, wherein the estimation condition is a condition that a difference between the input time series and a difference indicated by the second difference information is smaller than a predetermined difference.

[0014] One aspect of the present invention is a program for causing a computer to function as the above-described sound estimation model acquisition device.

[0015] One aspect of the present invention is a program for causing a computer to function as the above-described sound estimation device. [Effects of the Invention]

[0016] The present invention makes it possible to improve the accuracy of estimation of a time series representing a sound. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 is an explanatory diagram illustrating an overview of a sound estimation system according to an embodiment. [Figure 2] FIG. 1 is an explanatory diagram illustrating an example of machine learning processing according to an embodiment. [Figure 3] FIG. 1 is a diagram showing an example of the hardware configuration of a sound estimation model acquisition device according to an embodiment. [Figure 4] FIG. 2 is a diagram showing an example of the functional configuration of a control unit included in the sound estimation model acquisition device according to the embodiment. [Figure 5] 1 is a flowchart showing an example of the flow of processing executed by a sound estimation model acquisition device according to an embodiment. [Figure 6] FIG. 1 is a diagram showing an example of the hardware configuration of a sound estimation device according to an embodiment. [Figure 7] FIG. 2 is a diagram showing an example of the functional configuration of a control unit included in the sound estimation device according to the embodiment. [Figure 8] 1 is a flowchart showing an example of a flow of processing executed by a sound estimation device according to an embodiment. [Figure 9] FIG. 1 is a first diagram showing an example of an experimental result in the embodiment. [Figure 10] FIG. 2 is a second diagram showing an example of an experimental result in the embodiment. [Figure 11] FIG. 3 is a third diagram showing an example of an experimental result in the embodiment. [Figure 12] FIG. 10 is a diagram showing an example of the configuration of a learning network in a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0018] (Embodiment) 1 is an explanatory diagram illustrating an overview of a sound estimation system 100 according to an embodiment. The sound estimation system 100 includes a sound estimating model acquiring apparatus 1 and a sound estimating apparatus 2.

[0019] The sound estimation model acquisition device 1 obtains a sound estimation model based on a first sound time series that is a time series indicating a first sound, a second sound time series that is a time series indicating a second sound, and first difference information that indicates at least a part of the differences between the first sound and the second sound.

[0020] The first difference information is, for example, text data. When the first difference information is text data, the text indicated by the first difference information is, for example, the text "Add the sound of thunder."

[0021] The first difference information may be, for example, image data. When the first difference information is image data, the image indicated by the first difference information is, for example, an image of lightning.

[0022] The sound estimation model is a mathematical model that estimates a time series that satisfies an estimation condition (hereinafter referred to as an "estimated time series") from one or more pre-given time series. The estimation condition is a condition that the difference between the input time series and the difference indicated by the second difference information is smaller than a predetermined difference. For example, the estimation condition is a condition that the difference between the input time series is closest to the difference indicated by the second difference information.

[0023] The input time series is a time series input to the sound estimation model and indicates sound. Hereinafter, a time series indicating sound will be referred to as a sound time series. It can be said that the input time series is a sound time series input to the sound estimation model. The second difference information is information indicating at least a part of the differences between the input time series and the estimated time series.

[0024] The second difference information is information indicating at least a portion of the differences between the input time series and the estimated time series. Therefore, the second difference information input to the sound estimation model may be, for example, information indicating a portion of the differences between the input time series and the estimated time series.

[0025] For example, if the first difference information is text data, the second difference information is text data. For example, if the first difference information is image data, the second difference information is image data.

[0026] The sound estimation model acquisition device 1 acquires a sound estimation model by, for example, machine learning. More specifically, the sound estimation model acquisition device 1 updates the sound estimation model by machine learning until a predetermined termination condition related to learning (hereinafter referred to as a "learning termination condition") is satisfied, thereby acquiring a sound estimation model with higher estimation accuracy than before the learning termination condition was satisfied.

[0027] In machine learning, for example, a process of updating the sound estimation model is executed so as to reduce the difference between the time series estimated by the sound estimation model based on the first sound time series and the first difference information and the second sound time series. In this way, the accuracy of estimation of the sound estimation model is improved by machine learning.

[0028] The sound estimation model at the time when the learning termination condition is satisfied is the sound estimation model obtained by the sound estimation model acquisition device 1 based on the first sound time series, the second sound time series, and the first difference information. Hereinafter, the sound estimation model at the time when the learning termination condition is satisfied will be referred to as the trained sound estimation model. A trained sound estimation model is one type of sound estimation model.

[0029] FIG. 1 also shows an overview of a process for obtaining a trained sound estimation model through machine learning. Model M101 in FIG. 1 is an example of a sound estimation model before a learning termination condition is satisfied. Set G101 is an example of a set of candidate time series that can be estimated as an estimated time series by model M101. Time series belonging to set G101 are time series that indicate sounds. Information D101 is an example of a first sound time series, and information D102 is an example of first difference information.

[0030] Information D103 shown in Figure 1 is an example of the estimation result of model M101. That is, information D103 is an example of a time series that is estimated by model M101 to be an estimated time series from among time series belonging to set G101. Information D104 shown in Figure 1 is an example of a second sound time series. Figure 1 shows that model M101 is updated so that the difference between the estimation result of model M101 (i.e., information D103) and information D104 becomes smaller.

[0031] Here, a more detailed example of learning will be described. In the learning stage, the sound estimation model converts the first sound time series into a vector representing the first sound time series (hereinafter referred to as the "first sound vector"). In the learning stage, the sound estimation model converts the second sound time series into a vector representing the second sound time series (hereinafter referred to as the "second sound vector"). In the learning stage, the sound estimation model converts the first difference time series into a vector representing the first difference time series (hereinafter referred to as the "first difference vector").

[0032] In the learning stage, the sound estimation model estimates a time series that satisfies the estimation conditions from among the time series belonging to set G101, based on the first sound vector and the first difference vector. In the learning stage, the sound estimation model is updated so as to reduce the difference between the vector representing the estimated time series and the second sound vector.

[0033] The difference can be expressed by any index that indicates the distance between vectors in a vector space that represents the second sound vector. The distance between vectors can be expressed, for example, by the dot product between the vectors. The distance between vectors can be expressed, for example, by the cosine similarity between the vectors.

[0034] The distance between vectors may be expressed, for example, by the following formula (1), formula (2), or formula (3). The following formulas (1) to (3) all define the distance between vector x and vector y.

[0035]

number

[0036]

number

[0037]

number

[0038] Note that ||·||2 means the L2 norm.

[0039] The update of the sound estimation model specifically includes updating the content of the embedding process that converts the first sound time series into the first sound vector. The update of the sound estimation model specifically includes updating the content of the embedding process that converts the second sound time series into the second sound vector. The update of the sound estimation model specifically includes updating the content of the embedding process that converts the first difference time series into the first difference vector.

[0040] This updating will be explained using a formula. For example, updating is a mapping F in the process expressed by the following formula (4): q or a mapping F t This means updating the value of σ so as to satisfy the condition expressed by the following equation (4).

[0041]

number

[0042] a means the first sound time series, t means the first difference time series, and b means the second sound time series. Mapping F q represents a process executed on an estimated time series estimated by the sound estimation model based on a pair of the first and second arguments, and which obtains a vector representing the estimated time series.

[0043] Mapping F t represents the process of obtaining a vector that represents the sound time series indicated by the first argument. The mapping d means the process of obtaining the distance in vector space between the vector indicated by the first argument and the vector indicated by the second argument. Therefore, the mapping d represents the process of obtaining the dot product between the vector indicated by the first argument and the vector indicated by the second argument. Equation (1) means that updating is performed to minimize the distance d.

[0044] Mapping F q Specifically, for example, is expressed by the following formula (5).

[0045]

number

[0046] Mapping A represents the process of converting the sound time series of the first argument into the first sound vector. Mapping T represents the process of converting the first difference time series t into the first difference vector.

[0047] Mapping F t Specifically, for example, is expressed by the following equation (6).

[0048]

number

[0049] In the example of equation (6), the mapping F qThe image of the map F is expressed as the result of adding the image of the map A to the image of the map T. However, the map F q The image of a map F does not necessarily have to be expressed as the result of addition, as long as it is expressed as the result of a binary operation on the image of a map A and the image of a map T, as shown in the following equation (7). q The image of may be expressed as, for example, the image shown in the following equation (8).

[0050]

number

[0051]

number

[0052] The mapping f in equation (7) represents a binary operation on the first and second arguments. * in equation (8) means taking the element-wise product (i.e., the Hadamard product).

[0053] As described above, in the learning stage, the sound estimation model is updated so as to reduce the difference between the vector representing the estimated time series and the second sound vector. Therefore, for example, in learning the sound estimation model, the first difference vector is added to the first sound vector, and the sound time series having the feature vector closest to the sum is output. Then, in learning the sound estimation model, the sound estimation model is updated so as to reduce the difference between the vector representing the output sound time series and the second sound vector.

[0054] Fig. 2 is an explanatory diagram illustrating an example of machine learning processing in an embodiment. More specifically, Fig. 2 is a diagram illustrating an example of a learning network 10, which is a neural network that learns a sound estimation model.

[0055] The learning network 10 includes a first encoder 101, a second encoder 102, and a third encoder 103.

[0056] The first encoder 101 is an encoder whose processing content is represented by mapping A and into which the first sound time series is input. The second encoder 102 is an encoder whose processing content is represented by mapping T and into which first difference information is input. The third encoder 103 is an encoder whose processing content is represented by mapping A and into which the second sound time series is input.

[0057] The first encoder 101 and the third encoder 103 are neural networks with a layer that performs VGGish, a layer that performs pooling, and a layer that performs projection. The second encoder 102 is a neural network with a layer that performs DistiBERT and a layer that performs projection.

[0058] 2, the learning of the first encoder 101, the third encoder 103, and the second encoder 102 is performed, for example, by cross-modal contrastive learning. The distance between vectors used in the learning is expressed, for example, by the following equation (9) in FIG.

[0059]

number

[0060] The learning network 10 is updated by an updating unit 104. The updating unit 104 calculates a value of a loss function based on the output of the learning network 10. The updating unit 104 updates the learning network 10 based on the calculated value of the loss function. The loss function in the example of FIG. 2 is, for example, cosine similarity expressed by the following equation (10).

[0061]

number

[0062]

number

[0063] Note that τ represents a temperature parameter. The value of the temperature parameter τ may be 0. Note that i and j are indexes within the batch.

[0064] In the example of Figure 2, the parameter values ​​of VGGish and DistilBERT are fixed, and the parameter values ​​of VGGish and DistilBERT do not need to be updated by learning. In such a case, when updating the contents of mapping A and mapping T, for example, the parameter values ​​of projection are updated by learning, and the parameter values ​​of VGGish and DistilBERT are not updated.

[0065] Returning to the explanation of Fig. 1, the sound estimation device 2 estimates an estimated time series based on query information, using the sound estimation model acquired by the sound estimation model acquisition device 1. The query information includes a query time series, which is a sound time series input to the sound estimation device 2, and query difference information. The query difference information is second difference information that indicates at least a portion of the differences between the estimated time series and the query time series.

[0066] Note that the sound estimation model used by the sound estimation device 2 is a trained sound estimation model when the sound estimation model acquisition device 1 obtains the sound estimation model through machine learning. Note that the sound time series input to the sound estimation device 2 is input to the sound estimation model. Therefore, the query time series is also the sound time series input to the sound estimation model.

[0067] 3 is a diagram showing an example of the hardware configuration of a sound estimation model acquisition device 1 in an embodiment. The sound estimation model acquisition device 1 is equipped with a control unit 11 having a processor 91 such as a CPU (Central Processing Unit) and a memory 92 connected via a bus, and executes a program. By executing the program, the sound estimation model acquisition device 1 functions as a device having the control unit 11, input unit 12, communication unit 13, storage unit 14, and output unit 15.

[0068] More specifically, in the sound estimation model acquisition device 1, the processor 91 reads a program stored in the storage unit 14 and stores the read program in the memory 92. The processor 91 executes the program stored in the memory 92, causing the sound estimation model acquisition device 1 to function as a device including a control unit 11, an input unit 12, a communication unit 13, a storage unit 14, and an output unit 15.

[0069] The control unit 11 controls the operations of various functional units included in the sound estimation model acquisition device 1. The control unit 11 performs, for example, processing to acquire a sound estimation model.

[0070] The input unit 12 includes input devices such as a mouse, a keyboard, a touch panel, etc. The input unit 12 may also include an interface that connects these input devices to the sound estimation model acquisition device 1.

[0071] The communication unit 13 is configured to include an interface for connecting the sound estimation model acquisition device 1 to an external device. The communication unit 13 communicates with the external device via wired or wireless communication. The external device is, for example, a device that transmits a set of a first sound time series, first difference information, and a second sound time series. The communication unit 13 acquires the set of a first sound time series, first difference information, and a second sound time series by communicating with the device that transmits the set of a first sound time series, first difference information, and a second sound time series. The external device is, for example, the sound estimation device 2. The communication unit 13 transmits the sound estimation model to the sound estimation device 2 by communicating with the sound estimation device 2.

[0072] The storage unit 14 is configured using a computer-readable storage medium device such as a magnetic hard disk device or a semiconductor storage device. The storage unit 14 stores various types of information related to the sound estimation model acquisition device 1. The storage unit 14 stores, for example, various types of information generated as a result of processing executed by the control unit 11. The storage unit 14 stores in advance one or more time series that are candidates for an estimated time series, such as each time series belonging to set G101.

[0073] The output unit 15 is configured to include a display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, an organic EL (Electro-Luminescence) display, etc. The output unit 15 may be configured to include an interface that connects these display devices to the sound estimation model acquisition device 1.

[0074] 4 is a diagram illustrating an example of the functional configuration of the control unit 11 included in the sound estimation model acquisition device 1 of the embodiment. The control unit 11 includes an information acquisition unit 111, a model acquisition unit 112, a communication control unit 113, a memory control unit 114, and an output control unit 115.

[0075] The information acquisition unit 111 acquires data input to the communication unit 13. The information acquisition unit 111 acquires, for example, a first sound time series, first difference information, and a second sound time series.

[0076] The model acquisition unit 112 acquires a sound estimation model using one or more sets of a first sound time series, first difference information, and a second sound time series. The model acquisition unit 112 includes, for example, a learning network 10 and an update unit 104. In this case, the model acquisition unit 112 acquires a trained sound estimation model by updating the learning network 10 through machine learning until a learning termination condition is satisfied.

[0077] The communication control unit 113 controls the operation of the communication unit 13. The memory control unit 114 controls the operation of the memory unit 14. The output control unit 115 controls the operation of the output unit 15.

[0078] 5 is a flowchart showing an example of the flow of processing executed by the sound estimation model acquisition device 1 in an embodiment. The information acquisition unit 111 acquires one or more sets of a first sound time series, first difference information, and a second sound time series (step S101). Next, the model acquisition unit 112 acquires a sound estimation model using each set of the first sound time series, first difference information, and second sound time series acquired by the information acquisition unit 111 (step S102).

[0079] The sound estimation model obtained in this manner is used in the sound estimation device 2.

[0080] 6 is a diagram showing an example of the hardware configuration of a sound estimation device 2 according to an embodiment. The sound estimation device 2 includes a control unit 21 having a processor 93 such as a CPU and a memory 94 connected via a bus, and executes a program. By executing the program, the sound estimation device 2 functions as a device including the control unit 21, an input unit 22, a communication unit 23, a storage unit 24, and an output unit 25.

[0081] More specifically, in the sound estimation device 2, the processor 93 reads a program stored in the storage unit 24 and stores the read program in the memory 94. When the processor 93 executes the program stored in the memory 94, the sound estimation device 2 functions as a device including a control unit 21, an input unit 22, a communication unit 23, a storage unit 24, and an output unit 25.

[0082] The control unit 21 controls the operations of various functional units included in the sound estimation device 2. The control unit 21 executes, for example, a sound estimation model. The sound estimation model executed by the control unit 21 is, for example, a trained sound estimation model.

[0083] The input unit 22 includes input devices such as a mouse, a keyboard, and a touch panel. The input unit 22 may include an interface that connects these input devices to the sound estimation device 2. For example, a set of a query time series and query difference information is input to the input unit 22.

[0084] The communication unit 23 is configured to include an interface for connecting the sound estimation device 2 to an external device. The communication unit 23 communicates with the external device via wired or wireless communication. The external device is, for example, a device that has transmitted query information. The communication unit 23 acquires query information by communicating with the device that has transmitted the query information.

[0085] The external device is, for example, a sound estimation model acquisition device 1. The communication unit 23 acquires a sound estimation model by communicating with the sound estimation model acquisition device 1. The sound estimation model acquired by the communication unit 23 by communicating with the sound estimation model acquisition device 1 is, for example, a trained sound estimation model.

[0086] The storage unit 24 is configured using a computer-readable storage medium device such as a magnetic hard disk device or a semiconductor storage device. The storage unit 24 stores various information related to the sound estimation device 2. The storage unit 24 stores, for example, various information generated as a result of processing executed by the control unit 21. The storage unit 24 stores in advance, for example, one or more time series that are candidates for an estimated time series.

[0087] The output unit 25 includes a display device such as a CRT display, a liquid crystal display, an organic EL display, etc. The output unit 25 may also include an interface that connects these display devices to the sound estimation device 2.

[0088] 7 is a diagram illustrating an example of the functional configuration of the control unit 21 included in the sound estimation device 2 of the embodiment. The control unit 21 includes a query acquisition unit 211, an estimation unit 212, a communication control unit 213, a storage control unit 214, and an output control unit 215.

[0089] The query acquisition unit 211 acquires the query information input to the communication unit 23 .

[0090] The estimation unit 212 executes the sound estimation model on the query information acquired by the query acquisition unit 211. By executing the sound estimation model on the query information, the estimation unit 212 estimates a time series that satisfies the estimation conditions from one or more time series pre-stored in the storage unit 24. The sound estimation model executed by the estimation unit 212 is, for example, a trained sound estimation model.

[0091] The communication control unit 213 controls the operation of the communication unit 23. The memory control unit 214 controls the operation of the memory unit 24. The output control unit 215 controls the operation of the output unit 25.

[0092] 8 is a flowchart showing an example of the flow of processing executed by the sound estimation device 2 of the embodiment. The query acquisition unit 211 acquires query information (step S201). Next, the estimation unit 212 executes a sound estimation model on the query information acquired in step S201 (step S202). After step S202, the output control unit 215 controls the operation of the output unit 25 to cause the output unit 25 to output the result of estimation by the estimation unit 212 (step S203).

[0093] <Experimental Results> An example of the results of constructing an Audio Pair with Difference (APwD) dataset and verifying the effectiveness of the sound estimation system 100 is described below. Specifically, the sound estimation system 100 used in the experiment was a sound estimation system 100 that was equipped with the learning network 10 and update unit 104 shown in FIG. 2 and that obtained a sound estimation model through machine learning.

[0094] The APwD dataset is a dataset synthesized using audio files from FSD50k and ESC-50, which are datasets for audio tagging. It is a dataset consisting of triplets of pairs of similar but different acoustic signals and text describing the differences.

[0095] In the experiment, two scenes, Rain and Traffic, were prepared from the APwD dataset. For the Rain scene, background sounds labeled "rain" by FSD50k and event sounds labeled "dog", "chirping_bird", "thunder", and "footsteps" by ESC-50 were used for synthesis.

[0096] For the Traffic scene, background sound data labeled "car_passing" by FSD50k and event sound data labeled "dog", "chirping_bird", "car_horn", and "church_bell" by ESC-50 were used for synthesis. A training set of 50,000 pairs and an evaluation set of 1,000 pairs were prepared for each of the two scenes and used in the experiments.

[0097] In addition, to confirm whether the model can adapt to combinations of background sounds and event sounds that are not included in the training set, two scenes containing such combinations of event sounds, Rain with unknown events and Traffic with unknown events, were also synthesized.

[0098] For these scenes, background sounds labeled "rain" or "car_passing" by FSD50k were used, and event sounds labeled "dog", "chirping_bird", "thunder", "footsteps", "car_horn", and "church_bell" by ESC-50 were used. For these two scenes, only 1000 pairs were prepared for evaluation.

[0099] In the experiment, training was performed using three data sets: Rain-tr, Traffic-tr, and Both-tr, and the evaluation results were compared. Rain-tr contains 50,000 pairs of training sets from rain scenes. Traffic-tr contains 50,000 pairs of training sets from traffic scenes. Both-tr contains 25,000 pairs from the training set from rain scenes and 25,000 pairs from the training set from traffic scenes.

[0100] For training, 10% of each training set was used as a validation set. For optimization, Adam was used for 300 epochs, and the model with the smallest loss function value on the validation set was used for evaluation.

[0101] The evaluation used Recall@K, which is the percentage of correct candidates among the top K search candidates. In addition, a technology in which the text encoder part (i.e., the second encoder 102) was removed from the sound estimation system 100 (hereinafter referred to as the "comparison technology") was used as a baseline technology for comparison with the sound estimation system 100.

[0102] Fig. 9 is a first diagram showing an example of experimental results in an embodiment. More specifically, Fig. 9 shows the results of evaluating models trained using Rain-tr and Traffic-tr on the Rain and Traffic evaluation sets, respectively. The "Proposed" row in Fig. 9 shows the results of evaluating the sound estimation system 100. The "Baseline" row in Fig. 9 shows the results of evaluating the comparison technology.

[0103] The results in Figure 9 show that the estimation accuracy of the sound estimation system 100 exceeds that of the comparative technology under all conditions. Therefore, Figure 9 shows that the estimation accuracy is improved by using the first difference information and the second difference information. Note that the first difference information and the second difference information in the experiment are data expressed as text data.

[0104] Fig. 10 is a second diagram showing an example of experimental results in the embodiment. More specifically, Fig. 10 shows the results of comparing Recall@1 for each event where there is a difference in Recall@1. The results in Fig. 10 show that the estimation accuracy of the sound estimation system 100 exceeds that of the comparison technology.

[0105] Fig. 11 is a third diagram showing an example of experimental results in an embodiment. Specifically, Fig. 11 is a diagram showing an example of experimental results of an experiment to confirm whether a model can adapt to a combination of background sound and event sound that is not included in the training set in the sound estimation system 100. In the experiment to obtain the results of Fig. 11, a Rain-tr model, a Traffic-tr model, and a Both-tr model were evaluated using an evaluation set of Rain with unknown events and Traffic with unknown events.

[0106] The Rain-tr model is a mathematical model obtained by learning using Rain-tr. The Traffic-tr model is a mathematical model obtained by learning using Traffic-tr. The Both-tr model is a mathematical model obtained by learning using Both-tr.

[0107] Each value in Fig. 11 indicates Recall@1, and "*1" in Fig. 11 indicates an event sound that is not included in the Rain scene. "*2" in Fig. 11 indicates an event sound that is not included in the Traffic scene.

[0108] Figure 11 shows that among the evaluations of the models obtained by training using Rain-tr and Traffic-tr, the evaluations of car_horn or church_bell in Rain with unknown events and thunder or footsteps in Traffic with unknown events were lower than the results of the other evaluations.

[0109] The car_horn or church_bell in Rain with unknown events and the thunder or footsteps in Traffic with unknown events are information that is not included in the respective scenes.

[0110] On the other hand, Figure 11 shows that such a drop in evaluation is significantly reduced in the model obtained by training using Both-tr. Therefore, Figure 11 shows that the sound estimation system 100 can perform acoustic search by separating information on background sounds and differences and applying the learned differences to various background sounds.

[0111] The sound estimation model acquisition device 1 configured in this way acquires a sound estimation model using the first difference data, thereby improving the accuracy of estimation of a time series indicating sound.

[0112] Furthermore, the sound estimation device 2 configured in this manner estimates an estimation target using the mathematical model obtained by the sound estimation model acquisition device 1. Therefore, the sound estimation device 2 can improve the accuracy of estimation of a time series indicating a sound.

[0113] (Variation) The sound estimation model acquisition device 1 may acquire a sound estimation model based on the first sound time series, the second sound time series, the first difference information, and further based on the first sound event class label and the second sound event class label. That is, the model acquisition unit 112 may acquire a sound estimation model based on the first sound time series, the second sound time series, the first difference information, and further based on the first sound event class label and the second sound event class label. The first sound event class label indicates the classification result of the first sound time series by a predetermined classification process. The second sound event class label indicates the classification result of the second sound time series by a classification process similar to the classification process for the first sound time series. The predetermined classification process is a classification process that targets a sound time series as a classification object, and may be any classification process as long as it classifies the target of classification by a method that can classify a sound time series.

[0114] The predetermined classification process may be, for example, a process of classifying the first sound time series or the second sound time series according to a sound recognized by the user when the user hears the sound represented by the sound time series to be classified. The sound recognized by the user when the user hears the sound represented by the sound time series to be classified may be, for example, the sound of rain, a dog barking, or a person speaking. The predetermined classification process may also be, for example, a process of classifying the sound time series to be classified according to a location where the sound represented by the sound time series to be classified was recorded. The location where the sound represented by the sound time series to be classified may be, for example, a home, an airport, or a restaurant. The predetermined classification process may also be, for example, a process of classifying the sound time series to be classified according to a characteristic of speech contained in the sound represented by the sound time series to be classified. The characteristic of speech may be, for example, the gender of the speaker, the age of the speaker, or the language spoken.

[0115] Such a sound estimation model acquisition device 1 may include a learning network 10a instead of the learning network 10. In such a case, the target of update by the update unit 104 is the learning network 10a instead of the learning network 10. The learning network 10a differs from the learning network 10 of the embodiment in that it includes a classification network 105. The learning network 10a receives as input a first sound time series, a second sound time series, first difference information, a first sound event class label, and a second sound event class label.

[0116] Fig. 12 is a diagram showing an example of the configuration of a learning network 10a in a modified example. A classification network 105 classifies input information. The information input to the classification network 105 is the output of a first encoder 101 and the output of a third encoder 103. In the example of Fig. 12, the learning network 10a includes two classification networks 105, one of which receives the output of the first encoder and the other of which receives the output of the third encoder 103.

[0117] Since the classification network 105 classifies the input information, the output of the classification network 105 when the output of the first encoder 101 is input to the classification network 105 is the estimated result of the first sound event class label. The output of the classification network 105 when the output of the third encoder 103 is input to the classification network 105 is the estimated result of the second sound event class label. The classification network 105 is updated when the update unit 104 updates the learning network 10a.

[0118] The update unit 104 that updates the training network 10a uses a loss function including a first auxiliary loss function and a second auxiliary loss function to update the training network 10a. The first auxiliary loss function indicates the difference between the estimation result of the classification network 105 to which the output of the first encoder 101 is input and the first sound event class label. The second auxiliary loss function indicates the difference between the estimation result of the classification network 105 to which the output of the third encoder 103 is input and the second sound event class label.

[0119] The function representing the sum of the first sub-loss function and the second sub-loss function is, for example, a function expressed by the following equation (12).

[0120]

number

[0121] BCE stands for binary cross entropy. C(·) is a function representing the estimation process by the classification network 105. The update unit 104, which updates the learning network 10a, uses a loss function (hereinafter referred to as the "integrated loss function") that includes a function representing the sum of the first sub-loss function and the second sub-loss function in addition to the loss function used by the learning network 10 in equation (10).

[0122] The update unit 104, which updates the learning network 10a, updates the learning network 10a so as to reduce the integrated loss function. As the integrated loss function decreases, the values ​​of the first sub-loss function and the second sub-loss function also decrease. Therefore, the update unit 104, which updates the learning network 10a, updates the learning network 10a so as to reduce the values ​​of the first sub-loss function and the second sub-loss function.

[0123] Therefore, the integrated loss function is, for example, L overall =L+λL cl The loss function is expressed as L overall is the integrated loss function. λ is a weighting parameter. λ is a predetermined value.

[0124] In this way, by the sound estimation model acquisition device 1 using the learning network 10a instead of the learning network 10, the vectors output from the first to third encoders come to have features corresponding to the "different sound event types", which results in an effect of improving the estimation accuracy of the sound estimation device 2. In other words, by the model acquisition unit 112 obtaining a sound estimation model using the first sound time series, the second sound time series, the first difference information, and further the first sound event class label and the second sound event class label, the vectors output from the first to third encoders come to have features corresponding to the "different sound event types", which results in an effect of improving the estimation accuracy of the sound estimation device 2.

[0125] The sound estimation device 2 may estimate an estimation target using a mathematical model obtained by the sound estimation model acquisition device 1 that uses a learning network 10a instead of the learning network 10.

[0126] It should be noted that the sound estimation model acquisition device 1 and the sound estimation device 2 do not necessarily need to be implemented as different devices. The sound estimation model acquisition device 1 and the sound estimation device 2 may be implemented as, for example, a single device or system that combines the functions of both devices.

[0127] Furthermore, the functional units of the sound estimation model acquisition device 1 and the sound estimation device 2 may be implemented using a plurality of information processing devices communicably connected via a network.

[0128] Each of the sound estimation model acquisition device 1 and the sound estimation device 2 may be implemented using a plurality of information processing devices communicably connected via a network. All or part of the functions of each of the sound estimation model acquisition device 1 and the sound estimation device 2 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into a computer system. The program may be transmitted via a telecommunications line.

[0129] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Explanation of symbols]

[0130] 100...sound estimation system, 1...sound estimation model acquisition device, 2...sound estimation device, 11...control unit, 12...input unit, 13...communication unit, 14...storage unit, 15...output unit, 111...information acquisition unit, 112...model acquisition unit, 113...communication control unit, 114...storage control unit, 115...output control unit, 21...control unit, 22...input unit, 23...communication unit, 24...storage unit, 25...output unit, 211...query acquisition unit, 212...estimation unit, 213...communication control unit, 214...storage control unit, 215...output control unit, 91, 93...processor, 92, 94...memory, 10, 10a...learning network, 105...classification network

Claims

1. a model acquisition unit that acquires a mathematical model for estimating an estimated time series, which is a time series that satisfies a predetermined estimation condition, from one or more time series, based on a first sound time series that is a time series indicating a first sound, a second sound time series that is a time series indicating a second sound, and first difference information that indicates at least a part of the differences between the first sound and the second sound; Equipped with the mathematical model is a mathematical model that estimates the estimated time series based on an input time series that is a time series indicating an input sound and second difference information that is information indicating at least a part of a difference between the input time series and the estimated time series, the estimation condition is a condition that a difference between the difference with the input time series and the difference indicated by the second difference information is smaller than a predetermined difference. Sound estimation model acquisition device.

2. the first difference information and the second difference information are text data. The sound estimation model acquisition device according to claim 1 .

3. the first difference information and the second difference information are image data; The sound estimation model acquisition device according to claim 1 .

4. a query acquisition unit that acquires query information including a sound time series that is a time series indicating a sound and second difference information that indicates at least a part of differences between the sound time series and an estimated time series that is a time series that satisfies a predetermined estimation condition; an estimation unit that estimates a time series that satisfies the estimation condition from one or more time series, based on the query information acquired by the query acquisition unit, using the mathematical model acquired by a sound estimation model acquisition device, based on a first sound time series that is a time series indicating a first sound, a second sound time series that is a time series indicating a second sound, and first difference information that indicates at least a part of the difference between the first sound and the second sound, wherein the mathematical model estimates a time series that satisfies the estimation condition based on an input time series that is a time series indicating an input sound, and second difference information that is information that indicates at least a part of the difference between the input time series and a time series that satisfies the estimation condition, and A sound estimation device comprising:

5. a sound estimation model acquisition step of obtaining a mathematical model for estimating an estimated time series, which is a time series that satisfies a predetermined estimation condition, from one or more time series, based on a first sound time series that is a time series indicating a first sound, a second sound time series that is a time series indicating a second sound, and first difference information that indicates at least a part of the differences between the first sound and the second sound; and the mathematical model is a mathematical model that estimates the estimated time series based on an input time series that is a time series indicating an input sound and second difference information that is information indicating at least a part of a difference between the input time series and the estimated time series, the estimation condition is a condition that a difference between the difference with the input time series and the difference indicated by the second difference information is smaller than a predetermined difference. How to obtain a sound estimation model.

6. a query acquisition step of acquiring query information including second difference information indicating at least a part of a difference between a sound time series that is a time series indicating a sound and an estimated time series that is a time series that satisfies a predetermined estimation condition; an estimation step of estimating a time series that satisfies the estimation condition from one or more time series, based on the query information acquired in the query acquisition step, using the mathematical model acquired by the sound estimation model acquisition device; and a model acquisition unit that acquires a mathematical model that estimates a time series that satisfies the estimation condition from one or more time series, based on a first sound time series that is a time series indicating a first sound, a second sound time series that is a time series indicating a second sound, and first difference information that indicates at least a part of a difference between the first sound and the second sound, wherein the mathematical model estimates a time series that satisfies the estimation condition based on an input time series that is a time series indicating an input sound, and second difference information that is information that indicates at least a part of a difference between the input time series and a time series that satisfies the estimation condition, and the estimation condition is a condition that a difference between the input time series and the difference indicated by the second difference information is smaller than a predetermined difference. A sound estimation method having the following.

7. A program for causing a computer to function as the sound estimation model acquisition device according to any one of claims 1 to 3.

8. A program for causing a computer to function as the sound estimation device according to claim 4.

9. the model acquisition unit further obtains the mathematical model based on a first sound event class label indicating a classification result of the first sound time series and a second sound event class label indicating a classification result of the second sound time series. The sound estimation model acquisition device according to claim 1 .

10. the model acquisition unit further obtains the mathematical model based on a first sound event class label indicating a classification result of the first sound time series and a second sound event class label indicating a classification result of the second sound time series. The sound estimation device according to claim 4 .

Citation Information

Patent Citations

  • Effect sound retrieving device

    JP1997146580A

  • Audio signal retrieving device, audio signal retrieving method, data retrieving device, data retrieving method, and program

    WO2020241070A1