Signal processing device and method, learning device and method, and program
The learning device generates high-quality target sound signals by training an object sound source generator using sensor data from devices attached to objects, addressing the challenge of low signal-to-noise ratios and unintended interference in sound recordings without microphones, enhancing sound quality and device performance.
Patent Information
- Application Number
- JP2022557397
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-20
- Filing Date
- 2021-10-06
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2041-10-06
AI Technical Summary
Existing technologies face challenges in obtaining high-quality sound recordings of target sound sources, such as human voices and movement sounds, during events like sports and theater, due to the inability to attach microphones to the sound sources, leading to low signal-to-noise ratios and interference from unintended sounds.
A learning device and method that utilizes sensor signals from devices attached to objects, such as microphones, acceleration sensors, and GPS, to generate high-quality target sound signals by training an object sound source generator using machine learning, even without direct microphone recordings, by integrating motion and position information to separate and enhance desired sounds.
Enables the acquisition of high-quality target sound signals with improved signal-to-noise ratios and reduced physical burden on athletes or performers, extending device battery life and preventing unintended audio leakage during events.
Smart Images

Figure 0007754104000001 
Figure 0007754104000002 
Figure 0007754104000003
Abstract
Description
[Technical Field]
[0001] The present technology relates to a signal processing device and method, a learning device and method, and a program, and in particular to a signal processing device and method, a learning device and method, and a program that enable a high-quality target sound to be obtained. [Background technology]
[0002] When reproducing sound fields from any viewpoint, such as bird's-eye views or walk-throughs, it is important to record the sound of the target sound source with a high signal-to-noise ratio (SNR), and at the same time, it is necessary to obtain information indicating the position and direction of each sound source.
[0003] Specific examples of the target sound source include human voices, general human movement sounds such as walking and running sounds, and movement sounds specific to content such as sports and theater, such as the kicking of a ball.
[0004] Furthermore, for example, as a technology related to user behavior recognition, a technology has been proposed that enables the analysis of ranging sensor data detected by multiple ranging sensors to obtain behavior recognition results for one or more users (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2017-205213 Summary of the Invention [Problem to be solved by the invention]
[0006] However, when recording sports, theater, or other events as free viewpoint content, it can be difficult to obtain the desired sound source with a high signal-to-noise ratio, for example, because it is not possible to attach a device equipped with a microphone to the athletes who are the sound sources. In other words, it can be difficult to obtain high-quality desired sound.
[0007] The present technology has been made in view of such circumstances, and makes it possible to obtain a high-quality target sound. [Means for solving the problem]
[0008] A learning device according to a first aspect of the present technology includes a learning unit that performs learning based on one or more sensor signals obtained by one or more sensors attached to an object and a target signal related to the object that corresponds to a predetermined sensor, and generates coefficient data constituting a generator that receives the one or more sensor signals as input and outputs the target signal.
[0009] A learning method or program according to a first aspect of the present technology includes a step of performing learning based on one or more sensor signals obtained by one or more sensors attached to an object and a target signal related to the object that corresponds to a predetermined sensor, and generating coefficient data constituting a generator that receives the one or more sensor signals as input and outputs the target signal.
[0010] In a first aspect of the present technology, learning is performed based on one or more sensor signals obtained by one or more sensors attached to an object and a target signal related to the object that corresponds to a predetermined sensor, and coefficient data that constitutes a generator that takes the one or more sensor signals as input and outputs the target signal is generated.
[0011] A signal processing device according to a second aspect of the present technology includes an acquisition unit that acquires one or more sensor signals obtained by one or more sensors attached to an object, and a generation unit that generates a target signal for the object that corresponds to a predetermined sensor based on coefficient data constituting a generator that has been generated in advance by learning and the one or more sensor signals.
[0012] A signal processing method or program according to a second aspect of the present technology includes a step of acquiring one or more sensor signals obtained by one or more sensors attached to an object, and generating a target signal for the object corresponding to a predetermined sensor based on coefficient data constituting a generator generated in advance by learning and the one or more sensor signals.
[0013] In a second aspect of the present technology, one or more sensor signals obtained by one or more sensors attached to an object are acquired, and a target signal for the object corresponding to a specified sensor is generated based on coefficient data constituting a generator generated in advance by learning and the one or more sensor signals. [Brief explanation of the drawings]
[0014] [Figure 1] 10A and 10B are diagrams illustrating examples of device configurations during training and during content recording. [Figure 2] FIG. 10 illustrates an example of the configuration of an object sound source generator. [Figure 3] FIG. 2 is a diagram illustrating an example of the configuration of a recording device and a learning device. [Figure 4] 10 is a flowchart illustrating a recording process when generating learning data. [Figure 5] 10 is a flowchart illustrating a learning process. [Figure 6] FIG. 2 is a diagram illustrating an example of the configuration of a recording device and a sound source generating device. [Figure 7] 10 is a flowchart illustrating a recording process when an object sound source is generated. [Figure 8] 10 is a flowchart illustrating a sound source generation process. [Figure 9] FIG. 1 illustrates an example of the configuration of a learning device. [Figure 10] 10 is a flowchart illustrating a learning process. [Figure 11] FIG. 1 illustrates an example of the configuration of a learning device. [Figure 12] 10 is a flowchart illustrating a learning process. [Figure 13] FIG. 1 illustrates an example of the configuration of a sound source generating device. [Figure 14] 10 is a flowchart illustrating a sound source generation process. [Figure 15] 10 is a flowchart illustrating a learning process. [Figure 16] FIG. 1 illustrates an example of the configuration of a sound source generating device. [Figure 17] 10 is a flowchart illustrating a sound source generation process. [Figure 18] 10 is a flowchart illustrating a learning process. [Figure 19] FIG. 1 illustrates an example of the configuration of a sound source generating device. [Figure 20] 10 is a flowchart illustrating a sound source generation process. [Figure 21] FIG. 1 illustrates an example of the configuration of a learning device. [Figure 22] 10 is a flowchart illustrating a learning process. [Figure 23] FIG. 1 illustrates an example of the configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0016] First Embodiment About this technology This technology enables the acquisition of high-quality target signals by generating signals corresponding to sensor signals obtained from other sensors based on sensor signals obtained from one or more sensors.
[0017] The sensors referred to here include, for example, microphones, acceleration sensors, gyro sensors, geomagnetic sensors, distance measuring sensors, and image sensors.
[0018] In the following, an example will be described in which an object sound source signal corresponding to a microphone recorded signal obtained by a microphone when the microphone is attached to an object is generated as a target signal from sensor signals of one or more different types of sensors such as an acceleration sensor. Note that the target signal is not limited to an object sound source signal, and may be any signal such as a video signal of an animation or the like.
[0019] For example, few existing wearable devices have the ability to record voice and movement sounds in high quality while exercising. There are devices that combine a small, robust transmitter with a lavalier microphone, primarily used in the broadcasting industry. However, such devices do not have any sensors other than the microphone.
[0020] Furthermore, there are many wearable devices for motion analysis that are equipped with sensors that acquire positional and motion information during exercise, but these devices do not have the ability to acquire audio, or even if they do, they are not specialized for audio acquisition.
[0021] Therefore, there was no device that could simultaneously acquire audio, positional information, and motion information with high signal-to-noise ratio during exercise and use them to generate object sound sources.
[0022] Therefore, this technology makes it possible to generate an audio signal (acoustic signal) that reproduces the sound of the desired object sound source, i.e., an object sound source signal, from sensor signals such as acquired position information and movement information.
[0023] For example, assume that there are multiple objects in the same target space, and each of these objects has a recording device attached to or built in for recording content.
[0024] In this case, it is assumed that the sound emitted by an object to which a recording device is attached or built is recorded (collected) as the sound of the object sound source.
[0025] For example, the target space may be a space in which multiple athletes (players) or performers exist, such as in sports, opera, theater, or film.
[0026] Furthermore, for example, an object in the target space may be a moving object or a stationary object as long as it serves as a sound source (object sound source). More specifically, the object may be a person such as an athlete, or may be a robot, vehicle, drone, or other flying object equipped with or equipped with a recording device.
[0027] The recording device is equipped with, for example, a microphone for picking up the sound of the object sound source, a motion measurement sensor such as a 9-axis sensor for measuring the movement and direction (azimuth) of the object, a ranging sensor and a positioning sensor for measuring the position, and a camera (image sensor) for capturing images of the surrounding area.
[0028] Here, the ranging sensor (ranging device) or positioning sensor is, for example, a GPS (Global Positioning System) device for measuring the position of an object or a beacon receiver for indoor ranging, and the ranging sensor or positioning sensor can obtain location information indicating the position of the object.
[0029] In addition, the output of a motion measurement sensor provided in the recording device can provide motion information indicating the object's motion, such as speed and acceleration, the object's orientation (heading), and the object's movement.
[0030] The recording device uses a built-in microphone, motion measurement sensor, distance measurement sensor, and positioning sensor to obtain a microphone recording signal obtained by picking up sounds around the object, object position information, and object motion information. In addition, if the recording device is equipped with a camera, a video signal of the image around the object can also be obtained.
[0031] The microphone recording signal, position information, movement information, and video signal obtained for each object in this way can be used to obtain an object sound source signal, which is an acoustic signal of the sound of the object sound source, which is the target sound.
[0032] Here, the sound of the object sound source that is considered to be the target sound is, for example, the sound of the object (person) walking, running, breathing, clapping, and other movement sounds, and of course the sound of the object (person) speaking can also be considered to be the sound of the object sound source.
[0033] In such a case, it is conceivable to generate an object sound source signal of each sound source type for each object by, for example, detecting a time period in which a target sound such as the movement sound of each object exists using a microphone recorded signal, position information, and movement information, and performing signal processing to separate the target sound from the microphone recorded signal based on the detection result.Furthermore, it is conceivable to integrate and use position information and movement information obtained from multiple recording devices to generate an object sound source signal.
[0034] In this way, a high-quality signal can be obtained as an object sound source signal for free viewpoint playback.
[0035] However, in reality, when recording content, it may not be possible to acquire all types of sensor signals, such as microphone recording signals and acceleration sensor outputs.
[0036] Specifically, for example, there may be cases where the weight of the recording device is restricted so as not to interfere with the wearer's performance, or when the recorded audio is to be used over broadcast waves, a microphone cannot be used to prevent unintended audio, such as the sound of a strategy, from being broadcast.
[0037] Therefore, in this technology, by acquiring more sensor signals at a timing different from the time of recording the content than when the content is recorded and training the object sound source generator, it is possible to obtain object sound source signals that cannot be obtained when the content is recorded from the sensor signals obtained when the content is recorded.
[0038] A specific example would be a case where a sports match or a play is recorded as content.
[0039] In such cases, data is recorded by changing the device configuration of the recording device during training (practice) or rehearsals, including practice matches, and during the actual event, i.e., when recording content for a match or main performance.
[0040] As an example, a device configuration such as that shown in FIG. 1 can be used.
[0041] In the example of Figure 1, during training or rehearsal, the recording device is equipped with sensors for recording data, such as a microphone, an acceleration sensor, a gyro sensor, a geomagnetic sensor, and sensors for measuring position such as GPS and indoor ranging (ranging sensors and positioning sensors).
[0042] During training or rehearsals, the recording device can be heavy and the battery can be replaced during the training or rehearsal, so all sensors including microphones are installed in the recording device and all sensor signals that can be obtained are acquired. In other words, high-performance recording devices including microphones are attached to athletes or performers, and sensor signals are acquired (collected).
[0043] In this way, learning is performed based on the sensor signals obtained during training or rehearsal, and an object sound source generator is generated.
[0044] In contrast, during matches and performances, the recording devices are equipped with sensors for data recording, such as acceleration sensors, gyro sensors, geomagnetic sensors, and position measurement sensors such as GPS and indoor distance measurement. In other words, no microphones are installed during matches or performances.
[0045] During a match or performance, only a portion of the sensors that were installed during training or rehearsals are installed in the recording device, making the recording device lighter and increasing battery life. In other words, during a match or performance, the types and number of sensors installed are narrowed down, and sensor signals are acquired by a lightweight, low-power recording device.
[0046] In particular, it is assumed that batteries cannot be replaced during a match or a performance, and the battery is made to last from the start to the end of the match or performance. Also, in this example, microphones are not installed during a match or a performance, so it is possible to prevent the recording and leaking of inappropriate speech during live broadcasts, such as speech by players regarding strategy. In addition, in some sports, wearing microphones may be prohibited, but even in such cases, it is possible to obtain sensor signals from sensors other than microphones.
[0047] When sensor signals from sensors of a type other than microphones are obtained during a match or a performance, an object sound source signal is generated based on those sensor signals and an object sound source generator obtained through prior learning. This object sound source signal is an object sound source signal corresponding to a microphone recorded signal that would have been obtained if a microphone had been installed in the recording device during the match or performance, i.e., if a microphone had been attached to the object. The generated object sound source signal may be equivalent to the microphone recorded signal itself, or may be equivalent to (correspond to) a signal of the object sound source generated from the microphone recorded signal, or may be a signal for reproducing a sound equivalent to the sound of the object sound source generated from the microphone recorded signal.
[0048] When training the object sound source generator, sensor signals recorded (acquired) during training or rehearsals are utilized to obtain (learn) prior information such as the personalities of the athletes or performers wearing the recording devices and the environment in which the sensor signals were acquired.
[0049] Then, when the content is recorded, i.e., during a match or the actual performance, object sound source signals such as the movement sounds of objects such as players are estimated (restored) based on prior information such as individuality (object sound source generator) and the few sensor signals obtained during content recording.
[0050] By doing this, it is possible to reduce the physical burden on athletes and performers when recording content, extend the operating time of the recording device, and obtain a high-quality signal as an object sound source signal for the sound of the desired object sound source (target sound), such as the sound of an athlete's movements.
[0051] Furthermore, with this technology, even when a microphone cannot be installed in the recording device during content recording, when signals from the microphone or other sensors cannot be obtained due to a malfunction or the like, or when sensor signals with a high S / N ratio cannot be obtained due to a poor recording environment or the like, it is possible to robustly obtain high-quality object sound source signals.
[0052] The object sound generator will now be further explained.
[0053] The object sound source generator may be any type that can obtain an object sound source signal corresponding to a desired sensor signal from one or more sensor signals, such as one that uses a DNN (Deep Neural Network) or a parametric representation method.
[0054] For example, in a parametric representation object sound generator, the sound source type of the target object sound source is limited, and then the sound source type of the object sound source, such as a kick sound or a running sound, that is likely to be emitted is identified from one or more sensor signals.
[0055] Then, based on the sound source type identification result and the envelope of the signal strength (amplitude) of the sensor signal from the acceleration sensor, parametric signal processing is performed on the dry source waveform signal of the object sound source prepared in advance, and an object sound source signal is generated.
[0056] Such an object sound source generator can be configured as shown in FIG. 2, for example.
[0057] In the example shown in Figure 2, the object sound source generator 11 receives sensor signals from multiple sensors of different types, and if a microphone is provided in the recording device attached to the object, outputs an object sound source signal corresponding to the microphone recording signal obtained by the microphone.
[0058] The object sound source generator 11 includes a sound source type identifier 21, a sound source database 22, an envelope estimation unit 23, and a mixer 24.
[0059] Sensor signals acquired by sensors other than microphones provided in recording devices attached to objects such as players, which are object sound sources, are supplied to the sound source type classifier 21 and the envelope estimation unit 23. Here, it is assumed that the sensor signals from sensors other than microphones include at least sensor signals acquired by acceleration sensors.
[0060] In addition, if a microphone recording signal obtained by a microphone provided on a recording device attached to the object or a microphone provided on a recording device attached to another object is available, the obtained microphone recording signal is also supplied to the sound source type identifier 21 and the envelope estimation unit 23.
[0061] When one or more sensor signals including a microphone recorded signal are supplied to the sound source type classifier 21, the sound source type classifier 21 estimates the type of the object sound source, that is, the type of sound emitted from the object, by performing a calculation based on the supplied sensor signals and coefficients obtained in advance by learning. The sound source type classifier 21 supplies sound source type information indicating the estimation result (classification result) of the sound source type to the sound source database 22.
[0062] For example, the sound source type classifier 21 is a classifier using any machine learning algorithm such as a generalized linear discriminant or an SVM (Support Vector Machine).
[0063] The sound source database 22 stores, for each sound source type, dry source signals that are acoustic signals of dry source waveforms that reproduce the sound of an object emitted from an object sound source. That is, the sound source database 22 stores sound source type information and the dry source signals of the sound source type indicated by the sound source type information in association with each other.
[0064] For example, the dry source signal may be any signal obtained by actually recording the sound of an object sound source in an environment with a high S / N ratio, an artificially generated signal, a signal generated by extraction from a microphone signal recorded by a microphone of a recording device for learning purposes, etc. Furthermore, multiple dry source signals may be stored in association with one piece of sound source type information.
[0065] The sound source database 22 supplies one or more dry source signals of the sound source type indicated by the sound source type information supplied from the sound source type identifier 21 to the mixer 24, among the dry source signals of each sound source type held therein.
[0066] The envelope estimation unit 23 extracts the envelope of the sensor signal of the supplied acceleration sensor, and supplies envelope information indicating the extracted envelope to the mixer 24 .
[0067] For example, the amplitude of a predetermined component in a time interval of a predetermined length of the sensor signal of the acceleration sensor is time-smoothed and extracted as an envelope in the envelope estimation unit 23. The predetermined component here refers to one or more components including at least a component in the direction of gravity, for example, when generating an object sound source signal of walking sound as the sound of the object sound source.
[0068] Furthermore, when extracting the envelope, a microphone recording signal or a sensor signal from a sensor other than the acceleration sensor and microphone may also be used as needed.
[0069] The mixer 24 generates and outputs an object sound source signal by performing signal processing on one or more dry source signals supplied from the sound source database 22 based on the envelope information supplied from the envelope estimator 23.
[0070] For example, the signal processing performed by the mixing unit 24 may include modulation of the waveform of the dry source signal based on the signal strength of the envelope indicated by the envelope information, filtering of the dry source signal based on the envelope information, and mixing processing for mixing multiple dry source signals.
[0071] It should be noted that the method for generating an object sound source signal in the object sound source generator 11 is not limited to the example described here. For example, any type of classifier may be used to identify the type of sound source, any type of envelope estimation method, any type of method for generating an object sound source signal from a dry source signal, and the like. The configuration of the object sound source generator is not limited to the example shown in Fig. 2 and may be another configuration.
[0072] <Example of recording device and learning device configuration> Next, a more detailed embodiment of the present technology described above will be described.
[0073] First, we will explain an example of generating an object sound source signal when recording a performance (main performance) such as a sports match or a play as content, when a microphone cannot be used or the recording device does not have a microphone.
[0074] In this example, an object sound source generator 11 that generates an object sound source signal from a sensor signal of a type of sensor other than a microphone is generated by machine learning.
[0075] The recording device that acquires the sensor signals for training the object sound source generator 11 and the learning device that trains the object sound source generator 11 are configured as shown in Fig. 3. Although only one recording device is shown here, there may be one recording device or multiple recording devices for each object.
[0076] In the example of FIG. 3, a sensor signal acquired by a recording device 51 attached to or built into an object such as a moving object is transmitted as transmission data to a learning device 52 and used for learning of the object sound source generator 11.
[0077] The recording device 51 includes a microphone 61 , a movement measurement unit 62 , a position measurement unit 63 , a recording unit 64 , and a transmission unit 65 .
[0078] The microphone 61 picks up sounds around the recording device 51 and supplies a microphone recording signal, which is the resulting sensor signal, to the recording unit 64. The microphone recording signal may be a monaural signal or a multi-channel signal.
[0079] The movement measurement unit 62 consists of sensors for measuring the movement and orientation of an object, such as a 9-axis sensor, a geomagnetic sensor, an acceleration sensor, or a gyro sensor, and outputs a sensor signal indicating the measurement result (sensing value) to the recording unit 64 as movement information.
[0080] In particular, the motion measurement unit 62 measures the motion and orientation of the object while the microphone 61 is collecting sound, and outputs motion information indicating the results. The motion information may be a single sensor signal, or may be sensor signals from multiple sensors of different types.
[0081] The position measurement unit 63 consists of a ranging sensor or a positioning sensor, such as a GPS device or an indoor ranging beacon receiver, and measures the position of the object to which the recording device 51 is attached, and outputs a sensor signal indicating the measurement result to the recording unit 64 as position information.
[0082] The position information may be a single sensor signal or may be sensor signals from multiple sensors of different types. The position indicated by the position information, more specifically, the position of the object (recording device 51) determined from the position information, may be coordinate information based on a predetermined position in the space where recording is performed, such as a game venue or a theater.
[0083] The microphone signal, the motion information, and the position information are acquired simultaneously during the same period.
[0084] The recording unit 64 performs AD (Analog to Digital) conversion, etc., as appropriate on the microphone recording signal supplied from the microphone 61, the movement information supplied from the movement measurement unit 62, and the position information supplied from the position measurement unit 63, and supplies them to the transmission unit 65.
[0085] The transmission unit 65 performs compression processing on the microphone recording signal, movement information, and position information supplied from the recording unit 64 to generate transmission data containing the microphone recording signal, movement information, and position information, and transmits the transmission data to the learning device 52 via a network or the like.
[0086] Alternatively, an image sensor or the like may be provided in the recording device 51, and a video signal or the like may be included in the transmission data as a sensor signal used for learning. Furthermore, the learning data may be a plurality of sensor signals obtained by a plurality of sensors of different types, or may be a single sensor signal obtained by a single sensor.
[0087] The learning device 52 includes an acquisition unit 81 , a sensor database 82 , a microphone recording database 83 , a learning unit 84 , and a coefficient database 85 .
[0088] The acquisition unit 81 acquires transmission data, for example by receiving transmission data sent from the recording device 51, and performs decoding processing on the acquired transmission data as appropriate to obtain a microphone recording signal, movement information, and position information.
[0089] The acquisition unit 81 performs signal processing to extract the sound of the target object sound source from the microphone recorded signal as needed, and supplies the microphone recorded signal to be used as training data (correct data) for learning to the microphone recorded database 83. This training data is the object sound source signal to be generated by the object sound source generator 11.
[0090] The acquisition unit 81 also supplies the motion information and position information extracted from the transmitted data to the sensor database 82 .
[0091] The sensor database 82 records the motion information and position information supplied from the acquisition unit 81, and also supplies the information to the learning unit 84 as learning data as appropriate.
[0092] The microphone recording database 83 records the microphone recording signals supplied from the acquisition unit 81, and also supplies the recorded microphone recording signals to the learning unit 84 as training data as appropriate.
[0093] The learning unit 84 performs machine learning based on the motion information and position information supplied from the sensor database 82 and the microphone recorded signals supplied from the microphone recorded database 83, generates the object sound source generator 11, more specifically, coefficient data constituting the object sound source generator 11, and supplies the coefficient data to a coefficient database 85. The coefficient database 85 records the coefficient data supplied from the learning unit 84.
[0094] The object sound source generator generated by machine learning may be the object sound source generator 11 shown in Figure 2, or an object sound source generator with another configuration such as DNN, but in the following explanation we will assume that the object sound source generator 11 is generated.
[0095] In such a case, coefficient data consisting of coefficients used in the arithmetic processing (signal processing) in the sound source type identifier 21, envelope estimation unit 23, and mixer 24 is generated by machine learning and supplied to the coefficient database 85.
[0096] The training data during learning may be the microphone recorded signal itself obtained by the microphone 61, but it is also a microphone recorded signal (acoustic signal) that contains only the sound of the target object sound source, with unintended sounds such as inappropriate voices, that is, sounds that are not necessary as content, removed from the microphone recorded signal.
[0097] Specifically, in the case of soccer content, for example, sounds such as rustling of clothes and collision noises are removed as unnecessary sounds, and sounds such as ball kicks, voices, breathing sounds, and running sounds are extracted as sounds of the desired sound source type, that is, object sound sources, and the acoustic signals of the extracted sounds are used as training data.
[0098] When generating microphone-recorded signals that serve as training data for each sound source type, the desired and unwanted sounds will vary depending on the situation, such as for each piece of content or the country to which the content is distributed, and should be determined appropriately.
[0099] The process of removing unnecessary sounds from the microphone recording signal may be realized by any process, such as sound source separation using a DNN or other sound source separation. Furthermore, the process of removing unnecessary sounds from the microphone recording signal may use motion information, position information, microphone recording signals obtained by other recording devices 51, etc.
[0100] In addition, the training data is not limited to microphone-recorded signals or audio signals generated from microphone-recorded signals, but may be any audio signal (target signal) of a desired sound related to an object (object sound source) that corresponds to a microphone-recorded signal, such as a general kick sound generated in advance.
[0101] Furthermore, the learning device 52 may record learning data and teacher data for each individual, such as each player or performer who will be an object, and generate coefficient data for the object sound source generator 11 taking into consideration the individuality of each individual (person wearing the recording device 51). Conversely, learning data and teacher data obtained for multiple objects may be used to generate coefficient data for general-purpose object sound source generator 11.
[0102] Furthermore, in cases where the object sound source generator is configured using a DNN, a DNN may be prepared for each object sound source, or different sensor signals may be input to each DNN.
[0103] <Explanation of the recording process when generating learning data> Next, the operations of the recording device 51 and the learning device 52 shown in FIG. 3 will be described.
[0104] First, the recording process performed by the recording device 51 when generating learning data will be described with reference to the flowchart of FIG.
[0105] In step S11 , the microphone 61 picks up sounds around the recording device 51 and supplies the resulting microphone recording signal to the recording unit 64 .
[0106] In step S12, the recording unit 64 acquires the sensor signals output from the movement measuring unit 62 and the position measuring unit 63 as movement information and position information.
[0107] The recording unit 64 performs AD conversion, etc. as necessary, on the microphone recorded signal, movement information, and position information acquired through the above processing, and supplies the obtained microphone recorded signal, movement information, and position information to the transmission unit 65.
[0108] The transmitting unit 65 also generates transmission data consisting of the microphone recorded signal, the movement information, and the position information supplied from the recording unit 64. At this time, the transmitting unit 65 performs compression processing on the microphone recorded signal, the movement information, and the position information as necessary.
[0109] In step S13, the transmission unit 65 transmits the transmission data to the learning device 52, and the recording process ends. The transmission data may be transmitted sequentially in real time (online), or all the transmission data may be transmitted collectively offline after recording.
[0110] In this way, the recording device 51 acquires not only the microphone recorded signal but also the motion information and position information, and transmits them to the learning device 52. In this way, the learning device 52 can obtain an object sound source generator for obtaining the microphone recorded signal from the motion information and position information, and as a result, it becomes possible to obtain a high-quality target sound.
[0111] <Explanation of learning process> Next, the learning process performed by the learning device 52 will be described with reference to the flowchart of FIG.
[0112] In step S41, the acquisition unit 81 acquires the transmission data by receiving the transmission data transmitted from the recording device 51. Note that, if necessary, the acquired transmission data may be decrypted. Also, the transmission data may be acquired from a removable recording medium or the like, rather than being received via a network.
[0113] In step S42, the acquisition unit 81 performs labeling on the acquired transmission data.
[0114] For example, the acquisition unit 81 performs a labeling process in which, for each time interval of the microphone recording signal, movement information, and position information that make up the transmission data, it associates sound source type information that indicates which sound source type of object sound source each of those time intervals is a time interval for the sound of that object sound source.
[0115] The time interval corresponding to each sound source type may be manually input by a user, or sound source type information may be obtained as a result of signal processing such as sound source separation based on microphone recording signals, motion information, and position information.
[0116] The acquisition unit 81 supplies the labeled microphone recorded signals, i.e., the signals associated with sound source type information, to a microphone recorded database 83 for recording, and supplies each sensor signal constituting the labeled movement information and position information to a sensor database 82 for recording.
[0117] During learning, the sensor signals constituting the motion information and position information thus obtained are used as learning data, and the microphone-recorded signals are used as teacher data.
[0118] As described above, the training data may be a signal obtained by removing unnecessary sounds from a microphone-recorded signal. In such a case, for example, the acquisition unit 81 may input the microphone-recorded signal, motion information, and position information to a DNN previously obtained by learning, perform arithmetic processing, and obtain the microphone-recorded signal as training data as an output.
[0119] Furthermore, when a microphone recorded signal can be obtained as transmission data when using the object sound source generator 11, the microphone recorded signal, movement information, and position information that make up the transmission data can be used as learning data, and a signal obtained by removing unnecessary sounds from the microphone recorded signal can be used as training data.
[0120] In the learning device 52, a large amount of learning data and teacher data obtained at a plurality of different times are recorded in a sensor database 82 and a microphone recording database 83.
[0121] Furthermore, the learning data and the teacher data may be obtained from the same recording device 51, that is, for the same object, or may be obtained from a plurality of different recording devices 51 (objects).
[0122] Furthermore, as learning data corresponding to the teacher data obtained by a specified recording device 51, not only the motion information and position information obtained by that specified recording device 51, but also the motion information and position information obtained by other recording devices 51, and video signals obtained by a camera or the like other than the recording device 51 may be used.
[0123] In step S43, the learning unit 84 performs machine learning based on the learning data recorded in the sensor database 82 and the teacher data recorded in the microphone recording database 83, and generates coefficient data that constitutes the object sound source generator 11.
[0124] The learning unit 84 supplies the coefficient data thus obtained to the coefficient database 85 for recording, and the learning process ends.
[0125] In this way, the learning device 52 performs machine learning using the transmission data acquired from the recording device 51 and generates coefficient data. In this way, it becomes possible to use the object sound source generator 11 to obtain high-quality target sound even in a situation where a microphone-recorded signal cannot be obtained.
[0126] <Example of configuration of recording device and sound source generation device> Next, an example of the configuration of a recording device and sound source generation device for recording actual content during a match or a performance and generating an object sound source signal of the desired object sound source sound from the recording results is shown in Figure 6. Note that in Figure 6, parts that correspond to those in Figure 3 are given the same reference numerals, and their explanation will be omitted where appropriate.
[0127] In FIG. 6, a sensor signal acquired by a recording device 111 attached to or built into an object such as a moving body is transmitted as transmission data to a sound source generating apparatus 112, and an object sound source generator 11 generates an object sound source signal.
[0128] The recording device 111 includes a movement measurement unit 62 , a position measurement unit 63 , a recording unit 64 , and a transmission unit 65 .
[0129] The configuration of the recording device 111 differs from the configuration of the recording device 51 in that the recording device 111 does not include a microphone 61, but is otherwise the same as the recording device 51.
[0130] Therefore, the transmission data transmitted (output) by the transmission unit 65 of the recording device 111 consists of movement information and position information, and does not include a microphone recording signal. Note that the recording device 111 may be provided with an image sensor or the like, and a video signal or the like may be included in the transmission data as a sensor signal.
[0131] The sound source generating device 112 includes an acquisition unit 131 , a coefficient database 132 , and an object sound source generating unit 133 .
[0132] The acquisition unit 131 acquires coefficient data from the learning device 52 via a network or the like, and supplies and records it in the coefficient database 132. Note that the coefficient data may be acquired not only from the learning device 52 but also from another device via a wired or wireless connection, or from a removable recording medium or the like.
[0133] The acquisition unit 131 also receives transmission data sent from the recording device 111 and supplies the data to the object sound source generation unit 133. The acquisition unit 131 also performs decoding processing and the like on the received transmission data as appropriate.
[0134] The transmission data may be acquired from the recording device 111 in real time (online) or may be acquired offline after recording. The transmission data may be received directly from the recording device 111 or may be acquired from a removable recording medium or the like.
[0135] The object sound source generator 133 performs calculations based on the coefficient data supplied from the coefficient database 132, thereby functioning as the object sound source generator 11 shown in FIG.
[0136] That is, the object sound source generating unit 133 generates an object sound source signal based on the coefficient data supplied from the coefficient database 132 and the transmission data supplied from the acquiring unit 131, and outputs the generated signal to the subsequent stage.
[0137] <Explanation of recording process when generating object sound source> Next, the operation of the recording device 111 and the sound source generating device 112 when recording content, that is, when generating an object sound source signal, will be described.
[0138] First, the recording process performed by the recording device 111 when generating an object sound source will be described with reference to the flowchart of FIG.
[0139] When the recording process starts, motion information and position information are acquired in step S71, and transmission data is sent in step S72. These processes are similar to the processes in steps S12 and S13 in FIG. 4, so their description will be omitted.
[0140] However, the transmission data sent in step S72 includes only the motion information and position information, and does not include the microphone recorded signal.
[0141] Once the transmission data has been sent in this manner, the recording process ends.
[0142] In this way, the recording device 111 acquires the motion information and position information and transmits them to the sound source generation device 112. In this way, the sound source generation device 112 can obtain an object sound source signal corresponding to the microphone recorded signal from the motion information and position information. In other words, it becomes possible to obtain a high-quality target sound.
[0143] <Explanation of sound source generation process> Next, the sound source generation process performed by the sound source generation device 112 will be described with reference to the flowchart of FIG.
[0144] It is assumed that the coefficient data has already been acquired and recorded in the coefficient database 132 when the sound source generation process is started.
[0145] In step S101, the acquisition unit 131 receives transmission data transmitted from the recording device 111, acquires the transmission data, and supplies the data to the object sound source generation unit 133. Note that, if necessary, the acquired transmission data is also decoded.
[0146] In step S 102 , the object sound source generating unit 133 generates an object sound source signal based on the coefficient data supplied from the coefficient database 132 and the transmission data supplied from the obtaining unit 131 .
[0147] For example, the object sound source generator 133 functions as the object sound source generator 11 obtained in advance by learning using coefficient data from the coefficient database 132 .
[0148] In such a case, the sound source type identifier 21 performs calculations based on multiple sensor signals as motion information and position information constituting the transmission data supplied from the acquisition unit 131 and coefficients obtained in advance by learning, and supplies the resulting sound source type information to the sound source database 22.
[0149] The sound source database 22 supplies one or more dry source signals of the sound source type indicated by the sound source type information supplied from the sound source type identifier 21 to the mixer 24, among the dry source signals of each sound source type held therein.
[0150] The envelope estimation unit 23 extracts the envelope of the sensor signal of the acceleration sensor contained in the motion information constituting the transmission data supplied from the acquisition unit 131, and supplies envelope information indicating the extracted envelope to the mixer 24.
[0151] The mixing unit 24 generates an object sound source signal by performing signal processing such as modulation and filtering on one or more dry source signals supplied from the sound source database 22 based on the envelope information supplied from the envelope estimation unit 23.
[0152] The object sound source generating unit 133 may generate metadata for each object sound source signal at the same time as generating the object sound source signal for each object sound source.
[0153] In such a case, for example, the object sound source generation unit 133 generates metadata based on the transmission data, including sound source type information obtained by the sound source type identifier 21, and motion information and position information acquired as the transmission data.
[0154] In step S103, the object sound source generating unit 133 outputs the generated object sound source signal, and the sound source generation process ends. At this time, metadata may also be output together with the object sound source signal.
[0155] In this way, the sound source generator 112 generates and outputs an object sound source signal based on the coefficient data and transmission data.
[0156] In this way, even if a microphone recording signal cannot be obtained by the recording device 111, a high-quality target sound, that is, a high-quality object sound source signal can be obtained from a sensor signal of a sensor other than the microphone.
[0157] In other words, it is possible to obtain a robust object sound source signal of a target object sound source from sensor signals of a small number of types of sensors. Moreover, it is possible to obtain an object sound source signal that contains only the target sound, without including sounds such as noise and voice that are not desired to be reproduced.
[0158] Second Embodiment <Example of learning device configuration> As described above, the training data used for training the object sound source generator 11 may be a signal obtained by removing unnecessary sounds from a microphone-recorded signal, i.e., a signal obtained by extracting only the target sound.Moreover, in addition to the microphone-recorded signal, movement information and position information may also be used to generate the training data.
[0159] When motion information and position information are also used to generate training data, the learning device may be configured as shown in Fig. 9. In Fig. 9, parts corresponding to those in Fig. 3 are given the same reference numerals, and their explanation will be omitted as appropriate.
[0160] The learning device 161 shown in FIG. 9 includes an acquisition unit 81, a section detection unit 171, an integration unit 172, a signal processing unit 173, a sensor database 82, a microphone recording database 83, a learning unit 84, and a coefficient database 85.
[0161] The configuration of the learning device 161 differs from that of the learning device 52 in that it is newly provided with a section detection unit 171 through a signal processing unit 173, but in other respects it has the same configuration as the learning device 52.
[0162] When the acquisition unit 81 acquires the transmission data, it supplies the microphone recorded signals that make up the transmission data to the section detection unit 171, the integration unit 172, the signal processing unit 173, and the microphone recorded signals database 83.
[0163] In addition, the acquisition unit 81 supplies the movement information that constitutes the transmission data to the section detection unit 171, the integration unit 172, the signal processing unit 173, and the sensor database 82, and supplies the position information that constitutes the transmission data to the integration unit 172, the signal processing unit 173, and the sensor database 82.
[0164] The section detection unit 171 detects the type of sound of the object sound source contained in the microphone recorded signal and the time section in which the sound of the object sound source is contained, based on the microphone recorded signal and the movement information supplied from the acquisition unit 81, and supplies sound source type section information indicating the detection result to the integration unit 172.
[0165] For example, the section detection unit 171 performs threshold processing on the microphone recorded signal, performs calculations by substituting the microphone recorded signal and motion information into a classifier such as a DNN, or performs DS (Delay and Sum beamforming) or NBF (Null Beamformer) on the microphone recorded signal to identify the type of object sound source of the sound contained in each time section and generate sound source type section information.
[0166] As a specific example, if a sensor signal indicating a minute vertical displacement of an object measured by an acceleration sensor is used as motion information, the time period of the breathing sound of a person, which is the object, can be detected. In this case, for example, the time period in which the frequency of the sensor signal is about 0.5 Hz to 1 Hz is taken as the time period of the breathing sound of the object.
[0167] Alternatively, for example, by using NBF, it is possible to suppress the speech component of an object contained in a microphone-recorded signal. In this case, among the time intervals of the speech of the object detected in the microphone-recorded signal before suppression, the time intervals in which the speech of the object is not detected in the microphone-recorded signal after suppression are determined to be the final time intervals of the speech of the object.
[0168] The integration unit 172 generates final sound source type section information and sound source type information based on the microphone recording signal, movement information, and position information supplied from the acquisition unit 81 and the sound source type section information supplied from the section detection unit 171, and supplies them to the signal processing unit 173.
[0169] In particular, the integration unit 172 generates sound source type section information for the object to be processed (hereinafter also referred to as the target object) based on the microphone recording signal, movement information, and position information of the object and at least one of the microphone recording signals, movement information, and position information of other objects.
[0170] In this case, the integration unit 172 performs, for example, position information comparison processing, time interval integration processing, and interval smoothing processing to generate final sound source type interval information.
[0171] In the position information comparison process, the distance from the target object to other objects is calculated based on the position information of each object, and other objects that may have an effect on the sound of the object sound source of the target object, i.e., other objects near the target object, are selected as reference objects based on the calculated distance.
[0172] Next, in the time interval integration process, it is determined whether or not there is an object selected as a reference object.
[0173] Then, if no object is selected as the reference object, the sound source type section information of the target object obtained by the section detection unit 171 is output as is as the final sound source type section information to the signal processing unit 173. This is because, if there are no other objects near the target object, the sounds of other objects are not mixed into the microphone recorded signal.
[0174] On the other hand, if there are objects selected as reference objects, the position information and movement information of these reference objects are also used to generate the final sound source type section information of the target object.
[0175] Specifically, among the reference objects, a reference object having a time interval of the sound of the object sound source that overlaps with the time interval indicated by the sound source type interval information of the target object is selected as the final reference object.
[0176] Then, based on the position information and movement information of the reference object and the position information and movement information of the target object, relative orientation information indicating the relative direction of the reference object as seen from the target object in three-dimensional space is generated.
[0177] Furthermore, an NBF filter is formed based on the position information of the target object, the orientation of the target object indicated by the motion information, and the relative orientation information of each reference object. Furthermore, a convolution process is performed between the NBF filter and the time interval indicated by the sound source type interval information of the target object in the microphone recording signal of the target object.
[0178] Thereafter, based on the signal obtained by the convolution process and the motion information of the target object, sound source type section information is generated by the same process as that performed in the section detection unit 171. In this way, the sound emitted from the reference object is suppressed, and more accurate sound source type section information can be obtained.
[0179] Finally, the sound source type section information obtained by the time section integration process is subjected to section smoothing process, thereby obtaining final sound source type section information.
[0180] For example, for each type of object sound source, the average minimum duration of the sound when the sound of the object sound source of that type is generated is obtained in advance as the average minimum duration.
[0181] In the section smoothing process, smoothing is performed using a smoothing filter that connects the time sections of the sound of the object sound source that have been fragmented (divided) so that the length of the time section in which the sound of the object sound source is detected is equal to or longer than the average minimum duration.
[0182] Furthermore, the integration unit 172 generates sound source type information from the obtained sound source type section information, and supplies the sound source type section information and the sound source type information to the signal processing unit 173.
[0183] In this way, the integration unit 172 can remove information about the sounds of other objects that could not be completely removed (excluded) from the sound source type section information obtained by the section detection unit 171, thereby obtaining more accurate sound source type section information.
[0184] The signal processing unit 173 performs signal processing based on the microphone recording signal, movement information, and position information supplied from the acquisition unit 81 and the sound source type section information supplied from the integration unit 172, thereby generating an object sound source signal that serves as training data.
[0185] For example, the signal processing unit 173 performs signal processing such as sound quality correction processing, sound source separation processing, noise removal processing, distance correction processing, sound source replacement processing, and a combination of two or more of these processing.
[0186] Specifically, sound quality correction processing includes noise suppression processing such as filter processing and gain correction that suppress frequency bands where noise is dominant, processing that mutes noisy or unnecessary sections, and filter processing that increases frequency components that are easily attenuated.
[0187] For example, based on the sound source type section information, a time section in which sounds from multiple object sound sources are included in the microphone recorded signal is identified, and based on the identification results, a process based on independent component analysis is performed as sound source separation processing to separate the sounds from each object sound source according to differences in amplitude values and probability density distributions for each type of object sound source.
[0188] Furthermore, for example, in a microphone-recorded signal, when unwanted sounds such as background noise, steady noises such as cheers, or noises from wind are included in the time interval of the sound of the object sound source, a noise removal process is performed to suppress the noise for that time interval.
[0189] Furthermore, distance correction processing is performed to correct the influence of the distance attenuation and transfer characteristics from the object sound source to the position of the microphone 61 during recording on the absolute sound pressure of the sound emitted by the object sound source.
[0190] In this case, for example, in the distance correction process, the inverse characteristics of the transfer characteristics from the object sound source to the microphone 61 are added to the microphone recorded signal. This makes it possible to correct the deterioration in sound quality of the object sound source due to distance attenuation, transfer characteristics, etc.
[0191] In addition, a sound source replacement process is performed in which a part of a section of a microphone recorded signal is replaced with another acoustic signal that is prepared in advance or dynamically generated based on sound source type section information.
[0192] The signal processing unit 173 associates the object sound source signals, which are teacher data obtained as described above, with the sound source type information supplied from the integration unit 172, and supplies them to the microphone recording database 83.
[0193] Furthermore, the signal processing unit 173 supplies the sound source type information supplied from the integration unit 172 to the sensor database 82, and records the sound source type information in association with the motion information and position information.
[0194] In the learning device 161, sound source type information is generated and the sound source type information is associated with the object sound source signal that serves as teacher data, and the motion information and position information that serve as learning data, so there is no need for a user or the like to perform any operational input for the above-mentioned labeling process.
[0195] Furthermore, the microphone recording database 83 can supply the object sound source signal supplied from the signal processing unit 173 as training data to the learning unit 84, and can also supply the microphone recording signal supplied from the acquisition unit 81 as training data to the learning unit 84. Furthermore, the microphone recording signal may be used as training data.
[0196] <Explanation of learning process> Next, the learning process performed by the learning device 161 will be described with reference to the flowchart of FIG.
[0197] In step S131, the acquisition unit 81 acquires transmission data from the recording device 111.
[0198] The acquisition unit 81 supplies the microphone recorded signals that make up the transmission data to the section detection unit 171 , the integration unit 172 , the signal processing unit 173 , and the microphone recorded signals database 83 .
[0199] In addition, the acquisition unit 81 supplies the movement information that constitutes the transmission data to the section detection unit 171, the integration unit 172, the signal processing unit 173, and the sensor database 82, and supplies the position information that constitutes the transmission data to the integration unit 172, the signal processing unit 173, and the sensor database 82.
[0200] In step S132 , the section detection unit 171 generates sound source type section information based on the microphone recorded signal and the motion information supplied from the acquisition unit 81 , and supplies the information to the integration unit 172 .
[0201] For example, the section detection unit 171 performs threshold processing on the microphone recording signal or calculation processing based on a classifier such as a DNN to identify the type of object sound source of the sound included in each time section and generate sound source type section information.
[0202] In step S133, the integration unit 172 integrates information based on the microphone recording signal, movement information, and position information supplied from the acquisition unit 81 and the sound source type section information supplied from the section detection unit 171 to generate final sound source type section information and sound source type information, and supplies them to the signal processing unit 173.
[0203] For example, in step S133, the above-described position information comparison process, time interval integration process, and interval smoothing process are performed to integrate the information, and final sound source type interval information is generated.
[0204] In step S134, the signal processing unit 173 performs signal processing based on the microphone recording signal, movement information, and position information supplied from the acquisition unit 81 and the sound source type section information supplied from the integration unit 172, and generates an object sound source signal that serves as training data.
[0205] For example, the signal processing unit 173 performs signal processing such as sound quality correction processing, sound source separation processing, noise removal processing, distance correction processing, sound source replacement processing, or a combination of two or more of these processes, to generate an object sound source signal.
[0206] The signal processing unit 173 associates the object sound source signal, which is the teacher data, with the sound source type information supplied from the integration unit 172, and supplies the associated information to the microphone recording database 83 for recording. The signal processing unit 173 also supplies the sound source type information supplied from the integration unit 172 to the sensor database 82, and stores the sound source type information in association with the motion information and position information.
[0207] In step S135, the learning unit 84 performs machine learning based on the motion information and position information as learning data recorded in the sensor database 82 and the object sound source signal as teacher data recorded in the microphone recording database 83.
[0208] The learning unit 84 supplies the coefficient data obtained by machine learning to the coefficient database 85 for recording, and the learning process ends.
[0209] In this way, the learning device 161 generates training data using the transmission data acquired from the recording device 51, performs machine learning, and generates coefficient data. In this way, it becomes possible to use the object sound source generator 11 to obtain high-quality target sound even in a situation where a microphone-recorded signal cannot be obtained.
[0210] Third Embodiment <Example of learning device configuration> Incidentally, when obtaining an object sound source signal of the object sound source of the target object, it may be possible to obtain a microphone-recorded signal of the sound of the object sound source, although it does not have a high S / N ratio.
[0211] Such cases include cases where a microphone recorded signal can be obtained from a recording device 111 attached to an object other than the target object, such as a player or a referee other than the target player, or cases where sound is recorded by a separate microphone synchronized with the recording device 111. In addition, there may be cases where the microphone of the recording device 111 attached to the target object, which has a high priority, breaks down and a microphone recorded signal cannot be obtained.
[0212] Therefore, an object sound source signal of the object sound source intended for the target object may be generated by also using a microphone recording signal obtained by a recording device 111 attached to an object other than the target object.
[0213] In such a case, the learning device may be configured as shown in Fig. 11. In Fig. 11, the same reference numerals are used to designate parts that correspond to those in Fig. 3, and the description thereof will be omitted where appropriate.
[0214] 11 includes an acquisition unit 81, a correction processing unit 211, a sensor database 82, a microphone recording database 83, a learning unit 84, and a coefficient database 85. The configuration of the learning device 201 is the same as that of the learning device 52, except that a correction processing unit 211 is newly provided.
[0215] In this example, the correction processing unit 211 supplies the microphone recording signal obtained by the target recording device 51, which is supplied from the acquisition unit 81, to the microphone recording database 83 as training data.
[0216] The correction processing unit 211 is also supplied with the position information obtained by the target recording device 51, and the position information and microphone recording signals obtained by the other recording devices 51.
[0217] Note that, here, an example will be described in which position information and microphone recording signals obtained by another recording device 51 are used. However, this is not limiting, and the position information and microphone recording signals used by the correction processing unit 211 may be position information of a microphone that is not attached to the target object and a microphone recording signal obtained by that microphone.
[0218] The correction processing unit 211 performs correction processing on the microphone recording signals of the other recording devices 51 based on the location information of the target recording device 51 and the location information of the other recording devices 51, and supplies the resulting microphone recording signals to the microphone recording database 83 as learning data for the target recording device 51, where they are recorded.
[0219] In the correction process, processing is performed according to the positional relationship between the target recording device 51 and the other recording devices 51.
[0220] Specifically, for example, in the correction process, a process is performed to shift the microphone recorded signal in the time direction, such as compensating for propagation delay for the microphone recorded signal, depending on the distance from the target recording device 51 to other recording devices 51, which is determined from the position information.
[0221] In addition, based on the motion information of the target recording device 51 and the motion information of other recording devices 51, correction of the sound transmission characteristics may be performed as a correction process depending on the relative orientation (direction) of other objects as seen from the target object.
[0222] The learning unit 84 performs machine learning using the position information and movement information of the target recording device 51 and the microphone recording signals of other recording devices 51 after correction processing as learning data, and the microphone recording signals obtained from the target recording device 51 as training data.
[0223] That is, the coefficient data constituting the object sound source generator 11 is generated by machine learning using position information, movement information, and corrected microphone recorded signals as learning data as inputs and microphone recorded signals as output.
[0224] In this case, there may be one or more other recording devices 51. In addition, the motion information and position information obtained by the other recording devices 51 may also be used as learning data for the target object (recording device 51).
[0225] Alternatively, training data for a target object may be generated based on a microphone-recorded signal after correction processing of another recording device 51.
[0226] <Explanation of learning process> Next, the learning process performed by the learning device 201 will be described with reference to the flowchart of FIG.
[0227] The processes in steps S161 and S162 are similar to those in steps S41 and S42 in FIG. 5, and therefore will not be described further.
[0228] However, in step S162, the labeled movement information and position information are supplied to the sensor database 82, and the position information is supplied to the correction processing unit 211. In addition, the labeled microphone recorded signal is supplied to the correction processing unit 211.
[0229] The correction processing unit 211 supplies the labeled microphone recording signal of the target recording device 51 supplied from the acquisition unit 81 as it is to the microphone recording database 83 as training data, and causes it to be recorded.
[0230] In step S163, the correction processing unit 211 performs correction processing on the microphone recording signals of the other recording devices 51 supplied from the acquisition unit 81 based on the location information of the target recording device 51 and the location information of the other recording devices 51 supplied from the acquisition unit 81.
[0231] The correction processing unit 211 supplies the microphone recording signal obtained by the correction processing to the microphone recording database 83 as learning data for the target recording device 51, and causes it to be recorded.
[0232] After the correction process is performed, learning is performed in step S164 and the learning process ends. However, since the process in step S164 is the same as the process in step S43 in FIG. 5, a description thereof will be omitted.
[0233] However, in step S164, machine learning is performed using not only the position information and movement information of the target recording device 51, but also the corrected microphone recorded signals of other recording devices 51 as learning data.
[0234] In this way, the learning device 201 performs machine learning using the microphone recorded signals of the other recording devices 51 that have been subjected to the correction process as learning data, and generates coefficient data.
[0235] In this way, even in a situation where a microphone recorded signal cannot be obtained, a high-quality target sound can be obtained using the object sound source generator 11. In particular, in this example, a microphone recorded signal obtained by another recording device can be used as an input to the object sound source generator 11, so that a target sound of even higher quality can be obtained.
[0236] <Configuration example of sound source generation device> Moreover, a sound source generation device that generates an object sound source signal using coefficient data obtained by the learning device 201 may be configured as shown in Fig. 13. In Fig. 13, parts that correspond to those in Fig. 6 are given the same reference numerals, and their explanation will be omitted as appropriate.
[0237] The sound source generating device 241 shown in FIG. 13 includes an acquisition unit 131, a coefficient database 132, a correction processing unit 251, and an object sound source generating unit 133.
[0238] The sound source generating device 241 has a configuration in which a correction processing unit 251 is newly provided in the configuration of the sound source generating device 112 .
[0239] The acquisition unit 131 of the sound source generating device 241 acquires not only transmission data from the target recording device 111 but also transmission data from other recording devices 111. In this case, it is assumed that the other recording devices 111 are provided with microphones 61, and the transmission data from the other recording devices 111 also includes microphone recording signals.
[0240] The correction processing unit 251 does not perform correction processing on the transmission data acquired from the target recording device 111 supplied from the acquisition unit 131, and supplies the supplied transmission data to the object sound source generation unit 133 as is.
[0241] On the other hand, the correction processing unit 251 performs correction processing on the microphone recording signal included in the transmission data acquired from the other recording device 111, which is supplied from the acquisition unit 131, and supplies the microphone recording signal after the correction processing to the object sound source generation unit 133. Note that, similar to the correction processing unit 211, the position information and microphone recording signal used in the correction processing unit 251 may be the position information of a microphone that is not attached to the target object and the microphone recording signal obtained by that microphone.
[0242] <Explanation of sound source generation process> Next, the sound source generation process by the sound source generation device 241 will be described with reference to the flowchart of FIG.
[0243] The process of step S191 is the same as the process of step S101 in FIG. 8, and therefore a description thereof will be omitted.
[0244] At this time, the correction processing unit 251 supplies the transmission data of the target recording device 111 out of the transmission data supplied from the acquisition unit 131 to the object sound source generation unit 133 as is.
[0245] In step S192, the correction processing unit 251 performs correction processing on the microphone recording signal contained in the transmission data acquired from another recording device 111 and supplied from the acquisition unit 131, and supplies the microphone recording signal after the correction processing to the object sound source generation unit 133.
[0246] For example, in step S192, the same correction process as in step S163 of FIG. 12 is performed based on the location information of the target recording device 111 and the location information of the other recording devices 111.
[0247] After the correction process is performed, the processes of steps S193 and S194 are performed and the sound source generation process ends. However, since these processes are similar to the processes of steps S102 and S103 in FIG. 8, a description thereof will be omitted.
[0248] However, in step S193, for example, the sound source type identifier 21 of the object sound source generator 11 receives not only the position information and movement information of the target recording device 111, but also the microphone recording signal after correction processing supplied from the correction processing unit 251, and performs calculations.
[0249] That is, the object sound source generation unit 133 generates an object sound source signal corresponding to the microphone recording signal of the target recording device 111 based on the position information and movement information of the target recording device 111, the microphone recording signal after correction processing of the other recording devices 111, and coefficient data.
[0250] In this way, the sound source generating device 241 generates and outputs an object sound source signal using the microphone recorded signals of the other recording devices 111 that have been subjected to correction processing.
[0251] In this way, even if a microphone recording signal cannot be obtained from the target recording device 111, a high-quality target sound, that is, a high-quality object sound source signal can be obtained by using a microphone recording signal obtained from another recording device 111.
[0252] For example, even if only a microphone recorded signal with a low S / N ratio is obtained from the target recording device 111, the sound source generation device 241 may generate an object sound source signal using a microphone recorded signal obtained from another recording device 111. In this way, an object sound source signal with a higher S / N ratio, that is, a higher quality object sound source signal, can be obtained.
[0253] Fourth Embodiment <Explanation of learning process> For example, when recording content in a harsh environment such as a sports event, it is possible that some of the sensor signals used to generate the object sound source signal may not be obtained for some reason, such as a malfunction of sensors such as the motion measurement unit 62 or the position measurement unit 63 provided in the recording device 111.
[0254] Therefore, when actually generating an object sound source signal, it is possible to obtain the object sound source signal using obtainable sensor signals from the plurality of sensor signals, on the assumption that some of the sensor signals will not be obtained (will be missing).
[0255] In such a case, it is sufficient to train the object sound source generator 11 in advance, which inputs all the combinations of multiple sensor signals that can be acquired and outputs the object sound source signals. In this way, although the estimation accuracy may be lower, it is possible to obtain the object sound source signal from only some of the sensor signals.
[0256] For example, if a maximum of N types of sensor signals can be acquired, then Σ n=1 N N C n It is sufficient to train the object sound source generator 11 in advance.
[0257] It may not be practical to prepare object sound source generators 11 for all types (combinations) of sensor signals, that is, for all missing patterns. In such cases, learning may be performed only for patterns in which sensor signals with a high frequency of failure are missing, and for sensor signal missing patterns with a low frequency, no object sound source generator 11 may be prepared, or an object sound source generator 11 with a simpler configuration may be prepared.
[0258] In this way, when the object sound source generator 11, that is, the coefficient data, is prepared for each combination of sensor signals, the learning device 52 performs a learning process shown in FIG.
[0259] The learning process by the learning device 52 will be described below with reference to the flowchart in Fig. 15. Note that the processes in steps S221 and S222 are similar to the processes in steps S41 and S42 in Fig. 5, and therefore description thereof will be omitted.
[0260] In step S223, the learning unit 84 performs machine learning for each combination of multiple sensor signals that make up the learning data based on the learning data recorded in the sensor database 82 and the teacher data recorded in the microphone recording database 83.
[0261] That is, in machine learning for each combination, coefficient data is generated for the object sound source generator 11 using a predetermined combination of sensor signals as learning data as input and microphone recording signals as teacher data as output.
[0262] In this example as well, microphone recorded signals obtained by the target recording device 51 and microphone recorded signals obtained by other recording devices 51 may also be used as learning data.
[0263] The learning unit 84 supplies the coefficient data that constitutes the object sound source generator 11, obtained for each combination of sensor signals, to the coefficient database 85 for recording, and the learning process ends.
[0264] In this way, the learning device 52 performs machine learning for each combination of a plurality of sensor signals, and generates, for each combination, coefficient data that configures the object sound source generator 11. In this way, even in a situation where some sensor signals cannot be obtained, it becomes possible to robustly generate an object sound source signal from the obtained sensor signals.
[0265] <Configuration example of sound source generation device> Furthermore, when coefficient data is generated for each combination of sensor signals, the sound source generation device may be configured as shown in Fig. 16. In Fig. 16, parts corresponding to those in Fig. 6 are given the same reference numerals, and their explanation will be omitted as appropriate.
[0266] The sound source generating device 281 shown in FIG. 16 includes an acquisition unit 131, a coefficient database 132, a fault detection unit 291, and an object sound source generating unit 133.
[0267] The sound source generating device 281 has the same configuration as the sound source generating device 112, but with a new failure detection unit 291 added.
[0268] The fault detection unit 291 detects, for each sensor signal constituting the transmission data supplied from the acquisition unit 131, a fault (malfunction) of the sensor that acquired the sensor signal based on the sensor signal, and supplies the detection result to the object sound source generation unit 133.
[0269] Furthermore, the fault detection unit 291 configures the transmission data supplied from the acquisition unit 131 according to the fault detection result, that is, supplies only the sensor signals of sensors that are not faulty and are operating normally out of the sensor signals included in the transmission data to the object sound source generation unit 133.
[0270] Furthermore, the coefficient database 132 stores coefficient data generated for each combination of sensor signals.
[0271] Based on the failure detection result supplied from the failure detection unit 291, the object sound source generation unit 133 reads out from the coefficient database 132 coefficient data of the object sound source generator 11 that receives a healthy sensor signal as input and outputs an object sound source signal.
[0272] The object sound source generating unit 133 then generates an object sound source signal of the target object sound source based on the read coefficient data and the sensor signals of the non-faulty sensors supplied from the fault detecting unit 291.
[0273] In this example, the transmission data acquired by the acquisition unit 131 from the target recording device 111 may include a microphone recording signal as a sensor signal.
[0274] In such a case, for example, if the failure detection unit 291 does not detect a failure of the microphone corresponding to the microphone recorded signal, the object sound source generation unit 133 can output the microphone recorded signal as is or a signal generated from the microphone recorded signal as the object sound source signal.
[0275] On the other hand, if a microphone failure is detected, the object sound source generation unit 133 can generate an object sound source signal based on the sensor signals of sensors other than the microphone recording signals in which no failure has been detected.
[0276] <Explanation of sound source generation process> Next, the sound source generation process by the sound source generation device 281 will be described with reference to the flowchart of FIG.
[0277] The process of step S251 is the same as the process of step S101 in FIG. 8, and therefore a description thereof will be omitted.
[0278] In step S 252 , the failure detection unit 291 detects a sensor failure for each sensor signal constituting the transmission data supplied from the acquisition unit 131 , and supplies the detection result to the object sound source generation unit 133 .
[0279] For example, the fault detection unit 291 detects whether or not there is a fault for each sensor signal by performing a pure zero check on the sensor signal or detecting abnormal values, or by using a DNN that takes the sensor signal as input and outputs the detection result of whether or not there is a fault.
[0280] Furthermore, the failure detection unit 291 supplies only the sensor signals of sensors that are not malfunctioning, among the sensor signals included in the transmission data supplied from the acquisition unit 131, to the object sound source generation unit 133.
[0281] In step S 253 , the object sound source generator 133 generates an object sound source signal in accordance with the failure detection result supplied from the failure detector 291 .
[0282] That is, the object sound source generator 133 reads out from the coefficient database 132 the coefficient data of the object sound source generator 11 that uses a sensor signal from a sensor that is not malfunctioning as an input, that is, that does not use a sensor signal from a sensor in which a malfunction has been detected as an input.
[0283] The object sound source generator 133 then generates an object sound source signal based on the read coefficient data and the sensor signal supplied from the failure detector 291 .
[0284] After generating the object sound source signal, the object sound source generating unit 133 outputs the generated object sound source signal in step S254, and the sound source generation process ends.
[0285] In this way, the sound source generator 281 detects a sensor failure based on the transmitted data, and generates an object sound source signal using coefficient data according to the detection result.
[0286] In this way, even if some of the sensors fail for some reason, the object sound source signal can be robustly generated from the sensor signals of the remaining sensors. In other words, the object sound source signal can be robustly generated against changes in the situation.
[0287] Fifth Embodiment <Explanation of learning process> Furthermore, the surrounding environmental conditions may differ between when the learning data and teacher data used for learning are recorded (acquired) and when the actual content is recorded.
[0288] The environmental conditions referred to here are the environment surrounding the recording device, such as the sound absorption coefficient of the floor material, the type of shoes worn by the object athlete or performer, the reverberation characteristics of the space where recording takes place, the type of space where recording takes place (closed space or open space), the volume and 3D shape of the space where recording takes place (target space), weather, ground conditions, and type of ground.
[0289] By training the object sound source generator 11 for each of these different environmental conditions, it is possible to obtain an object sound source signal with a sound quality closer to reality, adapted to the environment in which the content was recorded.
[0290] In such a case, the learning device 52 shown in FIG. 3 performs the learning process shown in FIG.
[0291] The learning process performed by the learning device 52 will be described below with reference to the flowchart in Fig. 18. Note that the processes in steps S281 and S282 are similar to the processes in steps S41 and S42 in Fig. 5, and therefore description thereof will be omitted.
[0292] However, in step S282, the acquisition unit 81 acquires environmental condition information indicating the environment surrounding the recording device 51 by some method, and performs a labeling process in which not only the sound source type information but also the environmental condition information is associated with the microphone recording signal, movement information, and position information.
[0293] For example, the environmental condition information may be manually input by a user, or may be obtained by performing image recognition processing on a video signal obtained by a camera, or may be obtained by obtaining weather information from a server via a network.
[0294] In step S283, the learning unit 84 performs machine learning for each environmental condition based on the learning data recorded in the sensor database 82 and the teacher data recorded in the microphone recording database 83.
[0295] That is, the learning unit 84 performs machine learning using only the learning data and teacher data associated with environmental condition information indicating the same environmental condition, and generates coefficient data that configures the object sound source generator 11 for each environmental condition.
[0296] The learning unit 84 supplies the coefficient data for each environmental condition thus obtained to the coefficient database 85 for recording, and the learning process ends.
[0297] In this way, the learning device 52 performs machine learning for each environmental condition and generates coefficient data. In this way, it becomes possible to obtain an object sound source signal with sound quality closer to reality by using coefficient data according to the environmental condition.
[0298] <Configuration example of sound source generation device> Furthermore, when coefficient data is generated for each environmental condition, the sound source generation device may be configured as shown in Fig. 19. In Fig. 19, parts corresponding to those in Fig. 6 are given the same reference numerals, and their explanation will be omitted as appropriate.
[0299] The sound source generating device 311 shown in FIG. 19 includes an acquisition unit 131, a coefficient database 132, an environmental condition acquisition unit 321, and an object sound source generating unit 133.
[0300] The sound source generating device 311 has a configuration in which an environmental condition acquiring unit 321 is newly added to the configuration of the sound source generating device 112. Also, the coefficient database 132 records coefficient data for each environmental condition.
[0301] The environmental condition acquisition unit 321 acquires environmental condition information indicating the surrounding environment (environmental conditions) of the recording device 111 and supplies it to the object sound source generation unit 133.
[0302] For example, the environmental condition acquisition unit 321 may acquire information indicating the environmental conditions input by a user or the like, and use that information as the environmental condition information as is, or may acquire information indicating the weather around the recording device 111, such as sunny or rainy weather, acquired from an external server or the like, and use that information as the environmental condition information.
[0303] Furthermore, for example, the environmental condition acquisition unit 321 may acquire a video signal having the surroundings of the recording device 111 as a subject, identify the ground condition and type of ground around the recording device 111 by performing image recognition on the video signal or DNN arithmetic processing using the video signal as input, and generate environmental condition information indicating the identification results.
[0304] Here, the ground condition refers to the state of the ground determined by the weather, such as dry or after rain, and the ground type refers to the type of ground determined by the ground material, such as hard or grass. In addition, the environmental conditions are not limited to video signals, and may be identified by calculations using a classifier such as a DNN based on some observed value, such as temperature, a rain gauge, or a hygrometer, or by any signal processing such as image recognition or threshold processing.
[0305] The object sound source generation unit 133 reads out coefficient data corresponding to the environmental condition information supplied from the environmental condition acquisition unit 321 from the coefficient database 132, and generates an object sound source signal based on the read coefficient data and the transmission data supplied from the acquisition unit 131.
[0306] <Explanation of sound source generation process> Next, the sound source generation process by the sound source generation device 311 will be described with reference to the flowchart of Fig. 20. Note that the process of step S311 is the same as the process of step S101 in Fig. 8, and therefore description thereof will be omitted.
[0307] In step S 312 , the environmental condition acquisition unit 321 acquires environmental condition information and supplies it to the object sound source generation unit 133 .
[0308] For example, the environmental condition acquisition unit 321 acquires information indicating the weather from an external server as described above and uses it as environmental condition information, or performs image recognition on a video signal or the like to identify the environmental conditions, thereby obtaining environmental condition information.
[0309] In step S313, the object sound source generation unit 133 generates an object sound source signal in accordance with the environmental conditions.
[0310] That is, the object sound source generating unit 133 reads out from the coefficient database 132 the coefficient data generated for the environmental conditions indicated by the environmental condition information supplied from the environmental condition acquiring unit 321 .
[0311] Then, the object sound source generating unit 133 performs calculation processing by the object sound source generator 11 based on the read coefficient data and the transmission data supplied from the acquiring unit 131, thereby generating an object sound source signal.
[0312] In step S314, the object sound source generation unit 133 outputs the generated object sound source signal, and the sound source generation process ends.
[0313] In this way, the sound source generation device 311 generates an object sound source signal using coefficient data corresponding to the environmental conditions, thereby obtaining an object sound source signal that is suited to the environmental conditions and has a sound quality closer to reality.
[0314] Sixth Embodiment <Example of learning device configuration> Incidentally, the quality of the obtained sensor signal, such as the SN ratio, may differ between when learning data or teacher data used for learning is recorded (acquired) and when actual content is recorded.
[0315] For example, consider the case where a soccer match is recorded as content.
[0316] In this case, the learning data and teacher data used for learning are recorded during practice, so there is a high possibility that the microphone-recorded sensor signal obtained will not contain noise such as cheers from the surroundings or the rustling of clothes, resulting in a sensor signal with a high signal-to-noise ratio (high quality).
[0317] In contrast, since actual content is recorded during a match, the microphone-recorded sensor signal obtained is likely to contain noise such as cheers from the surrounding area and the rustling of clothing, resulting in a low signal-to-noise ratio (low-quality) sensor signal.
[0318] Furthermore, the reverberation characteristics of the surrounding space may differ between practice and game times, such as there being less reverberation during practice and more reverberation during game times.
[0319] Therefore, when the learning data and teacher data are recorded, the environment is lower in reverberation and noise than when the content is recorded, in other words, the content is recorded in a low S / N environment, and machine learning can be performed using learning data that simulates such a low S / N environment. By doing so, it is possible to obtain object sound source signals of higher quality.
[0320] In such a case, the learning device may be configured, for example, as shown in Fig. 21. In Fig. 21, the same reference numerals are used to designate parts that correspond to those in Fig. 3, and the description thereof will be omitted where appropriate.
[0321] The learning device 351 shown in FIG. 21 includes an acquisition unit 81, a superimposition processing unit 361, a sensor database 82, a microphone recording database 83, a learning unit 84, and a coefficient database 85.
[0322] The configuration of the learning device 351 differs from that of the learning device 52 in that a superimposition processing unit 361 is newly provided, but in other respects it has the same configuration as that of the learning device 52.
[0323] The superimposition processing unit 361 does not perform any processing on the microphone recording signal as training data supplied from the acquisition unit 81, but supplies the microphone recording signal as training data as is to the microphone recording database 83 for recording.
[0324] In addition, the superposition processing unit 361 performs a superposition process to convolve noise-added data for adding reverberation, noise, etc. onto the microphone-recorded signal as training data supplied from the acquisition unit 81, and supplies the resulting low-SN microphone-recorded signal with a low S / N ratio as learning data to the microphone-recorded database 83 for recording.
[0325] The noise addition data may be data that can add at least one of reverberation and noise to a microphone-recorded signal.
[0326] The low SNR obtained in this way, i.e., the low-quality low SNR microphone recording signal, is the signal that would be obtained in a low SNR environment.
[0327] Therefore, since a signal recorded by a low S / N microphone should be close to a signal actually recorded by a microphone when recording content, if such a signal recorded by a low S / N microphone is used as training data, it is possible to predict an object sound source signal with higher accuracy, that is, it is possible to obtain an object sound source signal of higher quality.
[0328] The superimposition processing unit 361 does not perform superimposition processing on the position information and movement information as learning data supplied from the acquisition unit 81, but supplies the position information and movement information as they are as learning data to the sensor database 82 for recording.
[0329] It is possible that noise not present when the learning data was recorded, such as when a player as an object comes into contact with another player, may be included in the sensor signal when the actual content is recorded. Therefore, noise-added data for adding such noise may be prepared in advance, and the superimposition processing unit 361 may convolve the noise-added data with the position information and movement information to provide the result as learning data to the sensor database 82.
[0330] <Explanation of learning process> Next, the learning process performed by the learning device 351 will be described with reference to the flowchart of FIG.
[0331] The processes in steps S341 and S342 are the same as those in steps S41 and S42 in FIG. 5, and therefore will not be described further.
[0332] In step S343, the superposition processing unit 361 performs a superposition process to convolve noise-added data onto the microphone recording signal supplied from the acquisition unit 81, and supplies the resulting low SN microphone recording signal as learning data to the microphone recording database 83 for recording.
[0333] The superimposition processing unit 361 also supplies the microphone recording signal supplied from the acquisition unit 81 as it is to the microphone recording database 83 as training data for recording, and also supplies the position information and movement information supplied from the acquisition unit 81 as it is to the sensor database 82 as learning data for recording. Note that, as described above, noise-added data may be convoluted with the position information and movement information.
[0334] After the superimposition process is performed, the process of step S344 is performed and the learning process ends. However, since the process of step S344 is the same as the process of step S43 in FIG. 5, a description thereof will be omitted.
[0335] However, in step S344, the low SN microphone recording signal and the position information and movement information, which are sensor signals other than the microphone recording signals to which reverberation and noise are added, among the multiple sensor signals included in the transmission data, are used as learning data.
[0336] Then, machine learning is performed based on such learning data and microphone-recorded signals as teacher data.
[0337] As a result, coefficient data for the object sound source generator 11 is generated, which receives position information, movement information, and a signal recorded by a low SN microphone as input and outputs an object sound source signal.
[0338] In this way, the learning device 351 generates a low-SN microphone recorded signal that simulates recording in a low-SN environment from the microphone recorded signal, and performs machine learning using the low-SN microphone recorded signal as learning data. In this way, it becomes possible to obtain a higher quality object sound source signal even in a low-SN environment when recording content.
[0339] In this case, a microphone 61 is provided in the recording device 111 when recording content, and the sound source generating device 112 acquires transmission data including a microphone recording signal, position information, and movement information from the recording device 111. This microphone recording signal is recorded in a low S / N environment, and therefore corresponds to the low S / N microphone recording signal described above.
[0340] In the sound source generating device 112, the processing of step S102 in FIG. 8, i.e., calculation processing based on the object sound source generator 11, is performed based on the microphone recording signal, position information, and movement information contained in the transmission data, and the coefficient data, and an object sound source signal is generated.
[0341] In this case, unnecessary noise and the like are suppressed, and a high-quality object sound source signal, that is, a high S / N ratio, can be obtained.
[0342] Any of the first to sixth embodiments described above may be combined.
[0343] <Example of computer configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.
[0344] FIG. 23 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes using a program.
[0345] In the computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are interconnected by a bus 504.
[0346] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0347] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a nonvolatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0348] In a computer configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.
[0349] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0350] In a computer, a program can be installed in a recording unit 508 via an input / output interface 505 by inserting a removable recording medium 511 into a drive 510. The program can also be received by a communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in advance in the ROM 502 or the recording unit 508.
[0351] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0352] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0353] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.
[0354] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.
[0355] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0356] Furthermore, the present technology can also be configured as follows.
[0357] (1) a learning unit that performs learning based on one or more sensor signals obtained by one or more sensors attached to an object and a target signal related to the object corresponding to a predetermined sensor, and generates coefficient data that constitutes a generator that receives the one or more sensor signals as input and outputs the target signal; Learning device. (2) The target signal is a sound source signal corresponding to a microphone as the predetermined sensor. The learning device according to (1). (3) The target signal is an acoustic signal generated based on a microphone signal obtained by the microphone serving as the predetermined sensor attached to the object. (2) A learning device according to the present invention. (4) The one or more sensors include at least one of a nine-axis sensor, a geomagnetic sensor, an acceleration sensor, a gyro sensor, a distance measurement sensor, a positioning sensor, an image sensor, and a microphone. A learning device according to any one of (1) to (3). (5) The one or more sensors are sensors of a type different from a microphone. A learning device according to any one of (1) to (3). (6) a correction processing unit that performs a correction process on a microphone recording signal obtained by another microphone not attached to the object in accordance with a positional relationship between the object and the other microphone; The learning unit performs the learning based on the microphone recorded signal after the correction processing, the one or more sensor signals, and the target signal, and generates the coefficient data constituting the generator that receives the microphone recorded signal after the correction processing and the one or more sensor signals as inputs and outputs the target signal. A learning device according to any one of (1) to (5). (7) The learning unit performs the learning for each combination of the plurality of sensors and generates the coefficient data. A learning device according to any one of (1) to (5). (8) The learning unit performs the learning for each environmental condition around the object and generates the coefficient data. A learning device according to any one of (1) to (5). (9) The system further includes a superimposition processing unit that adds reverberation or noise to a microphone-recorded signal as the sensor signal obtained by a microphone as the sensor, The learning unit performs the learning based on the microphone-recorded signal to which the reverberation or noise has been added, the sensor signals other than the microphone-recorded signal among the one or more sensor signals, and the target signal, and generates the coefficient data constituting the generator that receives the microphone-recorded signal and the sensor signals other than the microphone-recorded signal as inputs and outputs the target signal. A learning device according to any one of (1) to (4). (10) The learning device Learning is performed based on one or more sensor signals obtained by one or more sensors attached to an object and a target signal related to the object corresponding to a predetermined sensor, and coefficient data constituting a generator that receives the one or more sensor signals as input and outputs the target signal is generated. How to learn. (11) Learning is performed based on one or more sensor signals obtained by one or more sensors attached to an object and a target signal related to the object corresponding to a predetermined sensor, and coefficient data constituting a generator that receives the one or more sensor signals as input and outputs the target signal is generated. A program that causes a computer to perform a process. (12) an acquisition unit that acquires one or more sensor signals obtained by one or more sensors attached to the object; a generator that generates a target signal related to the object, corresponding to a predetermined sensor, based on coefficient data constituting a generator that has been generated in advance by learning and the one or more sensor signals; A signal processing device comprising: (13) The target signal is a sound source signal corresponding to a microphone as the predetermined sensor attached to the object. The signal processing device according to (12). (14) The one or more sensors include at least one of a nine-axis sensor, a geomagnetic sensor, an acceleration sensor, a gyro sensor, a distance measurement sensor, a positioning sensor, an image sensor, and a microphone. The signal processing device according to (12) or (13). (15) The one or more sensors are sensors of a type different from a microphone. The signal processing device according to (12) or (13). (16) the acquisition unit further acquires a microphone recording signal obtained by another microphone not attached to the object, a correction processing unit that performs a correction process on the microphone recording signal obtained by the other microphone in accordance with a positional relationship between the object and the other microphone; The generation unit generates the target signal based on the coefficient data, the microphone-recorded signal after the correction process, and the one or more sensor signals. A signal processing device according to any one of (12) to (15). (17) the acquisition unit acquires a plurality of the sensor signals; a failure detection unit that detects a failure of the sensor based on the plurality of sensor signals; The generator generates the target signal based on the sensor signal of the sensor that is not faulty among the plurality of sensor signals and the coefficient data constituting the generator that receives the sensor signal of the sensor that is not faulty as an input and outputs the target signal. A signal processing device according to any one of (12) to (15). (18) an environmental condition acquisition unit that acquires environmental condition information indicating an environmental condition around the object; The generation unit generates the target signal based on the coefficient data corresponding to the environmental condition information and the one or more sensor signals. A signal processing device according to any one of (12) to (15). (19) The signal processing device acquiring one or more sensor signals obtained by one or more sensors attached to the object; A target signal for the object corresponding to a predetermined sensor is generated based on coefficient data constituting a generator generated in advance by learning and the one or more sensor signals. Signal processing methods. (20) acquiring one or more sensor signals obtained by one or more sensors attached to the object; A target signal for the object corresponding to a predetermined sensor is generated based on coefficient data constituting a generator generated in advance by learning and the one or more sensor signals. A program that causes a computer to perform a process. [Explanation of symbols]
[0358] 11 object sound source generator, 51 recording device, 52 learning device, 81 acquisition unit, 84 learning unit, 111 recording device, 112 sound source generation device, 131 acquisition unit, 133 object sound source generation unit
Claims
1. a learning unit that performs learning based on one or more sensor signals obtained by one or more sensors attached to an object and a target signal related to the object that corresponds to a predetermined sensor attached to the object, and generates coefficient data that constitutes a generator that receives the one or more sensor signals as input and outputs the target signal; The target signal is a microphone-recorded signal obtained by a microphone serving as the predetermined sensor, or an acoustic signal generated based on the microphone-recorded signal. Learning device.
2. The one or more sensors include at least one of a nine-axis sensor, a geomagnetic sensor, an acceleration sensor, a gyro sensor, a distance measurement sensor, a positioning sensor, an image sensor, and a microphone. The learning device according to claim 1 .
3. The one or more sensors are sensors of a type different from a microphone. The learning device according to claim 1 .
4. a correction processing unit that performs a correction process on a microphone recording signal obtained by another microphone not attached to the object in accordance with a positional relationship between the object and the other microphone; The learning unit performs the learning based on the microphone recorded signal after the correction processing, the one or more sensor signals, and the target signal, and generates the coefficient data constituting the generator that receives the microphone recorded signal after the correction processing and the one or more sensor signals as inputs and outputs the target signal. The learning device according to claim 1 .
5. The learning unit performs the learning for each combination of the plurality of sensors and generates the coefficient data. The learning device according to claim 1 .
6. The learning unit performs the learning for each environmental condition around the object and generates the coefficient data. The learning device according to claim 1 .
7. The system further includes a superimposition processing unit that adds reverberation or noise to a microphone-recorded signal as the sensor signal obtained by a microphone as the sensor, The learning unit performs the learning based on the microphone recorded signal to which the reverberation or noise has been added, the sensor signal other than the microphone recorded signal among the one or more sensor signals, and the target signal, and generates the coefficient data constituting the generator that receives the microphone recorded signal and the sensor signal other than the microphone recorded signal as inputs and outputs the target signal. The learning device according to claim 1 .
8. The learning device performing learning based on one or more sensor signals obtained by one or more sensors attached to an object and a target signal related to the object that corresponds to a predetermined sensor attached to the object, and generating coefficient data constituting a generator that receives the one or more sensor signals as input and outputs the target signal; The target signal is a microphone-recorded signal obtained by a microphone serving as the predetermined sensor, or an acoustic signal generated based on the microphone-recorded signal. How to learn.
9. Learning is performed based on one or more sensor signals obtained by one or more sensors attached to an object and a target signal related to the object corresponding to a predetermined sensor attached to the object, and coefficient data constituting a generator that receives the one or more sensor signals as input and outputs the target signal is generated. Have the computer execute the process, The target signal is a microphone-recorded signal obtained by a microphone serving as the predetermined sensor, or an acoustic signal generated based on the microphone-recorded signal. program.
10. an acquisition unit that acquires one or more sensor signals obtained by one or more sensors attached to the object; a generator that generates a target signal for the object corresponding to a predetermined sensor not attached to the object, based on coefficient data constituting a generator that has been generated in advance by learning and the one or more sensor signals; Equipped with The target signal is a microphone-recorded signal obtained by a microphone serving as the predetermined sensor when the microphone is attached to the object, or an acoustic signal generated based on the microphone-recorded signal. Signal processing device.
11. The generation unit generates metadata of the target signal based on the one or more sensor signals. The signal processing device according to claim 10.
12. The one or more sensors include at least one of a nine-axis sensor, a geomagnetic sensor, an acceleration sensor, a gyro sensor, a distance measurement sensor, a positioning sensor, an image sensor, and a microphone. The signal processing device according to claim 10.
13. The one or more sensors are sensors of a type different from a microphone. The signal processing device according to claim 10.
14. the acquisition unit further acquires a microphone recording signal obtained by another microphone not attached to the object, a correction processing unit that performs a correction process on the microphone recording signal obtained by the other microphone in accordance with a positional relationship between the object and the other microphone; The generation unit generates the target signal based on the coefficient data, the microphone-recorded signal after the correction process, and the one or more sensor signals. The signal processing device according to claim 10.
15. the acquisition unit acquires a plurality of the sensor signals; a failure detection unit that detects a failure of the sensor based on the plurality of sensor signals; The generator generates the target signal based on the sensor signal of the sensor that is not faulty among the plurality of sensor signals and the coefficient data constituting the generator that receives the sensor signal of the sensor that is not faulty as an input and outputs the target signal. The signal processing device according to claim 10.
16. an environmental condition acquisition unit that acquires environmental condition information indicating an environmental condition around the object; The generation unit generates the target signal based on the coefficient data corresponding to the environmental condition information and the one or more sensor signals. The signal processing device according to claim 10.
17. The signal processing device acquiring one or more sensor signals obtained by one or more sensors attached to the object; generating a target signal for the object corresponding to a predetermined sensor not attached to the object based on coefficient data constituting a generator generated in advance by learning and the one or more sensor signals; The target signal is a microphone-recorded signal obtained by a microphone serving as the predetermined sensor when the microphone is attached to the object, or an acoustic signal generated based on the microphone-recorded signal. Signal processing methods.
18. acquiring one or more sensor signals obtained by one or more sensors attached to the object; A target signal for the object corresponding to a predetermined sensor not attached to the object is generated based on coefficient data constituting a generator generated in advance by learning and the one or more sensor signals. Have the computer execute the process, The target signal is a microphone-recorded signal obtained by a microphone serving as the predetermined sensor when the microphone is attached to the object, or an acoustic signal generated based on the microphone-recorded signal. program.
Citation Information
Patent Citations
Animal-machine audio interaction system
JP2011081383A
Information processing device, information processing method and program
JP2017205213A
Signal processing device, signal processing method, and program
WO2020208926A1