Signal processing device and method, learning device and method, and program

By installing multiple sensors on the object and generating an object sound source generator, the problem of microphones not being able to be installed in free viewpoint content recording is solved, achieving high-quality target sound acquisition and reducing user burden and device power consumption.

CN116324977BActive Publication Date: 2026-04-07SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-06
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In free-viewpoint content recording, especially in scenarios such as sports or games, it is difficult to obtain target sound with a high signal-to-noise ratio because the microphone device cannot be installed on the sound source, resulting in low target sound quality.

Method used

By installing various sensors on an object, such as microphones, accelerometers, gyroscopes, and rangefinders, sensor signals are acquired and a sound source generator is generated through a learning device to produce corresponding target signals, including sound source type identification, envelope estimation, and signal mixing. High-quality target sound can even be acquired without a microphone.

Benefits of technology

It enables the acquisition of high-quality target sound without installing a microphone, reducing reliance on microphones, lowering the burden on users, and extending the device's battery life. It can also robustly acquire high-quality target sound even in harsh environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116324977B_ABST
    Figure CN116324977B_ABST
Patent Text Reader

Abstract

The present technology relates to a signal processing apparatus and method, a learning apparatus and method, and a program capable of acquiring a target sound with high quality. The learning apparatus includes a learning unit configured to perform learning based on one or more sensor signals acquired by one or more sensors mounted on an object and a target signal related to the object and corresponding to a predetermined sensor, and generate coefficient data configuring a generator that receives one or more of a plurality of sensor signals as an input thereof and outputs a target signal. The present technology can be applied to the learning apparatus.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present technology relates to a signal processing apparatus and a signal processing method, a learning apparatus and a learning method, and a program, and more particularly, to a signal processing apparatus and a signal processing method, a learning apparatus and a learning method, and a program capable of acquiring a target sound with high quality. BACKGROUND

[0002] In sound field reproduction of a free viewpoint such as a bird's-eye view, a walk, and the like, it is important to record a target sound of a sound source with a high signal-to-noise ratio (SN ratio), while it is necessary to acquire information indicating a position and an orientation of each sound source.

[0003] As a specific example of the target sound of the sound source, for example, a human voice, a general operation sound of a human such as a walking sound and a running sound, an operation sound specific to sports content, a game, or the like such as a kick of a ball, and the like can be given.

[0004] Further, for example, as a technology related to recognition of a user's action, a technology capable of acquiring one or a plurality of results of recognition of a user's action by analyzing ranging sensor data detected by a plurality of ranging sensors has been proposed (for example, see PTL 1).

[0005] [LIST OF CITATIONS]

[0006] [PATENT LITERATURE]

[0007] [PTL 1]

[0008] Japanese Patent Publication No. 2017-205213. SUMMARY

[0009] [TECHNICAL PROBLEM]

[0010] However, in a case where a content recorded as a free viewpoint moves, a game, or the like, as in a case where a device loaded with a microphone cannot be loaded in a player or the like used as a sound source or the like, it is also difficult to acquire a target sound of a sound source with a high SN ratio. In other words, it is difficult to obtain a target sound with high quality.

[0011] The present technology, in view of such a situation, is capable of acquiring a target sound with high quality.

[0012] [SOLUTION TO PROBLEM]

[0013] According to a first aspect of the present technology, there is provided a learning device including: a learning unit configured to perform learning based on one or more sensor signals acquired by one or more sensors mounted on an object and a target signal related to the object and corresponding to a predetermined sensor, and generate coefficient data configuring a generator having as an input the one or more sensor signals and having as an output the target signal.

[0014] According to a first aspect of the present technology, there is provided a learning method or program including the steps of: performing learning based on one or more sensor signals acquired by one or more sensors mounted on an object and a target signal related to the object and corresponding to a predetermined sensor, and generating coefficient data configuring a generator having as an input the one or more sensor signals and as an output the target signal.

[0015] In the first aspect of the present technology, learning is performed based on one or more sensor signals acquired by one or more sensors mounted on an object and a target signal related to the object and corresponding to a predetermined sensor, and coefficient data configuring a generator having as an input the one or more sensor signals and having as an output the target signal is generated.

[0016] According to a second aspect of the present technology, there is provided a signal processing device including: an acquisition unit configured to acquire one or more sensor signals acquired by one or more sensors mounted on an object; and a generation unit configured to generate a target signal related to the object and corresponding to a predetermined sensor based on coefficient data configuring a generator generated in advance through learning and the one or more sensor signals.

[0017] According to a second aspect of the present technology, there is provided a signal processing method or program including the steps of: acquiring one or more sensor signals acquired by one or more sensors mounted on an object; and generating a target signal related to the object and corresponding to a predetermined sensor based on coefficient data configuring a generator generated in advance through learning and the one or more sensor signals.

[0018] In the second aspect of the present technology, one or more sensor signals acquired by one or more sensors mounted on an object are acquired, and a target signal related to the object and corresponding to a predetermined sensor is generated based on coefficient data configuring a generator generated in advance through learning and the one or more sensor signals. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 is a diagram showing an example of a configuration of a device at the time of training and at the time of content recording.

[0020] Figure 2 is a diagram showing a configuration example of an object sound source generator.

[0021] Figure 3 is a diagram showing a configuration example of a recording device and a learning device.

[0022] Figure 4 is a flowchart for explaining a recording process executed at the time of generating learning data.

[0023] Figure 5 is a flowchart for explaining a learning process.

[0024] Figure 6 is a diagram showing a configuration example of a recording device and a sound source generation device.

[0025] Figure 7 is a flowchart for explaining a recording process executed at the time of generating an object sound source.

[0026] Figure 8 is a flowchart for explaining a sound source generation process.

[0027] Figure 9 is a diagram showing a configuration example of a learning device.

[0028] Figure 10 is a flowchart for explaining a learning process.

[0029] Figure 11 is a diagram showing a configuration example of a learning device.

[0030] Figure 12 is a flowchart for explaining a learning process.

[0031] Figure 13 is a diagram showing a configuration example of a sound source generation device.

[0032] Figure 14 is a flowchart for explaining a sound source generation process.

[0033] Figure 15 is a flowchart for explaining a learning process.

[0034] Figure 16 is a diagram showing a configuration example of a sound source generation device.

[0035] Figure 17 is a flowchart for explaining a sound source generation process.

[0036] Figure 18 is a flowchart for explaining a learning process.

[0037] Figure 19 is a diagram showing a configuration example of a sound source generation device.

[0038] Figure 20 is a flowchart for explaining a sound source generation process.

[0039] Figure 21 is a diagram showing a configuration example of a learning device.

[0040] Figure 22 is a flowchart for explaining a learning process.

[0041] Figure 23 is a diagram showing a configuration example of a computer. DETAILED DESCRIPTION

[0042] Hereinafter, an embodiment to which the present technology is applied will be described with reference to the drawings.

[0043] <First Embodiment>

[0044] <The Present Technology>

[0045] The present technology can acquire a target signal having high quality by generating a signal corresponding to a sensor signal acquired by another sensor based on a sensor signal acquired by one or more sensors.

[0046] For example, the sensor described herein is, for example, a microphone, an acceleration sensor, a gyro sensor, a geomagnetic sensor, a distance measuring sensor, an image sensor, or the like.

[0047] Hereinafter, an example will be described in which a signal corresponding to an object sound source signal of a microphone recording signal acquired by a microphone in a state in which the microphone is mounted on an object is produced as a target from a sensor signal of one or more sensors of a type different from each other, such as an acceleration sensor. In addition, the signal targeted (target signal) is not limited to the object sound source signal, and can be any signal such as a video signal of animation or the like.

[0048] For example, there are a small number of existing wearable devices that have a function of recording a voice and an operation sound with high sound quality when moving. There are devices used mainly in the broadcasting industry in which a small and strong transmitter is configured in combination with a Lavalier microphone. However, in such devices, no sensor other than a microphone is provided.

[0049] Furthermore, although there are wearable devices for analyzing mobility in which a sensor that acquires position information and movement information when moving is mounted, such devices do not have a function of acquiring a voice or even if they have such a function, they are not exclusively used for acquiring a voice.

[0050] Therefore, there is no device that can simultaneously acquire speech, location, and motion information with a high SN ratio while in motion and use them to generate an object sound source.

[0051] Therefore, in this technology, the audio signal (acoustic signal) (i.e., the object sound source signal) used to reproduce the sound of the target object sound source is configured to be generated from acquired sensor signals (such as location information, motion information, etc.).

[0052] For example, suppose there are multiple objects in the same target space, and each object has a recording device for recording content installed or built into it.

[0053] At this point, it is assumed that the sound emitted by an object with a built-in recording device is recorded as the sound of the object's sound source (recording).

[0054] For example, the target space is considered as a space for sports, opera, drama, film, etc., in which multiple athletes, performers, etc. exist.

[0055] Additionally, for example, an object within the target space can be a moving or stationary body, as long as it serves as a sound source (object sound source). More specifically, for example, an object can be a person such as an athlete, a robot or a vehicle with recording devices installed or built into it, a flying object such as a drone, etc.

[0056] Recording devices may include, for example, a microphone for receiving sound from an object's sound source, a motion measurement sensor (such as a 9-axis sensor) for measuring the object's movement and orientation (azimuth angle), a range sensor and a positioning sensor for measuring position, and a camera (image sensor) for capturing video of the surrounding environment.

[0057] Here, the ranging sensor (ranging device) and the positioning sensor are, for example, a Global Positioning System (GPS) device used to measure the position of an object, an indoor ranging signal receiver, etc., and the indoor ranging sensor and the positioning sensor can be used to obtain position information representing the position of the object.

[0058] Furthermore, motion information, such as velocity, acceleration, orientation, and movement of an object, can be obtained from the output of the motion measurement sensor installed in the recording device.

[0059] The recording device, by using a built-in microphone, motion measurement sensor, range sensor, and positioning sensor, can acquire microphone recording signals obtained by receiving sound from the object's surrounding environment, the object's position information, and the object's movement information. Furthermore, if a camera is installed in the recording device, video signals of the object's surroundings can also be acquired.

[0060] The microphone recording signals, location information, movement information, and video signals acquired for each object in this manner can be used to acquire the object's sound source signal, which is the acoustic signal of the sound from the object's sound source, and the sound from the object's sound source is the target sound.

[0061] Here, for example, the sound of an object that is considered the target sound source is the operational sound of a person walking, running, breathing, clapping, etc., and obviously, in addition, the spoken voice of a person that is an object can also be considered as the sound source of the object.

[0062] In this scenario, for example, by using microphone recording signals, location information, and motion information to detect the time periods of target sounds, such as operational sounds, present for each object, and performing signal processing to separate the target sounds from the microphone recording signals based on the detection results, the object sound source signal for each sound source type can be considered as generated for each object. Furthermore, it is possible to consider using the location and motion information acquired by multiple recording devices holistically when generating the object sound source signal.

[0063] In this case, a high-quality signal can be obtained as the object sound source signal for free viewpoint reproduction.

[0064] However, in practice, when recording content, it may not be possible to acquire all types of microphone recording signals, sensor signals, accelerometer outputs, etc.

[0065] More specifically, for example, a microphone cannot be used in situations where there are weight restrictions on the recording device so as not to impede the performance of the person with the recording device installed; and in cases where recorded speech is used for broadcasting, a microphone cannot be used in order to avoid broadcasting unintended speech such as strategies.

[0066] Therefore, in this technology, by acquiring more sensor signals at a time different from those during content recording than during content recording, and learning the object sound source generator, object sound source signals that cannot be acquired during content recording can be obtained from the sensor signals acquired during content recording.

[0067] As a specific example, one could consider situations where sports games, competitions, etc., are recorded as content.

[0068] In this case, the device structure of the recording device is changed to record data during training (practice), rehearsals (including trial games, etc.), and performances (i.e., recording content during games, actual performances, etc.).

[0069] For example, one could use something like... Figure 1 The device configuration shown.

[0070] exist Figure 1 In the example shown, during training and rehearsals, a microphone, an accelerometer, a gyroscope, a geomagnetic sensor, and position measurement sensors (range sensors and positioning sensors) for GPS, indoor ranging, etc., are installed in the recording device as sensors for recording data.

[0071] During training and rehearsals, the recording device may be heavy, and battery changes can be performed midway. Therefore, all sensors, including the microphone, are mounted in the recording device, acquiring all available sensor signals. In other words, the advanced recording device, including the microphone, is mounted on the player or actuator and acquires (collects) sensor signals.

[0072] In this way, the object sound source generator is generated by learning from sensor signals acquired during training and rehearsals.

[0073] In contrast, during gameplay and actual performances, sensors used for data recording, such as accelerometers, gyroscopes, magnetometers, and position measurement sensors for GPS, indoor ranging, etc., are incorporated into the recording device. In other words, microphones are not used during gameplay and performances.

[0074] During gameplay and actual performance, only a portion of the sensors installed during training and rehearsals are incorporated into the recording device, resulting in a lighter recording device and increased battery life. In other words, the type and number of sensors installed are reduced during gameplay and actual performance, and sensor signals are acquired through a lightweight and low-power recording device.

[0075] Specifically, it is insufficient to assume that the battery is non-replaceable during gameplay or actual performance, and that the battery lasts from the start to the end of the game or performance. Furthermore, in this example, no microphone is installed during gameplay or actual performance, thus preventing the recording and leakage of inappropriate audio, such as players' statements regarding strategies, during live broadcasts. Additionally, while microphone installation may be prohibited in some sports, it is still possible to acquire sensor signals from sensors other than microphones in such cases.

[0076] When sensor signals from different types of sensors (other than microphones) are acquired during gameplay or actual performance, an object sound source signal is generated based on these sensor signals and a pre-acquired object sound source generator. In the case where a microphone is mounted on a recording device during gameplay or actual performance—in other words, when the microphone is mounted on an object—this object sound source signal corresponds to the microphone recording signal acquired by the microphone. Furthermore, the generated object sound source signal may correspond to the microphone recording signal, may correspond to (and be associated with) the sound of the object sound source generated from the microphone recording signal, or may be a signal used to reproduce the sound corresponding to the sound of the object sound source generated from the microphone recording signal.

[0077] In the learning of object sound source generators, sensor signals recorded (acquired) during training or rehearsals are utilized, and prior information such as the personality of the contestant or performer equipped with the recording device and the environment in which the sensor signals are acquired can be obtained (learned).

[0078] Then, when recording content, that is, during gameplay or actual performance, based on prior information about the individual (object sound source generator) and a small amount of sensor signals during content recording, the object sound source signal of the performer or other object's operation sound is estimated (recovered).

[0079] In this way, it is possible to obtain the sound source signal of the object sound source (target sound) with high quality signal as the target, such as the sound of the player's operation, while reducing the physical burden on the player or performer or further extending the driving time of the recording device when recording content.

[0080] Furthermore, in this technology, when recording content, even when a microphone cannot be set in the recording device, when a signal from a microphone or another sensor cannot be acquired due to a malfunction, or when a sensor signal with a high signal-to-noise ratio cannot be acquired due to a harsh recording environment, etc., it can robustly acquire object sound source signals with high quality.

[0081] Here, the object sound source generator will be described further.

[0082] The object sound source generator can be any type, such as one that uses a deep neural network (DNN) or a parameter representation scheme, as long as it can acquire an object sound source signal corresponding to a target sensor signal from one or more sensor signals.

[0083] For example, in an object sound source generator with a parameter representation scheme, the sound source type of the target object sound source is restricted, and the sound source type of the object sound source, such as the sound of kicking a ball or the sound of running, is identified from one or more sensor signals.

[0084] Then, based on the sound source type identification result and the envelope of the signal strength (amplitude) of the sensor signal from the accelerometer, parametric signal processing is performed on the pre-prepared object sound source waveform signal to generate the object sound source signal.

[0085] For example, such an object sound source generator can have, for example, Figure 2 The configuration shown.

[0086] exist Figure 2 In the example shown, for instance, the object sound source generator 11 has sensor signals from multiple sensors of different types as inputs, and when a microphone is set in a recording device installed in the object, it outputs an object sound source signal corresponding to a microphone recording signal acquired by the microphone.

[0087] The object sound source generator 11 includes a sound source type identification unit 21, a sound source database 22, an envelope estimation unit 23, and a mixing unit 24.

[0088] Sensor signals acquired by sensors other than the microphone, which are mounted on a recording device installed on an object such as a player, which is a sound source in the object, are provided to the sound source type identification unit 21 and the envelope estimation unit 23. Here, sensor signals acquired by at least the accelerometer are included as sensor signals from sensors other than the microphone.

[0089] Furthermore, when a microphone recording signal can be acquired by a microphone arranged in a recording device mounted on an object or a microphone arranged in a recording device mounted on another object, the acquired microphone recording signal is also provided to the sound source type identification unit 21 and the envelope estimation unit 23.

[0090] When one or more sensor signals, including those recorded by a microphone, are provided, the sound source type identification unit 21 estimates the type of the sound source from the object—in other words, the type of sound emitted from the object—by performing appropriate arithmetic operations based on the provided sensor signals and coefficients acquired in advance through learning. The sound source type identification unit 21 then provides the sound source type information, representing the estimation result (identification result), to the sound source database 22.

[0091] For example, suppose the sound source type recognition unit 21 is a recognition unit that uses any machine learning algorithm such as a generalized linear determinant or a support vector machine (SVM).

[0092] The sound source database 22 stores source signals, which are acoustic signals having source waveforms for reproducing the sound emitted from an object source for each sound source type. In other words, the sound source database 22 stores sound source type information and source signals representing the sound source type as associated with each other.

[0093] For example, the source signal can be any signal, such as a signal obtained by actually recording the sound of an object source in an environment with a high signal-to-noise ratio (SN ratio), an artificially generated signal, or a signal generated from a microphone recording signal recorded by a microphone of a recording device used for learning by extraction. Furthermore, multiple source signals can be stored in association with a single source type information.

[0094] The sound source database 22 provides one of the multiple source signals of the sound source type represented by the sound source type information provided by the sound source type identification unit 21 from the stored source source signals of each sound source type to the mixing unit 24.

[0095] Envelope estimation unit 23 extracts the envelope of the sensor signal provided by the accelerometer and provides the envelope information representing the extracted envelope to mixing unit 24.

[0096] For example, the amplitude of a predetermined component of the sensor signal from the accelerometer is smoothed relative to time within a time interval of a predetermined length and extracted by the envelope estimation unit 23 as an envelope. For example, in the case where a walking sound object source signal is generated as the sound of the object source, the predetermined component described herein is set to include at least one or more components including the component of the gravity direction.

[0097] In addition, when extracting the envelope, the microphone recording signal and sensor signals from other sensors besides the accelerometer and microphone can also be used as needed.

[0098] The mixing unit 24 generates and outputs an object sound source signal by performing signal processing on one or more dry source signals provided by the sound source database 22 based on the envelope information provided by the envelope estimation unit 23.

[0099] For example, the signal processing performed by the mixing unit 24 is configured to perform waveform modulation processing of the interfering source signal based on the signal strength of the envelope represented by the envelope information, filtering of the interfering source signal based on the envelope information, and mixing processing of multiple interfering source signals.

[0100] The techniques used to generate object sound source signals in the object sound source generator 11 are not limited to the examples described herein. For example, the identification unit for identifying sound source type, the technique for estimating envelope, and the technique for generating object sound source signals from source signals can be arbitrary, and the configuration of the object sound source generator is not limited to... Figure 2 The example shown can also be another configuration.

[0101] <Configuration Examples of Recording and Learning Devices>

[0102] The following will describe a more detailed implementation of the above-described technology.

[0103] First, in cases where performances such as sports or performances (actual performances) are recorded as content, an example of generating an object sound source signal will be described where a microphone cannot be used or no microphone is arranged in the recording device.

[0104] In this example, an object sound source generator 11 generates object sound source signals from sensor signals of other types of sensors different from microphones by using machine learning.

[0105] For example, the recording device for acquiring sensor signals for learning the object sound source generator 11 and the learning device for performing the learning of the object sound source generator 11 are configured as follows: Figure 3 As shown. Here, although only one recording device is shown, the number of recording devices can be one, or multiple recording devices can be provided for each object.

[0106] exist Figure 3 In the example shown, sensor signals acquired by a recording device 51 installed or built into an object (e.g., a moving body) are transmitted as data to a learning device 52 and used for learning by the object sound source generator 11.

[0107] The recording device 51 includes a microphone 61, a motion measurement unit 62, a position measurement unit 63, a recording unit 64, and a transmission unit 65.

[0108] Microphone 61 receives ambient sound from the recording device 51 and provides the microphone recording signal (i.e., the sensor signal acquired therefrom) to the recording unit 64. Furthermore, the microphone recording signal can be a mono signal or a multi-channel signal.

[0109] For example, the motion measurement unit 62 is formed by sensors (such as 9-axis sensors or geometric sensors, accelerometers and gyroscopes) for measuring the movement and orientation of an object, and outputs sensor signals representing its measurement results (sensed values) to the recording unit 64 as motion information.

[0110] Specifically, when microphone 61 is used to receive sound, motion measurement unit 62 measures the movement and orientation of the object and outputs motion information representing the result. Furthermore, the motion information can be a single sensor signal, or it can be sensor signals from multiple sensors of different types.

[0111] The position measurement unit 63 is formed, for example, by a range sensor and a positioning sensor (such as a GPS device and an indoor range signal receiver), measures the position of the object on which the recording device 51 is installed, and outputs the sensor signal representing the measurement result as position information to the recording unit 64.

[0112] Furthermore, the location information can be a single sensor signal, or it can be sensor signals from multiple sensors of different types. More specifically, the location represented by the location information, for example, is set as coordinate information of a predetermined location within a reference recording space (such as a game venue or theater).

[0113] Simultaneously acquire microphone recording signals, motion information, and location information within the same time period.

[0114] The recording unit 64 performs appropriate analog-to-digital (AD) conversion on the microphone recording signal provided from the microphone 61, the motion information provided from the motion measurement unit 62, and the position information provided from the position measurement unit 63, and provides the microphone recording signal, motion information, and position information to the transmission unit 65.

[0115] The transmission unit 65 generates transmission data including the microphone recording signal, motion information and location information by performing compression processing on the microphone recording signal, motion information and location information provided from the recording unit 64, and transmits the transmission data to the learning device 52 via a network or the like.

[0116] Furthermore, image sensors and the like can be arranged in the recording device 51, and video signals and the like can be included in the transmitted data as sensor signals for learning. The learning data can be multiple sensor signals acquired by multiple sensors of different types, or it can be a single sensor signal acquired by a single sensor.

[0117] The learning device 52 includes an acquisition unit 81, a sensor database 82, a microphone recording database 83, a learning unit 84, and a coefficient database 85.

[0118] The acquisition unit 81 acquires transmission data by receiving transmission data transmitted from the recording device 51, etc., and acquires microphone recording signals, motion information and position information by appropriately performing decoding processing on the acquired transmission data.

[0119] When necessary, the acquisition unit 81 performs signal processing on the microphone recording signal to extract the sound from the target object sound source, and provides the microphone recording signal, which is set as teaching data (correct answer data) during learning, to the microphone recording database 83. This teaching data is the object sound source signal generated by the object sound source generator 11.

[0120] In addition, the acquisition unit 81 provides the motion information and location information extracted from the transmitted data to the sensor database 82.

[0121] The sensor database 82 records the motion and location information provided by the acquisition unit 81, and appropriately provides the provided information as learning data to the learning unit 84.

[0122] The microphone recording database 83 records the microphone recording signals provided by the acquisition unit 81 and appropriately provides the recorded microphone recording signals to the learning unit 84 as teaching data.

[0123] Learning unit 84 performs machine learning based on motion and location information provided from sensor database 82 and microphone recording signals provided from microphone recording database 83, generating object sound source generator 11 (more specifically, configuring coefficient data for object sound source generator 11), and provides the generated coefficient data to coefficient database 85. Coefficient database 85 records the coefficient data provided by learning unit 84.

[0124] Although object sound source generators generated through machine learning can be Figure 2 The object sound generator 11 shown may be another object sound generator with a configuration such as a DNN, but in the following description, we will continue to describe it by assuming that the object sound generator 11 is generated.

[0125] In this case, coefficient data formed by the coefficients processed by arithmetic operations (signal processing) for each of the sound source type identification unit 21, envelope estimation unit 23 and mixing unit 24 is generated by machine learning and provided to coefficient database 85.

[0126] Although the teaching data during learning may be a microphone recording signal acquired by microphone 61, the teaching data during learning is assumed to be a microphone recording signal (acoustic signal), which includes only the sound of the target object sound source (obtained by removing sounds other than those that are the target, such as inappropriate speech, i.e., unnecessary sounds that are part of the content) from the microphone recording signal.

[0127] More specifically, in the case of football content, for example, rusting sounds, collision noises, etc. are removed as unnecessary sounds, and the sounds of kicking the ball, speech, breathing, running, etc. are extracted as sounds for each target sound source type (i.e., the sound of the object sound source), and the acoustic signals of the extracted sounds are set as teaching data.

[0128] Furthermore, when generating microphone recording signals used as teaching data for each sound source type, the target sound and unnecessary sound can be changed according to each content and situation (such as the country to which the content is delivered), and can be set appropriately.

[0129] Furthermore, for example, the process of removing unwanted sounds from the microphone recording signal can be implemented through specific processes such as sound source separation using a DNN, or any other sound source separation. Additionally, in the process of removing unwanted sounds from the microphone recording signal, motion information and position information acquired by another recording device 51, the microphone recording signal, etc., can be used.

[0130] Furthermore, the teaching data is not limited to microphone recorded signals and acoustic signals generated from microphone recorded signals, but can be any data such as pre-generated typical kicking sounds, as long as it is an acoustic signal (target signal) of a target sound related to the object (object sound source) corresponding to the microphone recorded signal.

[0131] Furthermore, the learning device 52 can record learning and teaching data for each individual, such as recording each contestant or performer as an object, and generating coefficient data for the object sound source generator 11, while also taking into account the individuality of each person (each on which the recording device 51 is mounted). Conversely, by using learning and teaching data acquired for multiple objects, coefficient data for a universal object sound source generator 11 can be generated.

[0132] Furthermore, when the object sound source generator is configured using DNNs, a DNN can be prepared for each object sound source, and the sensor signal used as input can be different for each DNN.

[0133] <Description of record processing during the generation of training data>

[0134] Then, the description Figure 3 The operation of the recording device 51 and the learning device 52 shown in the figure.

[0135] First, refer to Figure 4 The flowchart shown describes the recording process performed by the recording device 51 when generating learning data.

[0136] In step S11, the microphone 61 receives the ambient sound of the recording device 51 and provides the resulting microphone recording signal to the recording unit 64.

[0137] In step S12, the recording unit 64 acquires the sensor signals output from the motion measurement unit 62 and the position measurement unit 63 as motion information and position information.

[0138] Recording unit 64 performs AD conversion and other operations on the microphone recording signal, motion information and location information acquired in the above process as needed, and provides the acquired microphone recording signal, motion information and location information to transmission unit 65.

[0139] Furthermore, the transmission unit 65 generates transmission data formed from the microphone recording signal, motion information, and position information provided by the recording unit 64. At this time, the transmission unit 65 performs compression processing on the microphone recording signal, motion information, and position information as needed.

[0140] In step S13, the transmission unit 65 transmits the transmission data to the learning device 52, and the recording process ends. Furthermore, the transmission data can be transmitted sequentially in real time (online), or all transmission data can be transmitted offline after recording.

[0141] As described above, the recording device 51 acquires not only the microphone recording signal, but also motion information and location information, and transmits them to the learning device 52. In this way, the learning device 52 can obtain an object sound source generator for acquiring the microphone recording signal from the motion information and location information, and as a result, obtain a target sound with high quality.

[0142] <Description of learning processing>

[0143] Next, we will refer to Figure 5 The flowchart shown describes the learning process performed using the learning device 52.

[0144] In step S41, the acquisition unit 81 acquires transmission data by receiving transmission data transmitted from the recording device 51. Furthermore, it performs decoding processing on the acquired transmission data as needed. Additionally, the transmission data may be acquired from a removable recording medium or the like, without being received via a network.

[0145] In step S42, the acquisition unit 81 marks the acquired transmission data.

[0146] For example, the acquisition unit 81 performs a process that associates each time segment of the microphone recording signal, motion information, and location information configured for transmission data with sound source type information of the time segment indicating that such a time segment is a sound source of a specific sound source type, as a tagging process.

[0147] Furthermore, the association between each time segment and the sound source type can be manually entered by the user, or signal processing such as sound source separation can be performed based on microphone recording signals, motion information, and location information, and thus sound source type information can be obtained.

[0148] The acquisition unit 81 provides the tagged (in other words, associated with sound source type information) microphone recording signals to the microphone recording database 83, so that the microphone recording signals are recorded. It also provides each sensor signal with tagged configuration movement information and location information to the sensor database 82, and so that the sensor signals are recorded.

[0149] During learning, each sensor signal with motion and location information acquired in this manner is used as learning data, and the microphone recording signal is used as teaching data.

[0150] Furthermore, as described above, the teaching data can be configured to be a signal obtained by eliminating unwanted sounds from the microphone recording signal. In this case, for example, the acquisition unit 81 performs arithmetic operations by inputting the microphone recording signal, motion information, and position information into a pre-acquired DNN, and the microphone recording signal, which serves as the teaching data, can be acquired as the output.

[0151] Furthermore, when using the object sound source generator 11, the microphone recording signal can be acquired as transmission data, the microphone recording signal, motion information, and position information configured for transmission data can be set as learning data, and the signal acquired by eliminating unnecessary sounds from the microphone recording signal can be set as teaching data.

[0152] In the learning device 52, a large amount of learning and teaching data acquired at multiple different times are recorded in the sensor database 82 and the microphone recording database 83.

[0153] Furthermore, learning data and teaching data can be acquired for the same recording device 51 (i.e., the same object), or for multiple different recording devices 51 (objects).

[0154] Furthermore, as learning data corresponding to the teaching data acquired by the predetermined recording device 51, not only the movement information and location information acquired by the predetermined recording device 51, but also the movement information and location information acquired by another recording device 51, video signals acquired by cameras other than the recording device 51, etc., can be configured for use.

[0155] In step S43, the learning unit 84 performs machine learning based on the learning data recorded in the sensor database 82 and the teaching data recorded in the microphone recording database 83, thereby generating coefficient data for configuring the object sound source generator 11.

[0156] Learning unit 84 provides the coefficient data acquired in this manner to coefficient database 85, and the coefficient data is recorded, and the learning process ends.

[0157] In this way, the learning device 52 performs machine learning using the transmitted data acquired from the recording device 51, thereby generating coefficient data. With this configuration, high-quality target sound can be obtained even when a microphone recording signal cannot be obtained using the object sound source generator 11.

[0158] <Configuration Examples of Recording Devices and Sound Generating Devices>

[0159] Subsequently, Figure 6 The diagram illustrates a configuration example of a recording device and a sound source generating device for recording actual content during gameplay or actual performance, and generating an object sound source signal based on the recording results. It is important to note that... Figure 6 In, and in Figure 3 In cases where components correspond to each other, the same reference symbols are used, and their descriptions are appropriately omitted.

[0160] exist Figure 6 In this process, sensor signals acquired by a recording device 111 installed or built into an object (such as a moving body) are transmitted as transmission data to a sound source generator 112, and the object sound source signal is generated by the object sound source generator 112.

[0161] The recording device 111 includes a motion measurement unit 62, a position measurement unit 63, a recording unit 64, and a transmission unit 65.

[0162] The configuration of recording device 111 differs from that of recording device 51 in that it does not include microphone 61, but is otherwise identical to that of recording device 51.

[0163] Therefore, the transmission data transmitted (output) by the transmission unit 65 of the recording device 111 is formed from motion information and position information, and does not include microphone recording signals. Furthermore, by including an image sensor or the like in the recording device 111, video signals or the like can be included in the transmission data as sensor signals.

[0164] The sound source generating device 112 includes an acquisition unit 131, a coefficient database 132, and an object sound source generating unit 133.

[0165] The acquisition unit 131 acquires coefficient data from the learning device 52 via a network or the like, provides the acquired coefficient data to the coefficient database 132, and records the coefficient data. Alternatively, the coefficient data may not be acquired from the learning device 52, but rather from another device via a wired or wireless means, or from a removable recording medium or the like.

[0166] Additionally, the acquisition unit 131 receives transmission data transmitted from the recording device 111 and provides the received transmission data to the object sound source generating unit 133. Furthermore, the acquisition unit 131 performs appropriate decoding processing on the received transmission data.

[0167] Furthermore, data acquisition from the recording device 111 can be performed in real time (online), and data acquisition can also be performed offline after recording. Alternatively, instead of receiving data directly from the recording device 111, data can be acquired from a removable recording medium or the like.

[0168] The object sound source generating unit 133 is used, for example, as a device that performs arithmetic operations based on coefficient data provided from the coefficient database 132. Figure 2 The object sound source generator 11 is shown.

[0169] In other words, the object sound source generating unit 133 generates an object sound source signal based on the coefficient data provided from the coefficient database 132 and the transmission data provided from the acquisition unit 131, and outputs the generated object sound source signal to the later stage.

[0170] <Description of the recording processes performed when generating an object sound source>

[0171] The operation of the recording device 111 and the sound source generating device 112, which are performed when recording content, in other words, when generating an object sound source signal, will then be described.

[0172] First, refer to Figure 7 The flowchart shown describes the recording process performed by the recording device 111 when generating an object sound source.

[0173] When the recording process begins, although movement information and location information are acquired in step S71 and transmission data is transmitted in step S72, this processing is different from... Figure 4 The processes in steps S12 and S13 shown are similar, so their descriptions will be omitted.

[0174] Here, the transmission data transmitted in step S72 includes only motion information and location information, and does not include microphone recording signals.

[0175] When data is transmitted in this manner, the recording process ends.

[0176] The recording device 111 acquires the movement information and position information as described above, and transmits the acquired information to the sound source generating device 112. In this way, the sound source generating device 112 can acquire the object sound source signal corresponding to the microphone recording signal from the movement information and position information. In other words, a target sound with high quality can be obtained.

[0177] <Description of Sound Source Generation and Processing>

[0178] Then, refer to Figure 8 The flowchart shown describes the sound source generation process performed using the sound source generation device 112.

[0179] Furthermore, at the start of the sound source generation process, it is assumed that the coefficient data is pre-acquired and recorded in the coefficient database 132.

[0180] In step S101, the acquisition unit 131 acquires the transmission data received from the recording device 111 and provides the transmission data to the object sound source generating unit 133. Furthermore, decoding processing is performed on the acquired transmission data as needed.

[0181] In step S102, the object sound source generating unit 133 generates an object sound source signal based on the coefficient data provided from the coefficient database 132 and the transmission data provided from the acquisition unit 131.

[0182] For example, the object sound source generating unit 133 is used as an object sound source generator 11 that is pre-acquired by learning through coefficient data obtained from the coefficient database 132.

[0183] In this case, the sound source type identification unit 21 performs arithmetic operations based on multiple sensor signals and the motion information and location information of the transmission data provided by the acquisition unit 131 as configuration data, and learns the pre-acquired coefficients, and provides the sound source type information obtained therefrom to the sound source database 22.

[0184] The sound source database 22 provides one or more dry source signals of various sound source types, represented by the sound source type information provided by the sound source type identification unit 21, to the mixing unit 24.

[0185] Envelope estimation unit 23 extracts the envelope of the sensor signal of the accelerometer included in the motion information constituting the transmission data provided by acquisition unit 131, and provides the envelope information representing the extracted envelope to mixing unit 24.

[0186] The mixing unit 24 performs signal processing such as modulation and filtering on one or more dry source signals provided from the sound source database 22 based on the envelope information provided from the envelope estimation unit 23, thereby generating an object sound source signal.

[0187] In addition, the object sound source generating unit 133 can also generate metadata for each object sound source signal while generating the object sound source signal for each object sound source.

[0188] In this case, for example, the object sound source generating unit 133 generates metadata including sound source type information obtained by the sound source type identification unit 21, movement information and location information obtained as transmission data, etc., based on the transmission data.

[0189] In step S103, the object sound source generating unit 133 outputs the generated object sound source signal, and the sound source generation process ends. At this time, metadata can be configured to be output along with the object sound source signal.

[0190] As described above, the sound source generating device 112 generates an object sound source signal based on coefficient data and transmission data and outputs the object sound source signal.

[0191] In this way, even if the microphone recording signal cannot be acquired by the recording device 111, the target sound with high quality, that is, the object sound source signal with high quality, can be obtained from the sensor signals of other sensors besides the microphone.

[0192] In other words, the object sound source signal can be robustly obtained from sensor signals of a small number of sensor types. Furthermore, the object sound source signal can be obtained, which includes only the target sound and excludes unwanted sounds such as noise and speech.

[0193] <Second Implementation Method>

[0194] <Configuration Example of Learning Device>

[0195] As described above, the teaching data for learning the object sound source generator 11 can be set to a signal obtained by eliminating unnecessary sounds from the microphone recording signal, that is, a signal obtained by extracting only the target sound. Furthermore, in addition to the microphone recording signal, motion information and position information can also be used when generating the teaching data.

[0196] For example, when mobility and location information are also used to generate teaching data, such as Figure 9 The learning device is configured as shown. It should be noted that... Figure 9 In, and in Figure 3 In cases where components correspond to each other, the same reference symbols are used, and their descriptions are appropriately omitted.

[0197] exist Figure 9 The learning device 161 shown includes an acquisition unit 81, a segment detection unit 171, an integration unit 172, a signal processing unit 173, a sensor database 82, a microphone recording database 83, a learning unit 84, and a coefficient database 85.

[0198] The configuration of learning device 161 differs from that of learning device 52 in that it adds segment detection unit 171 to signal processing unit 173, and is otherwise similar to that of learning device 52.

[0199] When acquiring transmission data, the acquisition unit 81 provides the microphone recording signal constituting the transmission data to the segment detection unit 171, the integration unit 172, the signal processing unit 173, and the microphone recording database 83.

[0200] In addition, the acquisition unit 81 provides the movement information constituting the transmission data to the segment detection unit 171, the integration unit 172, the signal processing unit 173, and the sensor database 82, and provides the position information constituting the transmission data to the integration unit 172, the signal processing unit 173, and the sensor database 82.

[0201] The segment detection unit 171 detects the type of sound of the object sound source included in the microphone recording signal and the time segment of the sound source based on the microphone recording signal and motion information provided by the acquisition unit 81, and provides the sound source type segment information representing the detection result to the integration unit 172.

[0202] For example, by performing thresholding on the microphone recording signal, performing arithmetic operations by feeding the microphone recording signal and motion information into a recognition unit such as a DNN, or performing delay-sum beamforming (DS) or empty beamforming (NBF) on the microphone recording signal, the segment detection unit 171 identifies the type of object sound source included in each time segment and generates sound source type segment information.

[0203] As a concrete example, when a sensor signal representing a small displacement of an object in the vertical direction, as measured by, for example, an accelerometer, is used as motion information, a time segment of a person's breathing sound, which is the object, can be detected. In this case, for example, a time segment with a sensor signal frequency of approximately 0.5 Hz to 1 Hz is considered as the time segment of the object's breathing sound.

[0204] Furthermore, for example, by using NBF, the component of spoken speech of an object included in the microphone recording signal can be suppressed. In this case, the time segment in which spoken speech of the object is detected from the microphone recording signal before suppression, and the time segment in which spoken speech of the object is not detected from the microphone recording signal after suppression, are considered the final time segment of spoken speech of the object.

[0205] The integration unit 172 generates final sound source type segment information and sound source type information based on the microphone recording signal, motion information and location information provided by the acquisition unit 81 and the sound source type segment information provided by the segment detection unit 171, and provides the generated information to the signal processing unit 173.

[0206] Specifically, the integration unit 172 generates sound source type segment information of the target object based on at least one of the microphone recording signal, motion information, and position information of the object being processed (hereinafter also referred to as the target object), and the microphone recording signal, motion information, and position information of another object.

[0207] In this case, for example, the integration unit 172 performs position information comparison processing, time segment integration processing, and segment smoothing processing to generate the final sound source type segment information.

[0208] In the location information comparison processing, the distance between the target object and another object is calculated based on the location information of each object, and another object (i.e., other objects that exist near the target object) that may be affected by the sound source of the target object is selected as the reference object based on the obtained distance.

[0209] Next, in the time segment integration process, it is determined whether there are any objects selected as reference objects.

[0210] Then, in the absence of an object selected as a reference object, the sound source type segment information of the target object acquired by the segment detection unit 171 is directly output to the signal processing unit 173 as the final sound source type segment information. This is because, if another object is not present near the target object, the sound of that other object will not be mixed into the microphone recording signal.

[0211] On the other hand, when there is an object selected as a reference object, the position and movement information of this reference object are also used to generate the final sound source type segment information of the target object.

[0212] More specifically, among the reference objects, the reference object that has a segment that overlaps with the time segment represented by the sound source type segment information of the target object as the time segment of the sound source of the object is selected as the final reference object.

[0213] Then, based on the position and movement information of the reference object and the position and movement information of the target object, relative orientation information representing the relative direction of the reference object as seen from the target object in three-dimensional space is generated.

[0214] Furthermore, an NBF filter is formed based on the orientation of the target object, represented by the target object's position and movement information, and the relative orientation information of each reference object. Additionally, convolution processing is performed between the NBF filter and a time segment represented by the target object's sound source type segment information in the microphone recording signal.

[0215] Subsequently, sound source type segment information is generated by using a process similar to that performed by segment detection unit 171 based on the signal obtained through convolution processing and the movement information of the target object. This allows for the suppression of sound emitted from the reference object and enables the acquisition of more accurate sound source type segment information.

[0216] Finally, by performing segment smoothing processing on the sound source type segment information obtained through time segment integration processing, the final sound source type segment information is obtained.

[0217] For example, for each type of object sound source, when the sound of that type of object sound source is generated, the average time for the sound to continue at a minimum is obtained in advance as the average minimum duration.

[0218] In segment smoothing, smoothing is performed using a smoothing filter that connects the time segments of the sound from the divided object sound source, such that the length of the time segment in which the sound from the object sound source is detected becomes equal to or longer than the average minimum duration.

[0219] In addition, the integration unit 172 generates sound source type information from the acquired sound source type segment information and provides the sound source type segment information and sound source type information to the signal processing unit 173.

[0220] As described above, the integration unit 172 can eliminate information about the sound of another object that has not been eliminated (excluded) from the sound source type segment information obtained by the segment detection unit 171, and obtain sound source type segment information with higher precision.

[0221] The signal processing unit 173 performs signal processing based on the microphone recording signal, motion information and position information provided by the acquisition unit 81 and the sound source type segment information provided by the integration unit 172, thereby generating an object sound source signal that can be used as teaching data.

[0222] For example, the signal processing unit 173 performs sound quality correction processing, sound source separation processing, noise cancellation processing, distance correction processing, sound source replacement processing, or processing obtained by combining multiple processes such as signal processing.

[0223] More specifically, for example, performing noise suppression processing (such as filtering or gain correction to suppress noise-dominant frequency bands), silencing noisy or unnecessary sections, and adding filtering to frequency components that may be easily attenuated, as sound quality correction processing.

[0224] For example, as a sound source separation process, the following processing is performed: based on the sound source type segment information, the time segment in which the sound of multiple object sound sources is included in the microphone recording signal is identified, and based on the identification results, the independent component analysis of the sound of each object sound source is separated according to the differences between the amplitude values ​​or probability density distributions of each type of object sound source.

[0225] Furthermore, for example, in a microphone recording signal, when the time interval of the sound from an object source includes unwanted sounds (such as stable noise, mainly background noise, cheers, etc.) and noise such as wind, noise suppression processing is performed in the time interval as noise cancellation processing.

[0226] Furthermore, for example, for the absolute sound pressure level of the sound emitted by the object sound source, the process of correcting the distance attenuation from the object sound source to the microphone 61 during recording and the effect of convolution of the transmission characteristics are performed as distance correction processing.

[0227] In this case, for example, during distance correction processing, a process is performed to add the inverse characteristics of the transmission characteristics from the object sound source to the microphone 61 to the microphone recording signal. In this way, the sound quality degradation of the object sound source due to distance attenuation, transmission characteristics, etc., can be corrected.

[0228] In addition, for example, the process of replacing a segment of a microphone-recorded signal with another sound signal that is either pre-prepared or dynamically generated based on the segment information of the sound source type is performed as a sound source replacement process.

[0229] The signal processing unit 173 provides the microphone recording database 83 with the object sound source signal, which is the teaching data acquired as described above, and the sound source type information provided from the integration unit 172 in association with each other.

[0230] In addition, the signal processing unit 173 provides the sound source type information from the integration unit 172 to the sensor database 82, and records the sound source type information, movement information and location information in association with each other.

[0231] In the learning device 161, sound source type information is generated, and the object sound source signal used as teaching data and the movement information and position information used as learning data are associated with the sound source type information, so that users do not need to perform operation input for the above-mentioned marking processing.

[0232] Furthermore, the microphone recording database 83 can provide the object sound source signal provided by the signal processing unit 173 to the learning unit 84 as teaching data, and can also provide the microphone recording signal provided by the acquisition unit 81 to the learning unit 84 as teaching data. Additionally, the microphone recording signal can be used as learning data.

[0233] <Description of learning processing>

[0234] Next, we will refer to Figure 10 The flowchart shown describes the learning process performed using the learning device 161.

[0235] In step S131, the acquisition unit 81 acquires the transmission data from the recording device 111.

[0236] The acquisition unit 81 provides the microphone recording signal constituting the transmission data to the segment detection unit 171, the integration unit 172, the signal processing unit 173, and the microphone recording database 83.

[0237] In addition, the acquisition unit 81 provides the movement information constituting the transmission data to the segment detection unit 171, the integration unit 172, the signal processing unit 173, and the sensor database 82, and provides the position information constituting the transmission data to the integration unit 172, the signal processing unit 173, and the sensor database 82.

[0238] In step S132, the segment detection unit 171 generates sound source type segment information based on the microphone recording signal and motion information provided by the acquisition unit 81, and provides the generated sound source type segment information to the integration unit 172.

[0239] For example, the segment detection unit 171 performs threshold processing of the microphone recording signal, arithmetic operation processing based on a recognition unit such as DNN, etc., thereby identifying the type of object sound source included in each time segment and generating sound source type segment information.

[0240] In step S133, the integration unit 172 generates final sound source type segment information and sound source type information by integrating the microphone recording signal, motion information and position information provided by the acquisition unit 81 and the sound source type segment information provided by the segment detection unit 171, and provides the generated sound source type segment information and sound source type information to the signal processing unit 173.

[0241] For example, in step S133, the above-mentioned location information comparison processing, time segment integration processing, and segment smoothing processing are performed to integrate the information and generate the final sound source type segment information.

[0242] In step S134, the signal processing unit 173 performs signal processing based on the microphone recording signal, motion information and position information provided by the acquisition unit 81 and the sound source type segment information provided by the integration unit 172, and generates an object sound source signal to be used as teaching data.

[0243] For example, the signal processing unit 173 performs sound quality correction processing, sound source separation processing, noise cancellation processing, distance correction processing, or sound source exchange processing, or processing obtained by combining multiple such processes such as signal processing, and generates an object sound source signal.

[0244] The signal processing unit 173 associates the object sound source signal used as teaching data with the sound source type information provided from the integration unit 172 and provides it to the microphone recording database 83, where it records them. In addition, the signal processing unit 173 provides the sound source type information provided from the integration unit 172 to the sensor database 82, and records the sound source type information, motion information, and position information in association with each other.

[0245] In step S135, the learning unit 84 performs machine learning based on the motion and location information recorded in the sensor database 82 as learning data and the object sound source signals recorded in the microphone recording database 83 as teaching data.

[0246] Learning unit 84 provides the coefficient data obtained through machine learning to the coefficient database 85, records the coefficient data therein, and the learning process ends.

[0247] As described above, the learning device 161 uses the transmitted data acquired from the recording device 51 to generate teaching data and perform machine learning, thereby generating coefficient data. In this way, even when a microphone recording signal cannot be acquired, the object sound source generator 11 can be used to acquire target sound with high quality.

[0248] <Third Implementation Method>

[0249] <Configuration Example of Learning Device>

[0250] However, when acquiring the sound source signal of the target object, there are cases where the microphone recording signal of the sound source can be acquired even though the sound source does not have a high signal-to-noise ratio.

[0251] As an example, one could consider a situation where the microphone recording signal could be acquired by a recording device 111 installed on an object other than the target object (such as another player besides the target player, a referee, etc.), where a separate microphone synchronized with the recording device 111 is used to record speech. Furthermore, it could also be considered that the microphones of the high-priority recording device 111 installed on the target object are out of order, and the microphone recording signal cannot be acquired.

[0252] Therefore, the object sound source signal of the target object's sound source can be configured to also be generated by using a microphone recording signal acquired by a recording device 111 installed in another object different from the target object.

[0253] In this case, for example, such as Figure 11 As shown, a learning device is configured. It should be noted that in... Figure 11 In, corresponding to Figure 3 Those parts are indicated by the same reference numerals, and descriptions of those parts will be omitted appropriately.

[0254] exist Figure 11 The learning device 201 shown includes an acquisition unit 81, a correction processing unit 211, a sensor database 82, a microphone recording database 83, a learning unit 84, and a coefficient database 85. The configuration of the learning device 201 is a new arrangement of the correction processing unit 211 within the configuration of the learning device 52.

[0255] In this example, the correction processing unit 211 provides the microphone recording signal, which has been acquired by the target recording device 51 from the acquisition unit 81, as teaching data to the microphone recording database 83.

[0256] In addition, the location information obtained by the target recording device 51 and the location information and microphone recording signal obtained by the other recording device 51 are provided to the correction processing unit 211.

[0257] Here, an example using position information acquired by another recording device 51 and a microphone recording signal will be described. However, the configuration is not limited to this, and the position information and microphone recording signal used by the correction processing unit 211 may be the position information of a microphone not mounted on the target object and the microphone recording signal acquired by the microphone.

[0258] The correction processing unit 211 performs correction processing on the microphone recording signal of the other recording device 51 based on the position information of the target recording device 51 and the position information of the other recording device 51, and provides the microphone recording signal obtained therefrom to the microphone recording database 83 as learning data for the target recording device 51, and records the microphone recording signal therein.

[0259] In the calibration process, processing corresponding to the positional relationship between the target recording device 51 and the other recording device 51 is performed.

[0260] More specifically, for example, in the correction process, based on the distance between the target recording device 51 and another recording device 51 obtained from location information, a process of moving the microphone recording signal in the time direction is performed, such as compensation for the propagation delay of the microphone recording signal.

[0261] Additionally, as a correction process, corrections related to the sound transmission characteristics corresponding to the relative orientation (direction) of another object seen from the target object can be performed based on the movement information of the target recording device 51 and the movement information of other recording devices 51.

[0262] The learning unit 84 performs machine learning, which uses the position and movement information of the target recording device 51 and the microphone recording signal after correction processing by the other recording device 51 as learning data, and uses the microphone recording signal acquired by the target recording device 51 as teaching data.

[0263] In other words, the coefficient data of the object sound source generator 11 is generated by machine learning. The object sound source generator 11 has position information, movement information and microphone recording signal after correction processing as input, and microphone recording signal as teaching data as output.

[0264] Furthermore, in this case, the number of other recording devices 51 can be one or more. Additionally, the movement and position information acquired by the other recording devices 51 can also be used as learning data for the target object (recording device 51).

[0265] Furthermore, the teaching data for the target object can be configured to be generated based on the microphone recording signal after correction processing by another recording device 51.

[0266] <Description of learning processing>

[0267] Next, we will refer to Figure 12 The flowchart shown describes the learning process performed using the learning device 201.

[0268] It should be noted that the processing in steps S161 and S162 is related to... Figure 5 The processes in steps S41 and S42 are the same, therefore, their descriptions are omitted.

[0269] Here, in step S162, the tagged movement information and location information are provided to the sensor database 82, and the location information is provided to the correction processing unit 211. Additionally, the tagged microphone recording signal is provided to the correction processing unit 211.

[0270] The correction processing unit 211 provides the microphone recording signal marked by the target recording device 51 provided by the acquisition unit 81 directly to the microphone recording database 83 as teaching data, and records the microphone recording signal therein.

[0271] In step S163, the correction processing unit 211 performs correction processing on the microphone recording signal of the other recording device 51 provided by the acquisition unit 81 based on the position information of the target recording device 51 and the position information of the other recording device 51 provided by the acquisition unit 81.

[0272] The correction processing unit 211 supplies the microphone recording signal obtained by the correction processing to the microphone recording database 83 as learning data for the target recording device 51 and records the microphone recording signal therein.

[0273] When the correction process is performed, learning is then performed in step S164, and the learning process ends. The process in step S164 is then... Figure 5 The process of step S43 shown in the figure is similar, so its description will be omitted.

[0274] Here, in step S164, not only the position information and movement information of the target recording device 51 are used, but also the microphone recording signal after correction processing of another recording device 51 is used as learning data to perform machine learning.

[0275] As described above, the learning device 201 also uses the microphone recording signal of another recording device 51, which has already undergone correction processing, as learning data to perform machine learning and generate coefficient data.

[0276] In this way, even when a microphone recording signal cannot be obtained, the object sound generator 11 can be used to obtain a target sound with high quality. Specifically, in this example, since the microphone recording signal obtained by another recording device can be used as the input to the object sound generator 11, a target sound with higher quality can be obtained.

[0277] <Configuration Example of Sound Source Generator>

[0278] In addition, for example, such as Figure 13 As shown, a sound source generating device is configured to generate an object sound source signal using coefficient data acquired by the learning device 201. It should be noted that the same reference numerals will be applied to... Figure 6The situation in the middle corresponds to Figure 13 The part in the text will be omitted, and its description will be appropriately omitted.

[0279] Figure 13 The sound source generating device 241 shown includes an acquisition unit 131, a coefficient database 132, a correction processing unit 251, and an object sound source generating unit 133.

[0280] The configuration of the sound source generating device 241 is a new configuration of the correction processing unit 251 in the structure of the sound source generating device 112.

[0281] The acquisition unit 131 of the sound source generating device 241 acquires not only the transmission data from the target recording device 111, but also the transmission data from another recording device 111. In this case, the microphone 61 is arranged in the other recording device 111, and the microphone recording signal is also included in the transmission data from the other recording device 111.

[0282] The correction processing unit 251 does not perform correction processing on the transmission data provided by the acquisition unit 131 obtained from the target recording device 111, and directly provides the provided transmission data to the object sound source generating unit 133.

[0283] On the other hand, the correction processing unit 251 performs correction processing on the microphone recording signal, which includes the transmission data acquired from another recording device 111 and provided from the acquisition unit 131, and provides the microphone recording signal to the object sound source generating unit 133 after the correction processing. Furthermore, similar to the case of the correction processing unit 211, the position information and microphone recording signal used by the correction processing unit 251 may be the position information of a microphone not mounted on the target object and the microphone recording signal acquired by the microphone.

[0284] <Description of Sound Source Generation and Processing>

[0285] Then, refer to Figure 14 The flowchart shown describes the sound source generation process performed using the sound source generation device 241.

[0286] The process in step S191 is similar to Figure 8 The process of step S101 shown will therefore be omitted from description.

[0287] At this time, the correction processing unit 251 will directly provide the transmission data of the target recording device 111 from the transmission data provided by the acquisition unit 131 to the object sound source generating unit 133.

[0288] In step S192, the correction processing unit 251 performs correction processing on the microphone recording signal contained in the transmission data acquired from another recording device 111 and provided from the acquisition unit 131, and provides the microphone recording signal to the object sound source generating unit 133 after the correction processing.

[0289] For example, in step S192, based on the position information of the target recording device 111 and the position information of the other recording device 111, the following is performed: Figure 12 The correction process for step S163 shown in the figure is similar to the correction process.

[0290] When the correction process is complete, steps S193 and S194 are then executed, and the sound source generation process ends. This process is similar to... Figure 8 The processes of steps S102 and S103 shown are similar, so their descriptions will be omitted.

[0291] Here, in step S193, for example, not only the position information and movement information of the target recording device 111, but also the microphone recording signal after correction processing provided by the correction processing unit 251, are input to the sound source type identification unit 21 of the object sound source generator 11, and arithmetic operations are performed.

[0292] In other words, the object sound source generating unit 133 generates an object sound source signal corresponding to the microphone recording signal of the target recording device 111 based on the position and movement information of the target recording device 111, the microphone recording signal after correction processing by another recording device 111, and coefficient data.

[0293] As described above, the sound source generating device 241 also uses the microphone of another recording device 111, which has already undergone correction processing, to record signals to generate object sound source signals and outputs the generated object sound source signals.

[0294] In this way, even if the target recording device 111 cannot acquire the microphone recording signal, the microphone recording signal acquired by another recording device 111 can be used to acquire the target sound with high quality, that is, the object sound source signal with high quality.

[0295] Furthermore, for example, if the target recording device 111 or the like can only acquire a microphone recording signal with a low signal-to-noise ratio (SN ratio), the object sound source signal can be configured to be generated by the sound source generating device 241 using a microphone recording signal acquired by another recording device 111. In this case, an object sound source signal with a high SN ratio (i.e., higher quality) can be obtained.

[0296] <Fourth Implementation Method>

[0297] <Description of learning processing>

[0298] For example, in recording content in challenging environments such as motion, it is also possible to consider situations where some sensor signals used to generate the sound source signal of an object are not acquired due to any reason such as the failure of sensors such as the point measurement unit 62 and the position measurement unit 63 installed in the recording device 111.

[0299] Therefore, in the case of actually generating an object sound source signal, under the premise that some of the sensor signals among the multiple sensor signals are not available (insufficient), the object sound source signal can be configured to be acquired using the available sensor signals among the multiple sensor signals.

[0300] In this scenario, for all possible combinations of multiple sensor signals that can be acquired, a target sound source generator 11 can be learned in advance, having such sensor signals as input and the target sound source signal as output. In this case, although the estimation reliability decreases, the object sound source signal can be acquired from only some sensor signals.

[0301] For example, when up to N types of sensor signals can be obtained, Σ can be pre-learned. n=1 N N C n 11. Object sound source generator.

[0302] Furthermore, there are cases where it is not feasible to prepare an object sound source generator 11 for all types (combinations) of sensor signals (i.e., all defect modes). In such cases, learning is performed only for modes with insufficient sensor signals due to high-frequency faults, and for modes with insufficient sensor signals due to low-frequency faults, an object sound source generator 11 may not be prepared, or an object sound source generator 11 with a simpler configuration may be prepared.

[0303] In this way, while preparing the object sound source generator 11 (i.e., coefficient data) for the combination of sensor signals, the learning device 52 performs... Figure 15 The learning process is shown in the figure.

[0304] In the following text, reference will be made to Figure 15 The flowchart shown describes the learning process performed using the learning device 52. It should be noted that the processing of steps S221 and S222 is related to... Figure 5 The processes in steps S41 and S42 are the same, therefore, their descriptions are omitted.

[0305] In step S223, the learning unit 84 performs machine learning on the combination of multiple sensor signals constituting the learning data based on the learning data recorded in the sensor database 82 and the teaching data recorded in the microphone recording database 83.

[0306] In other words, in machine learning for combination, coefficient data of an object sound source generator 11 is generated, which produces sensor signals with a predetermined combination of learning data as input and microphone recording signals with teaching data as output.

[0307] Similarly, in this example, the microphone recording signal acquired by the target recording device 51 and the microphone recording signal acquired by another recording device 51 can also be used as learning data.

[0308] The learning unit 84 provides the coefficient data of the configured object sound source generator 11, which is obtained by combining sensor signals, to the coefficient database 85 and records the coefficient data therein, thus ending the learning process.

[0309] As described above, the learning device 52 performs machine learning on a combination of multiple sensor signals and generates coefficient data for configuring the object sound source generator 11 in such a combination. In this way, even when some sensor signals are unavailable, object sound source signals can be robustly generated from the acquired sensor signals.

[0310] <Configuration Example of Sound Source Generator>

[0311] Furthermore, in the case of generating coefficient data from combinations of sensor signals, for example, such as Figure 16 As shown, a sound source generating device is configured. Note that the same reference numerals will be applied to [other devices / familiarities]. Figure 6 The situation in the middle corresponds to Figure 16 The components in the document will be omitted as appropriate.

[0312] exist Figure 16 The sound source generating device 281 shown includes an acquisition unit 131, a coefficient database 132, a fault detection unit 291, and an object sound source generating unit 133.

[0313] The configuration of the sound source generating device 281 is the same as the configuration of the sound source generating device 112, with the addition of a fault detection unit 291.

[0314] The fault detection unit 291 detects faults (defects) of sensors that have acquired sensor signals for each sensor signal that constitutes the transmitted data provided from the acquisition unit 131 based on sensor signals, and provides the detection results to the object sound source generation unit 133.

[0315] Furthermore, based on the fault detection results, the fault detection unit 291 provides the object sound source generation unit 133 only with sensor signals from sensors that are not out of order and are operating normally, which constitute the transmission data supplied from the acquisition unit 131 (in other words, included in the transmission data).

[0316] In addition, the coefficient data generated from the combination of sensor signals is recorded in coefficient database 132.

[0317] The object sound source generating unit 133 reads the coefficient data of the object sound source generator 11 from the coefficient database 132 based on the fault detection result provided by the fault detection unit 291. The object sound source generator 11 has a fault-free sensor signal as input and an object sound source signal as output.

[0318] Then, the object sound source generating unit 133 generates the object sound source signal of the target object sound source based on the readout coefficient data provided by the fault detection unit 291 and the sensor signal of the fault-free sensor.

[0319] Furthermore, in this example, the microphone recording signal can be included as a sensor signal in the transmission data acquired by the acquisition unit 131 from the target recording device 111.

[0320] In this case, for example, if the fault of the microphone corresponding to the microphone recording signal is not detected by the fault detection unit 291, the object sound source generation unit 133 may output the microphone recording signal as is or the signal generated from the microphone recording signal as the object sound source signal.

[0321] On the other hand, if a microphone malfunction has been detected, the object sound source generating unit 133 can generate an object sound source signal based on the sensor signal of a sensor for which no malfunction has been detected, in addition to the microphone recording signal.

[0322] <Description of Sound Source Generation and Processing>

[0323] Then, refer to Figure 17 The flowchart shown describes the sound source generation process performed using the sound source generation device 281.

[0324] The process in step S251 is similar to Figure 8 The process of step S101 shown in the figure will therefore be omitted from the description.

[0325] In step S252, the fault detection unit 291 detects sensor faults for each sensor signal configured to receive transmission data from the acquisition unit 131, and provides the detection results to the object sound source generation unit 133.

[0326] For example, the fault detection unit 291 performs a pure zero check or anomaly detection on the sensor signals, or uses a DNN with sensor signals as input and detection results of fault presence / absence as output, thereby detecting the presence / absence of faults for each sensor signal.

[0327] In addition, the fault detection unit 291 provides only sensor signals from sensors that are not out of order among the sensor signals included in the transmitted data provided from the acquisition unit 131 to the object sound source generation unit 133.

[0328] In step S253, the object sound source generating unit 133 generates an object sound source signal based on the fault detection result provided by the fault detection unit 291.

[0329] That is, the object sound source generating unit 133 receives sensor signals that are not out of order, that is, it reads sensor signals from the coefficient database 132 without using the sensor signals that detected the fault as its input coefficient data to the object sound source generating unit 11.

[0330] Then, the object sound source generating unit 133 generates an object sound source signal based on the read coefficient data and the sensor signal provided by the fault detection unit 291.

[0331] When an object sound source signal is generated, in step S254, the object sound source generating unit 133 outputs the generated object sound source signal, and the sound source generation process ends.

[0332] As described above, the sound source generating device 281 detects the fault of the sensor based on the transmitted data and uses coefficient data based on the detection results to generate an object sound source signal.

[0333] In this way, even if some sensors are out of order for any reason, the object's sound source signal can be robustly generated from the sensor signals of the remaining sensors. In other words, the object's sound source signal can be robustly generated in response to changes in circumstances.

[0334] <Fifth Implementation Method>

[0335] <Description of learning processing>

[0336] Furthermore, the surrounding environmental conditions can differ when recording (acquiring) learning data and teaching data used for learning, as well as when recording actual content.

[0337] The environmental conditions described here include, for example, the surrounding environment of the recording device, such as the absorption coefficient of the floor material, the type of shoes worn by the athlete or performer, the reverberation characteristics of the space where the recording is performed, the type of space where the recording is performed (such as enclosed or open space), volume, 3D shape, weather, ground condition, and the type of ground of the space where the recording is performed (target space).

[0338] When the object sound source generator 11 is learned for each of these different environmental conditions, the object sound source signal with sound quality adapted to the environment at the time of recording can be obtained more closely to the real situation.

[0339] In this case, Figure 3 The learning device 52 shown in the figure performs the function of Figure 18 The learning process is shown in the figure.

[0340] In the following text, reference will be made to Figure 18 The flowchart shown describes the learning process performed using the learning device 52. It should be noted that the processing in steps S281 and S282 is related to... Figure 5 The processes in steps S41 and S42 are the same, so their descriptions will be omitted.

[0341] Here, in step S282, the acquisition unit 81 performs the following processing: it acquires environmental condition information representing the surrounding environment of the recording device 51 using a specific method, and associates not only the sound source type information but also the environmental condition information with the microphone recording signal, movement information and location information as a tag.

[0342] For example, environmental condition information can be configured to be manually entered by a user, can be configured to be obtained through an image recognition process using video signals acquired by a camera, or can be obtained by obtaining weather information from a server via a network.

[0343] In step S283, the learning unit 84 performs machine learning for each environmental condition based on the learning data recorded in the sensor database 82 and the teaching data recorded in the microphone recording database 83.

[0344] In other words, the learning unit 84 performs machine learning using only the learning data and teaching data associated with environmental condition information representing the same environmental condition, and generates coefficient data for configuring the object sound source generator 11 for each environmental condition.

[0345] The learning unit 84 provides the coefficient data for each environmental condition acquired in this manner to the coefficient database 85, records the coefficient data therein, and the learning process ends.

[0346] As described above, the learning device 52 performs machine learning for each environmental condition and generates coefficient data. In this way, object sound source signals with sound quality closer to real sound quality can be obtained using the coefficient data based on the environmental conditions.

[0347] <Configuration Example of Sound Source Generator>

[0348] In the case of generating coefficient data for each environmental condition, for example, the sound source generating device is configured as follows: Figure 19 As shown. Note that the same reference numerals will apply to [other locations / regions]. Figure 6 The situation in the middle corresponds to Figure 19 The part in the text will be omitted, and its description will be appropriately omitted.

[0349] Figure 19 The sound source generating device 311 shown includes an acquisition unit 131, a coefficient database 132, an environmental condition acquisition unit 321, and an object sound source generating unit 133.

[0350] The configuration of the sound source generating device 311 is a new arrangement of the environmental condition acquisition unit 321 within the configuration of the sound source generating device 112. Furthermore, coefficient data is recorded in the coefficient database 132 for each environmental condition.

[0351] The environmental condition acquisition unit 321 acquires environmental condition information (environmental condition) representing the surrounding environment of the recording device 111, and provides the acquired environmental condition information to the object sound source generating unit 133.

[0352] For example, the environmental condition acquisition unit 321 can acquire information representing the environmental condition input by the user and set the information as environmental condition information that does not change, or it can set the weather information (e.g., sunny or rainy) representing the surrounding environment of the recording device 111 acquired from an external server or the like as environmental condition information.

[0353] Alternatively, for example, the environmental condition acquisition unit 321 may acquire video signals with the surroundings of the recording device 111 as the subject, use image recognition of the video signals or DNN operation processing with the video signals as input to identify the ground conditions and ground types around the recording device 111, and generate environmental condition information representing the identification results.

[0354] Here, for example, ground condition is a ground condition determined by weather (such as drought or rain stopping), and ground type is a ground type determined by ground material (such as hard material or grass). Furthermore, environmental conditions are not limited to video signals and can be identified using recognition units such as DNNs based on arithmetic operations on specific observations obtained from thermometers, rain gauges, hygrometers, etc., or through arbitrary signal processing such as image recognition, thresholding, etc.

[0355] The object sound source generating unit 133 reads coefficient data corresponding to the environmental condition information provided by the environmental condition acquisition unit 321 from the coefficient database 132, and generates an object sound source signal based on the read coefficient data and the transmission data provided by the acquisition unit 131.

[0356] <Description of Sound Source Generation and Processing>

[0357] Next, we will refer to Figure 20 The flowchart shown describes the sound source generation process performed using the sound source generating device 311. The processing in step S311 is related to...Figure 8 The process of step S101 shown is similar, and its description will be omitted.

[0358] In step S312, the environmental condition acquisition unit 321 acquires environmental condition information and provides the acquired environmental condition information to the object sound source generation unit 133.

[0359] For example, as described above, the environmental condition acquisition unit 321 acquires environmental condition information by obtaining weather information from an external server and setting the acquired information as environmental condition information, or by performing image recognition on video signals and the like and identifying the environmental condition.

[0360] In step S313, the object sound source generating unit 133 generates an object sound source signal according to the environmental conditions.

[0361] In other words, the object sound source generating unit 133 reads coefficient data generated from the coefficient database 132 for the environmental conditions represented by the environmental condition information provided by the environmental condition acquisition unit 321.

[0362] Then, the object sound source generating unit 133 performs arithmetic operations on the object sound source generator 11 based on the read coefficient data and the transmission data provided by the acquisition unit 131, thereby generating an object sound source signal.

[0363] In step S314, the object sound source generating unit 133 outputs the generated object sound source signal, and the sound source generation process ends.

[0364] As described above, the sound source generating device 311 uses coefficient data corresponding to environmental conditions to generate an object sound source signal. In this way, an object sound source signal with sound quality that is more closely related to the real sound quality and is suitable for environmental conditions can be obtained.

[0365] <Sixth Implementation Method>

[0366] <Configuration Example of Learning Device>

[0367] Furthermore, the quality of the signal-to-noise ratio (SN ratio) of acquired sensor signals can differ from each other when recording (acquiring) learning data and teaching data used for learning, as well as when recording actual content.

[0368] For example, describe situations where football games are recorded as content.

[0369] In this case, the recording of learning data and teaching data for learning is performed during practice, and thus it is likely to obtain sensor signals with a high SN ratio (high quality), such as when the microphone recording signal does not include noise such as cheers or rustling sounds from the surrounding environment as the acquired sensor signal.

[0370] Conversely, since actual content is recorded during gameplay, there is a high probability of obtaining sensor signals with a low signal-to-noise ratio (low quality), where ambient noise such as shouts and rusting sounds are included in the microphone recording signal as the acquired sensor signal.

[0371] In addition, just as the reverberation is smaller during practice and larger during gameplay, the reverberation characteristics of the surrounding space can differ between practice and gameplay.

[0372] Therefore, when recording learning and teaching data, an environment with low reverberation and low noise compared to the recording process is assumed to be formed. In other words, a low-SN environment is assumed when recording content, and machine learning can be performed using learning data that simulates such a low-SN environment. In this way, high-quality object sound source signals can be obtained.

[0373] In this case, for example, the learning device is configured as follows: Figure 21 As shown. In Figure 21 In, corresponding to Figure 3 Those parts are indicated by the same reference numerals and symbols, and their descriptions will be omitted appropriately.

[0374] Figure 21 The learning device 351 shown includes an acquisition unit 81, an overlay processing unit 361, a sensor database 82, a microphone recording database 83, a learning unit 84, and a coefficient database 85.

[0375] The configuration of learning device 351 differs from that of learning device 52 in that a new overlay processing unit 361 is arranged, but otherwise the configuration is the same as that of learning device 52.

[0376] The overlay processing unit 361 does not perform any operation on the microphone recording signal, which is the teaching data provided by the acquisition unit 81, but directly provides the microphone recording signal, which is the teaching data, to the microphone recording database 83, and causes the microphone recording signal to be recorded in the microphone recording database 83.

[0377] Furthermore, the superposition processing unit 361 performs superposition processing by convolving the noise data used to add reverberation and noise with the microphone recording signal, which is the teaching data provided from the acquisition unit 81, and provides the low SN microphone recording signal obtained as a result to the microphone recording database 83 as learning data, and records the low SN microphone recording signal therein.

[0378] In addition, the data with added noise can be obtained by adding at least one of reverberation and noise to the microphone recording signal.

[0379] The low SN ratio obtained in this way means that the low-quality low SN microphone recording signal is the signal acquired in a low SN environment.

[0380] Therefore, the low-SN microphone recording signal is close to the actual microphone recording signal acquired during recording, and thus, by using such a low-SN microphone recording signal as learning data, object sound source signals can be predicted with higher accuracy. In other words, object sound source signals with higher quality can be obtained.

[0381] The overlay processing unit 361 does not perform overlay processing on the position information and movement information supplied from the acquisition unit 81 as learning data, but directly supplies the position information and movement information to the sensor database 82 as learning data, and records the position information and movement information therein.

[0382] Furthermore, for example, it can be considered that noise not present when recording learning data is included in the sensor signal when recording actual content, due to contact between the object and another object. Therefore, the data for adding such noise can be prepared in advance, and the superposition processing unit 361 can provide the data obtained by performing convolution of the data with added noise and position and movement information as learning data to the sensor database 82.

[0383] <Description of learning processing>

[0384] Next, we will refer to Figure 22 The flowchart shown describes the learning process performed using the learning device 351.

[0385] It should be noted that the processing in steps S341 and S342 is related to... Figure 5 The processes in steps S41 and S42 are the same, so their descriptions will be omitted.

[0386] In step S343, the superposition processing unit 361 performs superposition processing that convolves the data with added noise with the microphone recording signal provided by the acquisition unit 81, and provides the low SN microphone recording signal obtained as a result to the microphone recording database 83 as learning data, and records the low SN microphone recording signal therein.

[0387] Furthermore, the overlay processing unit 361 also directly supplies the microphone recording signal supplied from the acquisition unit 81 to the microphone recording database 83 as teaching data, recording the microphone recording signal therein. It also directly supplies the position information and motion information supplied from the acquisition unit 81 to the sensor database 82 as learning data, recording the position information and motion information therein. Additionally, as described above, convolution of the position information and motion information with the noise-added data can be performed.

[0388] When the overlay process is performed, the process in step S344 is then executed, and the learning process ends. The process in step S344 is related to... Figure 5 The process of step S43 shown in the figure is similar, so its description will be omitted.

[0389] Here, in step S344, position information and motion information are set as learning data among the low-SN microphone recording signal and multiple sensor signals included in the transmitted data. The position information and motion information are sensor signals other than the microphone recording signal to which reverberation or noise is added.

[0390] Then, machine learning is performed based on this learning data and the microphone recordings used as teaching data.

[0391] In this way, coefficient data is generated for an object sound source generator 11 that has location information, motion information, and low-SN microphone recording signals as its inputs and an object sound source signal as its output.

[0392] As described above, the learning device 351 generates a low-SN microphone recording signal that simulates recording in a low-SN environment from the microphone recording signal, and uses the low-SN microphone recording signal as learning data to perform machine learning. In this way, even when a low-SN environment is formed during recording, a high-quality object sound source signal can be obtained.

[0393] Furthermore, in this configuration, the microphone 61 is positioned within the recording device 111 during recording, and the sound source generating device 112 acquires transmission data from the recording device 111, including microphone recording signals, location information, and movement information. This microphone recording signal is recorded in a low-SN environment and therefore corresponds to the aforementioned low-SN microphone recording signal.

[0394] The sound source generating device 112 performs the operation shown in Figure 8 The processing in step S102 is that, based on the microphone recording signal, position information and movement information contained in the transmission data and coefficient data, the arithmetic operation processing of the object sound source generator 11 is performed to generate an object sound source signal.

[0395] In this case, unnecessary noise is suppressed, and a high-quality (i.e., high SN ratio) object sound source signal can be obtained.

[0396] Furthermore, any combination of embodiments can be used in the first to sixth embodiments described above.

[0397] <Computer Configuration Examples>

[0398] The above series of processes can also be performed by hardware or software. In the case where the series of processes are performed by software, a program for configuring the software is installed on the computer. Here, "computer" includes, for example, a computer built into dedicated hardware, or a general-purpose personal computer with various programs installed on it to perform various functions.

[0399] Figure 23 This is a block diagram illustrating a configuration example of computer hardware that uses a program to perform the above series of processes.

[0400] In a computer, the central processing unit (CPU) 501, read-only memory (ROM) 502, and random access memory (RAM) 503 are connected to each other via a bus 504.

[0401] The input / output interface 505 is further connected to the bus 504. The input unit 506, output unit 507, recording unit 508, communication unit 509, and driver 510 are connected to the input / output interface 505.

[0402] Input unit 506 includes a keyboard, mouse, microphone, imaging element, etc. Output unit 507 includes a display, speaker, etc. Recording unit 508 includes a hard disk, non-volatile memory, etc. Communication unit 509 includes a network interface, etc. Driver 510 drives a removable recording medium 511 such as a hard disk, optical disk, magneto-optical disk, or semiconductor memory.

[0403] In a computer configured as described above, for example, the CPU 501 loads the program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executes the program to perform the series of processes described above.

[0404] The program executed by the computer (CPU 501) can be recorded on, for example, a removable recording medium 511 as a packaging medium and provided in this state. The program can also be provided via wired or wireless transmission media (such as a local area network, the Internet, or digital satellite broadcasting).

[0405] In a computer, by installing a removable recording medium 511 in a drive 510, a program can be installed in a recording unit 508 via an input / output interface 505. Furthermore, the program can be received by a communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Additionally, the program can be pre-installed in a ROM 502 or in the recording unit 508.

[0406] It should be noted that a program executed by a computer may be a program that is processed sequentially in the order described in this specification, or it may be a program that is processed in parallel or at necessary timed intervals, such as when a call time is required.

[0407] The implementation of this technology is not limited to the above-described implementation, and various changes can be made within the scope of this technology without departing from its main purpose.

[0408] For example, this technology can be configured as cloud computing, where multiple devices share and collaborate on a function via a network.

[0409] In addition, each step described in the flowchart above can be performed by a single device or by multiple devices in a shared manner.

[0410] Furthermore, in cases where a step includes multiple processes, the multiple processes included in a step can be executed by a single device or by multiple devices in a shared manner.

[0411] In addition, this technology can also be configured as follows. (1)

[0413] A learning device includes a learning unit configured to perform learning based on one or more sensor signals acquired by one or more sensors mounted on an object and a target signal associated with the object and corresponding to a predetermined sensor, and to generate coefficient data configured with a generator having one or more of a plurality of sensor signals as its input and the target signal as its output. (2)

[0415] According to the learning device of (1), the target signal is a sound source signal corresponding to a microphone as a predetermined sensor. (3)

[0417] According to the learning device of (2), the target signal is an acoustic signal generated based on a microphone recording signal acquired by a microphone that is a predetermined sensor mounted on an object. (4)

[0419] The learning device according to any one of (1) to (3), wherein one or more sensors include at least one of a 9-axis sensor, a geomagnetic sensor, an accelerometer, a gyroscope sensor, a range sensor, a positioning sensor, an image sensor, and a microphone. (5)

[0421] The learning device according to any one of (1) to (3), wherein one or more sensors are sensors of a type different from microphones. (6)

[0423] The learning apparatus according to any one of (1) to (5) further comprises a correction processing unit that performs correction processing on microphone recording signals acquired by other microphones not mounted on the object as described in the positional relationship between the object and other microphones, the learning unit performs learning based on the corrected microphone recording signals, one or more sensor signals, and a target signal, and generates coefficient data for configuring a generator having the corrected microphone recording signals and one or more sensor signals as inputs and the target signal as its output. (7)

[0425] The learning apparatus according to any one of (1) to (5) wherein the learning unit performs learning for each combination of multiple sensors and generates coefficient data. (8)

[0427] The learning apparatus according to any one of (1) to (5) wherein the learning unit performs learning on each environmental condition around the object and generates coefficient data. (9)

[0429] The learning apparatus of any one of (1) to (4) further includes a superposition processing unit configured to add reverberation or noise to a microphone recording signal as a sensor signal acquired by the microphone as a sensor, wherein the learning unit performs learning based on the microphone recording signal to which reverberation or noise has been added, one or more sensor signals other than the microphone recording signal, and a target signal, and generates coefficient data of a generator configured to have the microphone recording signal and the sensor signal other than the microphone recording signal as its input and the target signal as its output. (10)

[0431] A learning method using a learning device, the learning method comprising: performing learning based on one or more sensor signals acquired by one or more sensors mounted on an object and a target signal related to the object and corresponding to a predetermined sensor; and generating coefficient data configured with a generator having one or more sensor signals as its input and the target signal as its output. (11)

[0433] A program that causes a computer to perform the following processing: performing learning based on one or more sensor signals acquired by one or more sensors mounted on an object and a target signal associated with the object and corresponding to a predetermined sensor, and generating coefficient data configured with a generator having one or more sensor signals as its input and the target signal as its output. (12)

[0435] A signal processing apparatus includes: an acquisition unit configured to acquire one or more sensor signals acquired by one or more sensors mounted on an object; and a generation unit configured to generate a target signal related to the object and corresponding to a predetermined sensor based on coefficient data of a generator pre-learned and generated by the one or more sensor signals. (13)

[0437] According to the signal processing apparatus of (12), the target signal is a sound source signal corresponding to a microphone which is a predetermined sensor mounted on an object. (14)

[0439] According to the signal processing apparatus of (12) or (13), one or more sensors include at least one of a 9-axis sensor, a geomagnetic sensor, an accelerometer, a gyroscope sensor, a ranging sensor, a positioning sensor, an image sensor, and a microphone. (15)

[0441] According to the signal processing apparatus of (12) or (13), one or more sensors are sensors of a type different from microphones. (16)

[0443] The signal processing apparatus according to any one of (12) to (15) wherein the acquisition unit further acquires a microphone recording signal acquired by another microphone not mounted on the object, and the signal processing apparatus further includes a correction processing unit configured to perform correction processing on the microphone recording signal acquired by the other microphone as described in the positional relationship between the object and the other microphone, wherein the generation unit generates a target signal based on coefficient data, the corrected microphone recording signal, and one or more sensor signals. (17)

[0445] The signal processing apparatus according to any one of (12) to (15) includes an acquisition unit that acquires multiple sensor signals, and the signal processing apparatus further includes a fault detection unit configured to detect sensor faults based on the multiple sensor signals, wherein the generation unit generates a target signal based on the sensor signal of a fault-free sensor among the multiple sensor signals and coefficient data of a generator configured to have the fault-free sensor signal as its input and a target signal as its output. (18)

[0447] The signal processing apparatus according to any one of (12) to (15) further includes an environmental condition acquisition unit, the environmental condition acquisition unit being configured to acquire environmental condition information representing the environmental conditions surrounding the object, wherein the generation unit generates a target signal based on coefficient data corresponding to the environmental condition information and the one or more sensor signals. (19)

[0449] A signal processing method using a signal processing apparatus includes: acquiring one or more sensor signals acquired by one or more sensors mounted on an object; and generating a target signal related to the object and corresponding to a predetermined sensor by learning coefficient data of a pre-generated generator and the one or more sensor signals based on a configuration. (20)

[0451] A program that causes a computer to perform the following processes: acquiring one or more sensor signals acquired by one or more sensors mounted on an object; and generating a target signal that is object-related and corresponds to a predetermined sensor by learning coefficient data of a pre-generated generator and the one or more sensor signals based on a configuration.

[0452] [List of Reference Numbers]

[0453] 11. Object Sound Source Generator

[0454] 51 Recording device

[0455] 52 Learning Devices

[0456] 81 Acquisition Unit

[0457] 84 Learning Units

[0458] 111 Recording device

[0459] 112 Sound source generating device

[0460] 131 Acquisition Unit

[0461] 133. Object sound source generating unit.

Claims

1. A learning device comprising a learning unit and a correction processing unit, the correction processing unit being configured to perform correction processing on microphone recording signals acquired by other microphones not mounted on an object, based on the positional relationship between the object and the other microphones. in, The learning unit performs learning based on the microphone recording signal after the correction process, one or more sensor signals acquired by one or more sensors mounted on the object, and a target signal associated with the object and corresponding to a predetermined sensor, and generates coefficient data for configuring a generator, which uses the microphone recording signal after the correction process and the one or more sensor signals as inputs to the generator and the target signal as outputs to the generator.

2. The learning device according to claim 1, wherein, The target signal is a sound source signal corresponding to the microphone, which is the predetermined sensor.

3. The learning device according to claim 2, wherein, The target signal is an acoustic signal generated based on a microphone recording signal acquired by the microphone, which is the predetermined sensor mounted on the object.

4. The learning device according to claim 1, wherein, The one or more sensors include at least one of a 9-axis sensor, a geomagnetic sensor, an accelerometer, a gyroscope, a ranging sensor, a positioning sensor, an image sensor, and a microphone.

5. The learning device according to claim 1, wherein, The one or more sensors are sensors of a different type than microphones.

6. The learning device according to claim 1, wherein, The learning unit performs learning for each combination of multiple sensors and generates the coefficient data.

7. The learning device according to claim 1, wherein, The learning unit performs learning for each environmental condition of the object's surrounding environment and generates the coefficient data.

8. The learning device according to claim 1, further comprising: The overlay processing unit is configured to add reverberation or noise to a microphone recording signal, which is a sensor signal acquired by a microphone that serves as the sensor. The learning unit performs learning based on the microphone recording signal with added reverberation or noise, sensor signals other than the microphone recording signal from the one or more sensor signals, and the target signal, and generates the coefficient data configuring the generator. The generator uses the microphone recording signal and the sensor signals other than the microphone recording signal as inputs to the generator, and uses the target signal as the output of the generator.

9. A learning method using a learning device, the learning method comprising: For microphone recording signals acquired by other microphones not mounted on the object, a correction process is performed based on the positional relationship between the object and the other microphones. Learning is performed based on the corrected microphone recording signals, one or more sensor signals acquired by one or more sensors mounted on the object, and a target signal associated with the object and corresponding to a predetermined sensor. Coefficient data for configuring a generator is generated, wherein the corrected microphone recording signals and the one or more sensor signals are used as inputs to the generator and the target signal is used as the output of the generator.

10. A computer-readable storage medium storing a program that, when executed, causes a computer to perform the following processes: performing correction processing on microphone recording signals acquired by other microphones not mounted on an object, based on the positional relationship between the object and the other microphones; performing learning based on the corrected microphone recording signals, based on one or more sensor signals acquired by one or more sensors mounted on the object, and a target signal associated with the object and corresponding to a predetermined sensor; and generating coefficient data for configuring a generator, the generator taking the corrected microphone recording signals and the one or more sensor signals as inputs to the generator and the target signal as outputs to the generator.

11. A signal processing apparatus, comprising: The acquisition unit is configured to acquire one or more sensor signals acquired by one or more sensors mounted on the object; as well as The generation unit is configured to generate a target signal related to the object and corresponding to a predetermined sensor by learning coefficient data of a pre-generated generator and the one or more sensor signals based on the configuration. The acquisition unit also acquires a microphone recording signal from another microphone not mounted on the object. The signal processing device further includes a correction processing unit configured to perform correction processing on the microphone recording signal acquired by the other microphone, based on the positional relationship between the object and the other microphone. The generation unit generates the target signal based on the coefficient data, the microphone recording signal after correction processing, and the one or more sensor signals.

12. The signal processing apparatus according to claim 11, wherein, The target signal is a sound source signal corresponding to a microphone that is a predetermined sensor mounted on the object.

13. The signal processing apparatus according to claim 11, wherein, The one or more sensors include at least one of a 9-axis sensor, a geomagnetic sensor, an accelerometer, a gyroscope, a ranging sensor, a positioning sensor, an image sensor, and a microphone.

14. The signal processing apparatus according to claim 11, wherein, The one or more sensors are sensors of a different type than microphones.

15. The signal processing apparatus according to claim 11, in, The acquisition unit acquires signals from multiple sensors. The signal processing device further includes a fault detection unit configured to detect sensor faults based on the signals from the plurality of sensors. The generation unit generates the target signal based on the sensor signal of the fault-free sensor among the plurality of sensor signals and the coefficient data configuring the generator. The generator uses the sensor signal of the fault-free sensor as the input of the generator and the target signal as the output of the generator.

16. The signal processing apparatus according to claim 11, further comprising: The environmental condition acquisition unit is configured to acquire environmental condition information representing the environmental conditions of the object's surrounding environment. The generation unit generates the target signal based on the coefficient data corresponding to the environmental condition information and the one or more sensor signals.

17. A signal processing method using a signal processing apparatus, the signal processing method comprising: Acquire one or more sensor signals obtained by one or more sensors mounted on the object; as well as Based on the configuration, by learning the coefficient data of the pre-generated generator and the signals of the one or more sensors, a target signal related to the object and corresponding to a predetermined sensor is generated; It also acquires microphone recording signals obtained by another microphone not mounted on the object. The microphone recording signal acquired by the other microphone is corrected according to the positional relationship between the object and the other microphone. The target signal is generated based on the coefficient data, the microphone recording signal after the correction process, and the one or more sensor signals.

18. A computer-readable storage medium storing a program that, when executed, causes a computer to perform the following processes: Acquire one or more sensor signals from one or more sensors mounted on the object; and Based on the configuration, by learning the coefficient data of the pre-generated generator and the signals of the one or more sensors, a target signal related to the object and corresponding to a predetermined sensor is generated; It also acquires microphone recording signals obtained by another microphone not mounted on the object. The microphone recording signal acquired by the other microphone is corrected according to the positional relationship between the object and the other microphone. The target signal is generated based on the coefficient data, the microphone recording signal after the correction process, and the one or more sensor signals.

Citation Information

Patent Citations

  • Information processing device, information processing method and program

    JP2017205213A

  • Recognition system and recognition method

    JP2020162765A

  • DC Converter

    JP7388432B2

  • Animal-machine audio interaction system

    US20110082574A1