Signal processing device, method and program

By using sensors to detect the time interval of sound and combining sensor signals when there are other moving objects around the moving object, the problem of poor target sound quality in a multi-sound source environment is solved, and high-precision, high-quality sound extraction is achieved.

CN114402390BActive Publication Date: 2025-09-16SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080064274.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-18
Filing Date
2020-09-04
Publication Date
2025-09-16
Estimated Expiration
2040-09-04

AI Technical Summary

Technical Problem

When recording content such as sports and drama from a free viewpoint, the complex movements of multiple sound sources make it difficult to obtain high-quality target sounds with a high signal-to-noise ratio.

Method used

By detecting the time interval of sound using a sensor attached to a moving object when there are other moving objects around the moving object, and combining the sensor signal with the recorded signal, high-quality target sound can be distinguished and extracted.

Benefits of technology

It achieves high-precision extraction of high-quality target sounds in complex sound source environments, improving the signal-to-noise ratio and sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114402390B_ABST
    Figure CN114402390B_ABST
Patent Text Reader

Abstract

This technology relates to a signal processing device, method, and program that enable the acquisition of high-quality target sounds. The signal processing device includes a time interval detection unit for detecting a time interval of sounds emitted by a moving object, contained in a recorded signal obtained by collecting sounds around the moving object and a sensor signal output from a sensor mounted on the moving object, when other moving objects are present around the moving object. This technology can be applied to recording systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to a signal processing device, method, and program, and particularly, to a signal processing device, method, and program that make it possible to obtain high-quality target sound. Background Art

[0002] To reproduce a sound field generated from a free viewpoint such as a bird's-eye view and a walkthrough view, it is important to record the sound from a target sound source with a high SN ratio (Signal-to-Noise Ratio), and at the same time, it is necessary to acquire information indicating the position and orientation of the corresponding sound source.

[0003] Specific examples of sounds from the target sound source include voices from humans, general human action sounds (such as walking sounds and running sounds), and action sounds specific to content such as sports and games (such as the sound of a ball being kicked).

[0004] Furthermore, as a technology associated with user behavior recognition, for example, a technology of obtaining behavior recognition results of one or more users by analyzing distance measurement sensor data detected by a plurality of distance measurement sensors has been proposed (for example, see PTL 1).

[0005] [Citation List]

[0006] [Patent Document]

[0007] [PTL1]

[0008] JP2017-205213A Summary of the Invention

[0009] [Technical Issues]

[0010] Meanwhile, when capturing content such as sports and drama from a free viewpoint, the recording space includes multiple sound sources. These sound sources may sometimes make complex movements. In such cases, it is difficult to capture the target sound source with a high signal-to-noise ratio. Consequently, it is difficult to obtain high-quality target sound.

[0011] The present technology was developed in consideration of the above circumstances and aims to obtain high-quality target sounds.

[0012] [Solution to the problem]

[0013] A signal processing device according to one aspect of the present technology includes an interval detection unit configured to detect a time interval containing a sound emitted from a mobile object, where the sound is included in a recorded signal obtained by collecting sounds around the mobile object in a state where other mobile objects are present around the mobile object, and to detect the time interval based on the recorded signal and a sensor signal output from a sensor attached to the mobile object.

[0014] A signal processing method or program according to one aspect of the present technology includes the steps of detecting a time interval containing a sound emitted from a mobile body, wherein the sound is included in a recorded signal obtained by collecting sounds around a mobile body in a state where other mobile bodies are present around the mobile body, and detecting the time interval based on the recorded signal and a sensor signal output from a sensor attached to the mobile body.

[0015] According to aspects of the present technology, a time interval containing sound emitted from a moving object is detected based on a recorded signal and a sensor signal output from a sensor attached to the moving object, and the sound is included in the recorded signal obtained by collecting sounds around the moving object in a state where other moving objects exist around the moving object. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a diagram depicting a configuration example of an ingestion system.

[0017] Figure 2 This is a diagram explaining objects and object sound sources.

[0018] Figure 3 is a diagram depicting an example of sound source classification section information.

[0019] Figure 4 This is a diagram explaining the generation of sound source classification section information.

[0020] Figure 5 is a diagram explaining selection of removal target object.

[0021] Figure 6 This is a flowchart explaining the collection process.

[0022] Figure 7 This is a flowchart explaining the data generation process.

[0023] Figure 8 is a diagram depicting a configuration example of an ingestion system.

[0024] Figure 9 This is a flowchart explaining the data generation process.

[0025] Figure 10 is a diagram depicting a configuration example of an ingestion system.

[0026] Figure 11 This is a flowchart explaining the collection process.

[0027] Figure 12 This is a flowchart explaining the data generation process.

[0028] Figure 13 is a diagram depicting a configuration example of an ingestion system.

[0029] Figure 14 This is a flowchart explaining the data generation process.

[0030] Figure 15 is a diagram depicting a configuration example of an ingestion system.

[0031] Figure 16 is a diagram depicting a configuration example of a computer. DETAILED DESCRIPTION

[0032] Hereinafter, embodiments to which the present technology is applied will be described in detail with reference to the accompanying drawings.

[0033] <First embodiment>

[0034] <Configuration example of the recording system>

[0035] This technology is used to obtain high-quality target sound by attaching a microphone, a distance measuring device, a camera, etc. to each of multiple mobile objects in a target space, and by extracting the sound of the own mobile object, while distinguishing the sound of the own mobile object from the sounds of other mobile objects based on the sound recording signal, position information associated with the mobile object, motion information associated with the mobile object, surrounding images, etc.

[0036] Specifically, examples of content to which the present technology is appropriately applied include the following items.

[0037] - Content that recreates a field where team sports are played

[0038] - Content that reproduces musical performances (e.g., performances by an orchestra, military band, etc.)

[0039] - Content that recreates a space where multiple performers are present, such as in musicals, operas, and plays

[0040] - Reproduce content in any space such as sports days, live performances, various types of events, theme park parades, etc.

[0041] Hereinafter, the space to be included is referred to as the target space.

[0042] It is particularly assumed here that a plurality of mobile bodies exist in the same target space, and a collection device for collecting content is attached to or built into each of these mobile bodies.

[0043] In this case, a mobile body to which a recording device is separately attached or a mobile body in which a recording device is separately built-in is assumed to be an object, and sound emitted from each object is recorded (collected) as sound of a corresponding object sound source.

[0044] For example, each object (moving body) in the target space may be a person such as a sports player, or may be a robot, a vehicle, or a flying object such as a drone, to which a recording device is attached or in which a recording device is built.

[0045] For example, in the case where the subject is a person, it is preferable that the recording device attached to the person is as small as possible to avoid affecting the person's performance and not be visually recognized by the surrounding environment.

[0046] In addition, for example, the collection device includes a microphone for collecting sound from the object sound source, a sensor such as a nine-axis sensor for measuring the movement or direction (orientation) of the object, a distance measuring device for measuring position, a camera or other device for capturing surrounding images.

[0047] For example, the distance measurement device herein is a GPS (Global Positioning System) device, an indoor distance measurement beacon receiver, or other device for measuring the position of an object, and position information indicating the position of the object can be acquired using the distance measurement device.

[0048] Furthermore, based on the output from the sensor provided on the recording device, motion information indicating the motion of the object (such as speed and acceleration) or indicating the direction (orientation) of the object can be acquired.

[0049] The recording device uses a built-in microphone, sensor, and distance measurement device to obtain a recorded signal, including position information and motion information associated with the object. The recorded signal is an audio signal obtained by collecting the sounds surrounding the object. Furthermore, if the recording device is equipped with a camera, it can also obtain an image signal of the image surrounding the object.

[0050] An object sound source signal, which is an audio signal of a sound as a target sound generated from the object sound source, is obtained using the recorded signal, position information, motion information, and image signal of each object thus obtained.

[0051] Examples of the sound from the object sound source as the target sound herein include a voice uttered by a person as the object, a walking sound or a running sound of the object, and an action sound such as applause.

[0052] The recorded signal obtained for each object includes not only the sound emitted by the object itself, but also sounds emitted by other nearby objects. Furthermore, the recorded signal includes sounds belonging to the object but emitted from multiple different object sound sources, i.e., sounds of different classifications, such as the object's own voice and movement sounds.

[0053] In this technology, by using position information, motion information, and image signals obtained for each object as needed, it is possible to distinguish (identify) sounds contained in the recorded signal and generated from various target sound sources, and extract the target sound source signal for each target sound source from the recorded signal.

[0054] Specifically, for example, by specifying the motion state of the object based on the motion information, a time interval including and containing the sound of the corresponding object sound source can be detected in the recorded signal.

[0055] Therefore, for example, by extracting the signal within the sound interval of the target sound source from the recorded signal and performing signal processing such as sound quality correction, sound source separation and noise removal on the extracted signal as needed, a high-quality target sound source signal exhibiting a high SN ratio can be obtained.

[0056] Furthermore, by integrating information such as position information, motion information, and image signals obtained for each of a plurality of objects, a higher quality object sound source signal can be obtained, thereby improving the accuracy of the detection result of the time interval of the sound from the object sound source.

[0057] Hereinafter, the present technology will be described in more detail.

[0058] Figure 1 is a diagram depicting a configuration example of a recording system according to an embodiment to which the present technology is applied.

[0059] exist Figure 1 In the example depicted in , the recording system includes a recording device 11 attached to an object as a mobile body, and a server 12 that receives transmission data from the recording device 11 and generates an object sound source signal.

[0060] Note that the recording apparatus 11 may be built in a mobile body. However, in the following description, it is assumed that the recording apparatus 11 is attached to the mobile body.

[0061] The collection device 11 is attached to a mobile body that moves freely in a target space as a collection target. The collection device 11 generates transmission data including a collection signal, position information, and motion information, and transmits the generated transmission data to the server 12.

[0062] Note that although only one receiving apparatus 11 is depicted herein, in actual cases there are a plurality of receiving apparatuses 11 , and each receiving apparatus is attached to a corresponding one of a plurality of objects that are different from each other.

[0063] The server 12 outputs object sound source data including an object sound source signal and metadata for each object sound source as content data based on the transmission data received from the plurality of recording devices 11. Note that the server 12 is not required to be arranged in the target space.

[0064] Furthermore, the recording device 11 includes a microphone 21 , a motion measurement unit 22 , a position measurement unit 23 , a recording unit 24 , and a transmission unit 25 .

[0065] The microphone 21 collects sounds around the recording device 11 and supplies a recording signal obtained as a result of the sound collection to the recording unit 24. Note that the recording signal may be a monaural signal. However, in this description, it is assumed that the recording signal is a multi-channel signal.

[0066] The recording device 11 collects sounds using the microphone 21 in a state where not only the object to which the recording device 11 is attached but also other objects exist around the recording device 11. Therefore, the sounds associated with the recording signal include sounds from a plurality of sound sources.

[0067] The motion measurement unit 22 includes a sensor for measuring the motion or direction of the object, such as a nine-axis sensor, a magnetic field sensor, an acceleration sensor, or a gyro sensor, and outputs a sensor signal indicating a measurement result (sensing value) as motion information to the recording unit 24 .

[0068] Specifically, the motion measurement unit 22 measures the motion or direction of the object during the period in which the microphone 21 collects sound, and outputs motion information indicating the measurement result.

[0069] Note that, described here is an example in which the sensor signal is used as motion information without change. However, as needed, motion information can be generated from the sensor signal by performing signal processing on the sensor signal using the recording unit 24.

[0070] Furthermore, the motion measurement unit 22 may be provided outside the recording apparatus 11 and attached to a position different from the position where the recording apparatus 11 is attached to the object.

[0071] For example, the position measurement unit 23 includes a distance measurement device such as a GPS device and an indoor distance measurement beacon receiver. The position measurement unit 23 measures the position of the object to which the recording device 11 is attached, and outputs position information indicating the measurement result to the recording unit 24.

[0072] Note that the recorded signal, motion information, and position information are acquired simultaneously within the same time period.

[0073] As needed, the recording unit 24 performs AD (Analog-to-Digital) conversion on the recording signal provided from the microphone 21 , the motion information provided from the motion measurement unit 22 , and the position information provided from the position measurement unit 23 , and provides the processed signals and information to the transmission unit 25 .

[0074] The transmission unit 25 generates transmission data including the recorded signal, motion information, and position information provided from the recording unit 24 by performing compression processing on the recorded signal, motion information, and position information. The transmission unit 25 then transmits the obtained transmission data to the server 12 via a wireless network or the like.

[0075] Furthermore, the server 12 includes a receiving unit 31 , a section detecting unit 32 , and a target sound source data generating unit 33 .

[0076] The receiving unit 31 receives transmission data transmitted from each of the plurality of collection devices 11 and extracts a collection signal, position information, and motion information from the transmission data.

[0077] The receiving unit 31 supplies the recorded signal to the section detecting unit 32 and the target sound source data generating unit 33. The receiving unit 31 also supplies the motion information to the section detecting unit 32 and the motion information and position information to the target sound source data generating unit 33.

[0078] Based on the recorded signal and motion information supplied from the receiving unit 31 , the section detecting unit 32 detects, for each recorded signal, the classification (type) of the sound generated from the target sound source and contained in the recorded signal, i.e., the classification of the target sound source, and the time section containing the sound of the target sound source.

[0079] The section detection unit 32 supplies the target sound source data generation unit 33 with sound source classification section information indicating the classification and time section of the sound of the target sound source detected from the recorded signal.

[0080] Furthermore, the section detection unit 32 provides the target sound source data generation unit 33 with sound source classification information indicating the object corresponding to the recorded signal and the classification of the sound of the target sound source detected from the recorded signal. In other words, the sound source classification information is information indicating the classification of the target sound source, which is the source of the sound based on the target sound source signal, and the object corresponding to the source of the sound.

[0081] The object sound source data generation unit 33 generates object sound source data based on the recorded signal, motion information, and position information supplied from the receiving unit 31, and based on the sound source classification section information and sound source classification information supplied from the section detection unit 32. The object sound source data generation unit 33 then outputs the generated object sound source data to a reproduction device or the like arranged in the next stage.

[0082] The object sound source data generating unit 33 includes a signal processing unit 41 and a metadata generating unit 42 .

[0083] The signal processing unit 41 performs predetermined signal processing on the recorded signal supplied from the receiving unit 31 based on the sound source classification section information supplied from the section detection unit 32 and the motion information and position information supplied from the receiving unit 31. This generates a target sound source signal.

[0084] For example, based on the sound source classification interval information, the object sound source signal is generated by performing one or more types of signal processing (for example, processing of extracting the time interval of the sound of the object sound source from the recorded signal and processing of silencing the time interval of the sound of the object sound source that does not contain the recorded signal).

[0085] Furthermore, the metadata generating unit 42 generates metadata for each target sound source (ie, each target sound source signal) including the sound source classification information supplied from the section detecting unit 32 and the motion information and position information supplied from the receiving unit 31 .

[0086] The object sound source data including the object sound source signal and metadata obtained thereby is output from the object sound source data generating unit 33 to the next stage.

[0087] <Server Units>

[0088] Next, each unit included in the server 12 will be described in more detail.

[0089] First, the section detecting unit 32 will be described.

[0090] Note that, hereinafter, a predetermined object of interest will also be referred to as a target object, and objects other than the target object will also be referred to as other objects, where appropriate.

[0091] The section detection unit 32 distinguishes sounds emitted by the target object contained in the recorded signal from sounds emitted by other objects contained in the recorded signal, specifies the category of the sounds emitted by the target object, and detects the time section of the sounds emitted by the target object.

[0092] As described above, the section detection unit 32 receives the recorded signal and motion information as input, and outputs the sound source classification section information and the sound source classification information corresponding to the input.

[0093] It is assumed here that the mobile body to which the recording apparatus 11 is attached is an object, and each part of the object serves as an object sound source, for example, Figure 2 As depicted in . It is assumed that the sound of the object sound source is emitted from the object defined in this way. Note that, more specifically, it is also assumed that a musical instrument or the like carried by the object can be used as the object sound source.

[0094] Furthermore, in the recording device 11 and the server 12 , categories of target sound sources are defined in advance.

[0095] For example, it is assumed that some classifications of object sound sources, that is, some classifications of sounds of object sound sources are common to all contents, and other classifications are different for each content.

[0096] Specifically, if Figure 2 Examples of the classification of the sound of the target sound source defined as a classification common to all contents, depicted in the right part of , include speech uttered by a person as a target, and the walking sound, running sound, and applause of the person.

[0097] Furthermore, examples of the categories of sounds of the object sound sources defined as content-specific categories associated with sports include the sounds of passing balls, shots, and whistles. Examples of the categories of sounds of the object sound sources defined as content-specific categories associated with music include the sounds of musical instruments. Furthermore, examples of the categories of sounds of the object sound sources defined as content-specific categories associated with drama, dance, etc. include sounds associated with the actions of actors, such as the rustling of clothes and the sound of footsteps.

[0098] The section detection unit 32 generates sound source classification section information indicating to which category the sound of the target sound source belongs and in which time section of the recorded signal the sound is included.

[0099] The sound source classification interval information can be in any form, for example, Figure 3 As depicted in , for example, binary information indicated by 0 or 1 and probability information represented by continuous values. In addition, the sound source classification interval information may be interval information for a time signal or interval information for each frequency bin.

[0100] For example, in Figure 3 In the example depicted in the upper left portion of , the sound source classification section information is binary information defined for each target sound source, and indicates whether the sound of the target sound source is included at each time point in the recorded signal given as a time signal.

[0101] In this example, the corresponding lines indicate whether "background noise," "walking or running sound," "shooting sound," and "voice" are included as target sound sources at each time point.

[0102] Specifically, the horizontal direction of each line indicates time, and an interval in which a line protrudes upward indicates that a sound of the target sound source is included in the interval.

[0103] In addition, Figure 3In the example depicted in the upper right part of , the sound source classification interval information is continuous value information defined for each object sound source, and the continuous value information indicates a probability value representing the probability of containing the sound of the object sound source at each time point in the recorded signal given as a time signal.

[0104] In this example, corresponding curves indicate probability values ​​representing the probability of including "background noise," "walking or running sound," "shooting sound," or "voice" as the object sound source at each time point.

[0105] For example, when detection of an object sound source is set as a recognition problem from multiple classes, each of the continuous probability values ​​indicating the probability of a sound containing the corresponding object sound source is an output value of a DNN (Deep Neural Network) obtained by machine learning or the like.

[0106] In addition, Figure 3 In the example depicted in the lower left portion of , the sound source classification section information is binary information in the form of a time-frequency mask generated for each object sound source classification.

[0107] This binary information in the form of a time-frequency mask uses binary values ​​to indicate, for each component of the time-frequency interval of the recorded signal, whether the sound of the target sound source is included in each time interval (time point) of the recorded signal. Specifically, in this example, the vertical axis indicates the time-frequency interval, while the horizontal axis indicates time.

[0108] In addition, Figure 3 In the example depicted in the lower right portion of , the sound source classification interval information is continuous value information in the form of a time-frequency mask generated for each object sound source classification. Similar to the above, in this example, the vertical axis indicates the time-frequency interval, and the horizontal axis indicates time.

[0109] This continuous value information in the form of a time-frequency mask expresses, using continuous values, the probability that the sound of the target sound source is included in each time interval (time point) of the recorded signal for each component of the time-frequency interval of the recorded signal.

[0110] Note that the sound source classification interval information is not limited to Figure 3 The example depicted in , and may be any form of information. It suffices to appropriately determine which form of sound source classification section information is used according to the signal processing performed by the signal processing unit 41 arranged in the next stage.

[0111] Furthermore, to generate the sound source classification section information, the section detection unit 32 detects the sound of the target sound source for each classification in each time section of the recorded signal. In other words, the section detection unit 32 detects the sound of the target sound source for each classification in each time section of the recorded signal.

[0112] The motion information obtained by the collection device 11 is information indicating the motion or direction of the subject during sound collection performed by the microphone 21 to obtain a collection signal.

[0113] Therefore, by detecting each time interval containing the sound of the object sound source based on the motion information, it is possible to determine whether the sound contained in each time interval of the recorded signal is emitted from the object or from surrounding objects.

[0114] Examples of the sound of the object sound source include various types of motion sounds such as walking sounds, running sounds, applause, shooting sounds when playing soccer, and footsteps when dancing.

[0115] For example, as a method for detecting the time interval of the action sound, a method of detecting the time interval of the action sound using a simple algorithm such as threshold processing using a threshold value may be adopted.

[0116] In this case, for example, a time interval in which the sensing value of the sensor as motion information falls within a specific range defined for the operation sound to be detected is set as the time interval of the operation sound.

[0117] Furthermore, for example, an identifier such as a DNN can be created through multimodal learning and can be used to detect the time interval of action sounds.

[0118] In this case, an identifier such as a DNN is created through learning. This identifier receives the recorded signal and sensor values ​​of sensors such as an acceleration sensor, a magnetic field sensor, and a gyroscope sensor as input, obtains the sensor values ​​as motion information, and outputs the presence or absence of an action sound in each time interval of the recorded signal.

[0119] Note that as the above-mentioned identifier, for example, such an identifier that sets a plurality of action sounds common to all contents as detection targets or such an identifier that sets an action sound unique to the contents as detection targets may be used.

[0120] Here, a specific example of a time interval for detecting an action sound will be described.

[0121] For example, in the case of detecting the time interval of the walking sound or the running sound of the object as the action sound, it is sufficient to use the sensor value indicating the acceleration of the object in the up-down direction measured by the acceleration sensor as the motion information.

[0122] In this case, the subject's walking or running can be detected based on changes in sensor values. For example, time intervals where the frequency (i.e., oscillation frequency) of the sensor value's temporal waveform is approximately 2 Hz or lower can be identified as intervals where the subject is walking, i.e., time intervals where the subject is making walking sounds. Similarly, time intervals where the sensor value's oscillation frequency is approximately between 3 Hz and 4 Hz can be identified as intervals where the subject is running, i.e., time intervals where the subject is making running sounds.

[0123] Furthermore, when detecting the time interval of a kicking sound as an action sound or the time interval of a sound associated with a shot during a ball game, it is sufficient to use information related to the rotation angle, etc., which is primarily measured by a gyro sensor and indicates the rotation of the subject, as motion information. This information is useful as motion information because the subject rotates his or her body during the kicking action or the shooting action.

[0124] Furthermore, in the case of detecting the time interval of finger clicks as action sounds or the time interval of sounds produced when an object hits his or her body, for example, it is sufficient to use changes in sensor values ​​of an acceleration sensor, gyro sensor, magnetic field sensor, or other sensors.

[0125] In this case, for example, an acceleration sensor, a gyro sensor, or a magnetic field sensor used as the motion measurement unit 22 is attached to a body part, a wrist, an arm, etc. of a person serving as an object to detect the movement of the object's body or hand based on the amount of change in the sensor value corresponding to the attached part.

[0126] Furthermore, when detecting the time interval of breathing sounds of a subject person as the time interval of the motion sounds, it is sufficient to use a sensor value indicating a small displacement of the subject in the up-down direction measured by an acceleration sensor as motion information.

[0127] In this case, the subject's breathing action can be detected based on changes in the sensor value. For example, a time interval in which the sensor value oscillates at a frequency approximately in the range of 0.5 Hz to 1 Hz is identified as a time interval in which breathing action is being performed, allowing the recording of breathing sounds at an audible level, i.e., a time interval in which the subject's breathing sounds are being heard.

[0128] In addition, the time interval of the sound of each object sound source can be detected using an identifier such as a DNN, which receives the recorded signal and motion information as input and outputs the presence or absence of the sound of the object sound source based on the characteristics of each action when the sound is emitted from the corresponding object as described above.

[0129] For example, Figure 4As depicted in FIG, by using the recorded signal and the sensor value (sensor signal) of the acceleration sensor as an identifier of motion information, the time period of the walking sound as the sound from the target sound source can be accurately detected.

[0130] exist Figure 4 In the figure, the portion indicated by arrow Q11 shows the time waveform of the recorded signal, while the portion indicated by arrow Q12 shows the frequency spectrum of the recorded signal. Furthermore, the portion indicated by arrow Q13 shows the time waveform of the sensor signal from the acceleration sensor, while the portion indicated by arrow Q14 shows the frequency spectrum of the sensor signal. Note that the horizontal direction of each portion indicated by arrows Q11 to Q14 in the figure represents time.

[0131] In this example, for example, the walking sound of the target object and the walking sounds of other objects located around the target object are mixed in the portion indicated by arrow A11 and the like in the recorded signal.

[0132] In this case, it is difficult to determine whether the walking sound component included in the recorded signal belongs to the target object or another object based only on the time waveform and spectrum of the recorded signal.

[0133] Therefore, in this example, whether such a component belongs to the target object or to another object is determined (discrimination is performed) based not only on the recorded signal but also on the sensor signal (motion information).

[0134] The time waveform of the sensor signal indicated by arrow Q13 fluctuates periodically in the vertical direction. The value of this time waveform, that is, the value of the component in the vertical direction, represents the vertical component of the floor reaction force of the target object.

[0135] In particular, in this case, a protruding portion protruding upward in the figure, such as the portion indicated by arrow A12, corresponds to a body motion of the target object in one step. Obviously, the sensor signal contains information indicating the body motion of the target object at a high SN ratio.

[0136] Furthermore, it is apparent that the bright and dark pattern of the frequency spectrum of the sensor signal indicated by the arrow Q14 also shows a clear correspondence with the temporal waveform of the sensor signal indicated by the arrow Q13 .

[0137] As described above, the sensor signal contains information indicating the body motion of the target subject at a high SN ratio, but contains no information indicating the body motion of other subjects at all.

[0138] Therefore, by using the recorded signal and the motion information, the time period of the sound of the target sound source of the target object can be accurately detected.

[0139] Specifically, for example, if the recorded signal contains sounds of other objects at the same sound pressure as the target object's sound, it is difficult to accurately detect the time period of the target object's sound using only the recorded signal. However, by using not only the recorded signal but also motion information, the time period of the target object's sound can be accurately detected.

[0140] Generally, in the fields of behavior recognition and the like, a method of estimating the behavior of an object based on sensor values ​​of an acceleration sensor, a gyroscope sensor, a magnetic field sensor, or other sensors is often proposed.

[0141] On the other hand, as described above, for each classification of the sound of the object sound source, the section detection unit 32 distinguishes the sound emitted from the target object from the sound emitted from other objects by using the recorded signal and the motion information.

[0142] Note that, for example, walking and running are both defined as continuous actions in fields such as behavior recognition and physical therapy, and are therefore often described as continuous state transitions, such as the stance phase and the swing phase.

[0143] On the other hand, for example, the interval detection unit 32 detects a time interval in which the walking sound or the running sound is actually generated, that is, the time interval from the ground contact of the foot (more specifically, the heel or the toe) of the person as the subject to the ground leaving is detected as the time interval of the walking sound or the running sound.

[0144] Furthermore, the time period of the voice, which is the target sound source, uttered by the object can also be accurately detected based on the motion information.

[0145] For example, in a case where the motion measurement unit 22 is attached to a portion around the neck or head of a person as an object, when the target object makes a sound, information indicating the body movement caused by the sound is observed with a high SN ratio in the sensor signal corresponding to the motion information.

[0146] Therefore, similar to the case of action sounds, by using the recorded signal and motion information, it is possible to achieve highly accurate distinction between the speech emitted from the target object and the speech emitted from other objects within the time interval of the emitted speech.

[0147] Note that, in certain cases, it may be difficult to obtain motion information including information indicating the body motion of the target object when uttering a sound with a high SN ratio.

[0148] However, in this case, it is sufficient to utilize such a property that, for example, when the voice emitted from the target object attached to the collection device 11 is collected by the plurality of microphones included in the microphone 21, the propagation direction of the emitted voice toward each microphone becomes substantially constant.

[0149] Specifically, for example, the section detection unit 32 performs DS (Delay and Sum Beamforming) on ​​the recorded signal obtained by each of the plurality of microphones to emphasize the component in the azimuth of the voice propagation of the target object in the recorded signal.

[0150] In this way, by using the recorded signal and motion information obtained as described above, it is possible to accurately achieve distinction between the speech uttered by the target subject and the speech uttered by other subjects.

[0151] Furthermore, for example, the section detecting unit 32 may reduce the component of the voice uttered by the target subject and included in the recorded signal by using NBF (Null Beam Former).

[0152] In this case, a time interval of the target subject's speech detected from the pre-component-reduced recorded signal is compared with a time interval of the target subject's speech detected from the post-component-reduced recorded signal. Thereafter, the time interval detected from the pre-component-reduced recorded signal but not from the post-component-reduced recorded signal is determined as the final time interval of the target subject's speech.

[0153] Next, the processing performed by the signal processing unit 41 will be described in more detail.

[0154] The signal processing unit 41 performs signal processing based on the sound source classification section information, motion information, position information, and recorded signal obtained by the section detection unit 32 to generate a target sound source signal that is an audio signal classified for each target sound source.

[0155] For example, the signal processing unit 41 performs, as signal processing on the recorded signal, sound quality correction processing, sound source separation processing, noise removal processing, distance correction processing, sound source replacement processing, or a combination of these processing.

[0156] More specifically, for example, what is performed as sound quality correction processing is processing for improving the sound quality (sound quality) of the object sound source by reducing sounds other than the target (for example, noise generated at the contact portion between the recording device 11 and the object due to the movement of the object).

[0157] Specifically, examples of sound quality correction processing include noise reduction processing, such as filtering processing and gain correction for reducing the noise-dominant frequency band, and processing for muting intervals containing a lot of noise, unnecessary intervals, and intervals containing inappropriate voices during content viewing and listening.

[0158] Incidentally, it is conceivable that a time section containing inappropriate speech is detected based on sound source classification section information or by performing speech recognition processing on a recorded signal, for example.

[0159] Furthermore, for example, the sound quality correction processing may include filtering to enhance the high-frequency components in a time interval containing the sound of the target sound source, where the high-frequency band of the recorded signal is susceptible to attenuation, thereby improving the quality of the sound of the target sound source. In this case, for example, the sound quality correction processing may be performed by performing processing configured for each target sound source category based on the sound source category information for each time interval of the recorded signal.

[0160] Furthermore, for example, the time interval in which the sounds of a plurality of target sound sources are included in the recorded signal may be specified with reference to the sound source classification interval information.

[0161] Therefore, based on the results of this specification, sound source separation processing based on independent component analysis can be performed on the recorded signal to separate the sounds of each target sound source based on the difference in amplitude value and probability density distribution of each target sound source classification.

[0162] Furthermore, by performing beamforming or the like as sound source separation processing based on the direction difference between the target sound sources as viewed from the object, the sound signals of the respective target sound sources can be separated from the recorded signal.

[0163] Furthermore, when a time section of the recorded signal includes sound of only one target sound source specified by the sound source classification section information, a process of cutting out the signal in the time section as the target sound source signal is executed as the sound source separation process.

[0164] These processes allow a signal containing the sound of only one object sound source to be acquired, and this signal can be used as the object sound source signal.

[0165] In addition, when unnecessary sounds mainly including static noise such as background noise or cheering and wind noise are included in the time interval of the sound of the object sound source in the recorded signal, a process of reducing the noise included in the time interval can be performed as a noise removal process, which is similar to the sound quality correction process.

[0166] Furthermore, for example, whether there are other objects around the target object, the relative positions of the other objects relative to the target object, and the distances between the target object and the other objects can be specified based on the position information and motion information associated with each object.

[0167] As a result, based on these designation results and the sound source classification interval information, it is possible to determine whether the time interval containing the sound of the target object's target sound source contains the sound of other objects. Therefore, for example, by performing sound source separation using a DNN, it is possible to extract (separate) only the sound of the target object's target sound source.

[0168] Note that, in order to perform the above-described sound source separation and the like, for example, other objects located within a circular region R11 having a predetermined radius, the center of which is located at Figure 5 At the target object OB11 depicted in FIG. Figure 5 Each point in represents an object.

[0169] Furthermore, processing such as sound source separation is performed to remove the sound of the target object in the time interval containing the sound of the target sound source in the recorded signal of target object OB11, taking into account the distance from the target object and the relative orientation of the target object. In other words, the signal of the sound of the target sound source of target object OB11 is extracted.

[0170] At this time, based on the position information associated with these objects, the distance from the target object OB11 to the removal target object can be obtained. In addition, the relative orientation of the removal target object viewed from the target object OB11 can be obtained based on the direction indicated by the motion information and position information associated with these objects.

[0171] Furthermore, an object located outside the region R11 , that is, an object located a predetermined distance or longer from the target object OB11 , is not set as a removal target object.

[0172] The sound generated by a distant object and mixed into the recorded signal of the target object OB11 decreases due to distance attenuation. Therefore, there is no need to consider the voice or action sound generated by such an object, and the object is not removed.

[0173] In addition, when removing (separating) the sound of the removal target object, the gain or intensity of the sound of the removal target object during separation can be changed according to the distance from the target object OB11 to the removal target object. In other words, the mixing volume (contribution ratio) can be treated as a continuously variable factor according to the distance.

[0174] Furthermore, for example, the distance correction processing performed as signal processing is processing for correcting the influence generated by distance attenuation or transfer characteristics from the target sound source to the position of the microphone 21 and convoluted in the absolute sound pressure of the sound emitted from the target sound source during recording.

[0175] Specifically, for example, a process of adding an inverse feature of a transfer feature from the target sound source to the microphone 21 to the collected signal may be performed as the distance correction process.

[0176] In this way, the sound quality degradation of the sound of the object sound source caused by distance attenuation, transfer characteristics, etc. can be corrected, and the relative relationship between the absolute sound pressures of the sounds of the respective object sound sources according to the positional relationship between the respective object sound sources can be restored when the content is reproduced.

[0177] In addition, for example, the sound source replacement processing performed as signal processing is a processing in which a sound of a predetermined target sound source classification indicated by the sound source classification section information is replaced with a sound different from the recorded sound (for example, a sound prepared in advance), and the replaced sound is used as the target sound source signal.

[0178] In other words, in the sound source replacement process, a partial section of the recorded signal or a partial section of the target sound source signal obtained from the recorded signal is replaced with another audio signal prepared in advance or generated dynamically based on the sound source classification section information.

[0179] For example, here, based on the target sound source classification, a pre-prepared sound signal with a high S / N ratio can be used as the target sound source signal. This sound source replacement process is particularly effective when the amplitude of the sensor value serving as motion information is large (i.e., when the object's motion is large) and the sound quality of the recorded sound of the target sound source is low. Therefore, for example, whether to perform sound source replacement processing can be determined based on the results of threshold processing performed on the motion information.

[0180] Furthermore, in the sound source replacement process, for example, a sound signal parameterized and generated by replacing acceleration with motion information of a function may be used as the target sound source signal.

[0181] Furthermore, in the sound source replacement process, for example, when a time interval containing inappropriate speech during content viewing and listening exists as the sound of the target sound source, a signal of a predetermined sound prepared in advance can be used as the target sound source signal in the time interval.

[0182] Note that the object sound source signal obtained by the signal processing unit 41 may be a signal only in the time interval containing the sound of the object sound source, or a signal corresponding to the entire time interval but appearing as a silent signal in the time interval not containing the sound of the object sound source.

[0183] Furthermore, the aforementioned sound quality correction, sound source separation, noise removal, distance correction, and sound source replacement processes can be performed in any of the following situations: online processing for each frame of the recorded signal; processing using forward frames; offline processing; and other situations. In this case, as needed, it is sufficient to retain the recorded signal, sound source classification interval information, motion information, position information, and the like for frames preceding the target frame of the recorded signal.

[0184] <Description of the Inclusion Process>

[0185] Next, the operations of the incorporating device 11 and the server 12 will be described.

[0186] First, a description will be given of the operation of the recording device 11. The recording device 11 is attached to an object, and performs recording processing in a predetermined time period (for example, a time period in which the object is performing a performance or playing a game).

[0187] The following will refer to Figure 6 The flowchart in describes the ingestion processing performed by the ingestion device 11.

[0188] In step S11 , the recording unit 24 records ambient sounds.

[0189] Specifically, when the microphone 21 collects ambient sound and outputs a resultant recorded signal, the recording unit 24 acquires the recorded signal output from the microphone 21 to obtain a recorded signal of the recorded sound.

[0190] In step S12 , the recording unit 24 acquires motion information and position information from the motion measurement unit 22 and the position measurement unit 23 , respectively.

[0191] The collection unit 24 performs AD conversion or other processing on the collection signal, motion information, and position information obtained in the above-described manner as necessary, and supplies the thus processed signal and information to the transmission unit 25 .

[0192] Furthermore, the transmission unit 25 generates transmission data including the recorded signal, motion information, and position information supplied from the recording unit 24. At this time, the transmission unit 25 performs compression processing on the recorded signal, motion information, and position information as necessary.

[0193] In step S13 , the transmission unit 25 sends the transmission data to the server 12 .

[0194] Note that an example will be described here in which transmission data obtained by ingestion is sequentially transmitted in real time (online) to the server 12 during ingestion. However, transmission data may be accumulated during ingestion and all transmitted together offline to the server 12 after ingestion.

[0195] In step S14, the recording unit 24 determines whether to end the processing. For example, when an instruction to end recording is issued by operating a button (not shown) provided on the recording device 11, the processing is determined to be ended.

[0196] In a case where it has not been determined in step S14 that the processing is to be ended, the processing returns to step S11 to repeat the above-described processing.

[0197] On the other hand, in the case where it is determined in step S14 that the processing is to be ended, the respective units of the incorporating apparatus 11 stop the current operation, and the incorporating processing ends.

[0198] By performing the processing in the above manner, the recording device 11 collects sound and measures the motion and position of the object, and then sends transmission data containing the recording signal, motion information, and position information to the server 12. In this way, the server 12 is allowed to obtain high-quality target sound.

[0199] <Description of Data Generation Processing>

[0200] In addition, when transmission data is sent from each recording device 11 to the server 12, the server 12 performs data generation processing to output the target sound source data. Figure 7 The flowchart in describes the data generation process performed by the server 12.

[0201] In step S41 , the receiving unit 31 receives the transmission data transmitted from the recording device 11 .

[0202] Furthermore, the receiving unit 31 performs decompression processing on the received transmission data as needed to extract the collection signal, motion information, and position information from the transmission data.

[0203] Thereafter, the receiving unit 31 supplies the recorded signal and motion information to the section detecting unit 32 , supplies the recorded signal, motion information, and position information to the signal processing unit 41 , and supplies the motion information and position information to the metadata generating unit 42 .

[0204] In step S42 , the section detection unit 32 generates sound source classification section information for each object (recording device 11 ) based on the recorded signal and motion information associated with the corresponding object provided from the receiving unit 31 , and provides the generated sound source classification section information to the signal processing unit 41 .

[0205] For example, the interval detection unit 32 specifies the target sound source class included in each time interval by performing threshold processing on the recorded signal, assigning the recorded signal and motion information to an identifier such as a DNN for calculation, and performing DS or NBF on the recorded signal in the aforementioned manner. Thus, the interval detection unit 32 generates sound source class interval information.

[0206] Furthermore, the section detecting unit 32 generates sound source classification information indicating the target sound source classification of sounds and objects included in the recorded signal based on the specified result of the target sound source classification included in each time section, and supplies the generated sound source classification information to the metadata generating unit 42 .

[0207] In step S43 , the signal processing unit 41 generates a target sound source signal based on the recorded signal, motion information, and position information supplied from the receiving unit 31 , and based on the sound source classification section information supplied from the section detecting unit 32 .

[0208] Specifically, the signal processing unit 41 performs the aforementioned sound quality correction, sound source separation, noise removal, distance correction, and sound source replacement processing on the recorded signal as needed to generate a target sound source signal for each object. In this case, the target sound source signal is generated not only by using the motion information, position information, and sound source classification interval information associated with the target object, but also by using the motion information, position information, and sound source classification interval information associated with other objects.

[0209] In step S44 , the metadata generating unit 42 generates metadata including the sound source classification information supplied from the section detecting unit 32 and the motion information and position information supplied from the receiving unit 31 for each target sound source of the object.

[0210] When the object sound source signal and metadata for each object sound source are obtained in this manner, the object sound source data generating unit 33 outputs the object sound source data including the object sound source signal and metadata to the next stage of each object sound source.

[0211] In step S45, the server 12 determines whether to end the processing. For example, in the case where the processing of all the transmission data received from the ingestion device 11 is completed, it is determined in step S45 that the processing is to be ended.

[0212] In the event that it has not been determined in step S45 that the processing is to be ended, the processing then returns to step S41 to repeat the above-described processing.

[0213] On the other hand, in the case where it is determined in step S45 that the processing is to be ended, each unit of the server 12 stops the processing currently being executed, and the data generation processing ends.

[0214] Note that, an example has been described herein in which transmission data is sequentially transmitted from the recording apparatus 11 in real time and in which target sound source data is also sequentially generated from the transmission data using the server 12 .

[0215] However, the transmission data received from the collection device 11 may be accumulated and the accumulated transmission data may be collectively processed to generate the target sound source data. In addition, when the transmission data is collectively sent from the collection device 11, only the received transmission data needs to be collectively processed to generate the target sound source data.

[0216] In the above-described manner, the server 12 receives transmission data from the plurality of recording devices 11 , generates target sound source data using these pieces of transmission data, and outputs the generated target sound source data.

[0217] At this time, by generating sound source classification section information using not only the recorded signal but also motion information and generating target sound source data using the generated sound source classification section information, a high-quality target sound, ie, a high-quality target sound signal, can be acquired.

[0218] <Second embodiment>

[0219] <Configuration example of the recording system>

[0220] Note that the above has described an example in which information obtained for other objects is not used when generating sound source classification section information associated with each object. However, for example, the accuracy of sound source classification section information can be improved by integrating information obtained for each object.

[0221] In this case, for example, Figure 8 The configuration described in the collection system. Note that Figure 8 Zhongyu Figure 1 Corresponding parts in are given the same reference symbols, and descriptions of these parts are omitted where appropriate.

[0222] Figure 8 The collection system described in the embodiment includes a collection device 11 and a server 12. The collection device 11 has Figure 1 Same configuration as depicted in .

[0223] on the other hand, Figure 8 The server 12 of the recording system shown in FIG. 1 includes a receiving unit 31 , a section detecting unit 32 , an integrating unit 71 , and a target sound source data generating unit 33 . Furthermore, the target sound source data generating unit 33 includes a signal processing unit 41 and a metadata generating unit 42 .

[0224] The configuration of server 12 here is the same as Figure 1The server 12 depicted in FIG. 1 differs in that an integrated unit 71 is provided and is otherwise identical to the server 12 depicted in FIG. Figure 1 The configuration of the server 12 is the same as that depicted in FIG.

[0225] In this example, the sound source classification section information generated by the section detection unit 32 is supplied to the integration unit 71. In addition to the sound source classification section information from the section detection unit 32, the integration unit 71 is supplied with a recorded signal, motion information, and position information from the reception unit 31.

[0226] The integration unit 71 generates final sound source classification interval information based on the received recorded signal, sound source classification interval information, motion information, and position information, and supplies the final sound source classification interval information to the signal processing unit 41. The integration unit 71 also generates sound source classification information and supplies the sound source classification information to the metadata generation unit 42.

[0227] Specifically, the integration unit 71 integrates various information, such as motion information and position information obtained by each recording device 11, to generate more accurate sound source classification interval information.

[0228] Note that the following will describe an example in which the integration unit 71 is provided separately from the section detection unit 32. However, the integration unit 71 may be provided on the section detection unit 32. In this case, the section detection unit 32 performs the following processing, which is described later, performed by the integration unit 71, simultaneously with the above-mentioned processing to generate sound source classification section information and sound source classification information.

[0229] The integrated unit 71 will be described in more detail here.

[0230] For example, the section detecting unit 32 detects, for each object (ie, each recording device 11 ), a time section estimated to include the object's motion sound or voice, and generates sound source classification section information based on the estimated time section.

[0231] However, in this case, a time interval containing the action sound or voice of another object may be mistakenly detected as a time interval containing the action sound or voice of the target object, or a time interval containing the action sound or voice of the target object that needs to be detected may not be detected.

[0232] Therefore, the integration unit 71 integrates the information obtained by each recording device 11 to generate more accurate sound source classification section information.

[0233] Specifically, for example, the integration unit 71 performs position information comparison processing, time interval integration processing and interval smoothing processing for each frame with a predetermined time length based on the sound source classification interval information, recorded signal, motion information and position information to obtain the final sound source classification interval information.

[0234] In other words, the integration unit 71 generates sound source classification section information associated with the target object based on the recorded signals, motion information, and position information associated with other objects and at least any one of the recorded signals, motion information, and position information associated with the target object.

[0235] Examples of the above-mentioned position information comparison processing, time interval integration processing, and interval smoothing processing will be further described below.

[0236] First, all objects are selected as target objects in sequence, and position information comparison processing, time interval integration processing, and interval smoothing processing are performed on each target object.

[0237] In the position information comparison process, the distance between the target object and other objects is calculated based on the position information associated with each object.

[0238] Thereafter, based on the calculated distance, other objects that may affect the sound of the object sound source of the target object, ie, other objects located near the target object, are selected as reference objects.

[0239] Specifically, for example, an object located at a predetermined threshold distance or less from the target object is selected as a reference object. In this example, the recording device 11 is attached to each object, so each distance between the recording devices 11 is substantially equivalent to each distance between the objects. Therefore, the distance calculated based on the position information is used to select the reference object.

[0240] Note that an example will be described here in which a reference object is selected based on a distance and in which a time interval integration process is performed using information associated with the selected reference object.

[0241] However, all objects may be used as reference objects, and the time interval integration process may be performed using information associated with the reference objects weighted according to respective distances from the target object.

[0242] In the time interval integration process, it is initially determined whether or not there is an object selected as a reference object by the position information comparison process.

[0243] Thereafter, if no object is selected as a reference object, the sound source classification interval information associated with the target object and obtained by the interval detection unit 32 is output unchanged to the signal processing unit 41 as final sound source classification interval information. This information is output because the sounds of other objects are not mixed into the recorded signal if no other objects are near the target object.

[0244] On the other hand, if there is an object selected as a reference object, the sound source classification section information associated with the target object is also updated using the position information and motion information associated with the selected reference object. In other words, final sound source classification section information is generated.

[0245] Specifically, a reference object having a time interval overlapping with the time interval indicated by the sound source classification interval information associated with the target object as a time interval containing the sound of the target sound source is selected from the respective reference objects as the final reference object.

[0246] Specifically, even if an object is selected as a reference object through position information comparison processing, if the selected object has a time interval indicated by the sound source classification interval information, and the time interval does not overlap with the time interval indicated by the sound source classification interval information associated with the target object, the object is excluded from the reference object.

[0247] Subsequently, based on the position information and motion information associated with the reference object, and based on the position information and motion information associated with the target object, the relative orientation (direction) of the reference object as viewed from the target object in three-dimensional space is estimated. Relative orientation information indicating the result of this estimation is then generated. More specifically, for example, the direction (orientation) of the reference object's mouth as viewed in front of the target object is estimated. Note that relative orientation information can be generated using only position information without using motion information.

[0248] Furthermore, an NBF filter is formed based on position information associated with the target object, a direction of the target object indicated by the motion information, and relative orientation information associated with respective reference objects.

[0249] The NBF filter is a filter that implements beamforming for reducing the sound coming in the direction indicated by the relative orientation information while maintaining the gain of the sound coming in the direction of the mouth of the target object indicated by the direction of the target object.

[0250] The integration unit 71 performs convolution processing for convolving the NBF filter obtained in this manner with a time section included in the recorded signal of the target object and indicated by the sound source classification section information associated with the target object.

[0251] Furthermore, the integration unit 71 performs processing similar to that performed by the section detection unit 32, based on the signal obtained through the convolution processing and the motion information associated with the target object, such as threshold processing and calculation processing using an identifier such as a DNN, to generate sound source classification section information. In this way, the sound emitted from the reference object is reduced, and therefore, more accurate sound source classification section information can be obtained.

[0252] Note that motion information, position information, recorded signals, and the like associated with the reference object may be input to an identifier such as a DNN to perform calculation processing.

[0253] Finally, the integration unit 71 performs interval smoothing processing on the sound source classification interval information obtained through the time interval integration processing to obtain final sound source classification interval information.

[0254] For example, the average time of the minimum durations of sounds generated from the corresponding classes is obtained in advance as the average minimum duration of each object sound source class.

[0255] In the interval smoothing processing, smoothing is performed using a smoothing filter that connects the segmented (divided) time intervals of each sound containing the object sound source so that the length of each time interval of the detected sound containing the object sound source lasts for an average minimum duration or longer.

[0256] In other words, in the interval smoothing process, multiple consecutively aligned time intervals, each containing detected sounds of the target sound source of the same classification in the recorded signal, are concatenated into a single final time interval. In this case, the multiple time intervals to be concatenated include at least one time interval whose duration is shorter than the average minimum duration.

[0257] For example, the integration unit 71 retains in advance a smoothing filter formed based on the average minimum duration of each object sound source class.

[0258] As a section smoothing process, the integration unit 71 performs filtering (filtering process) on the sound source classification section information obtained through the time section integration process using a smoothing filter to obtain final sound source classification section information. The integration unit 71 then provides the final sound source classification section information to the signal processing unit 41. In the section smoothing process, in some cases, filtering is performed on the sound source classification section information associated with multiple consecutive frames based on the target sound source classification (i.e., the average minimum duration).

[0259] Furthermore, the integration unit 71 generates sound source classification information based on the obtained sound source classification section information, and supplies the generated sound source classification information to the metadata generation unit 42 .

[0260] In the above manner, the integration unit 71 removes information associated with sounds of other objects that are not removed (excluded) based on the sound source classification section information obtained by the section detection unit 32 , thereby making it possible to obtain more accurate sound source classification section information.

[0261] For example, depending on the situation, the section detection unit 32 performs DS or NBF on the recorded signal as needed, as described above.

[0262] However, for example, the DS may not be able to sufficiently emphasize the component in the direction from which the target object's voice comes. In this case, when the volume of the voice of other objects is loud, it may be difficult to obtain correct sound source classification section information.

[0263] In addition, for example, if another object is located near the target object and in a direction close to the direction of the target object's sound, and makes a sound almost simultaneously with the target object, NBF may not be able to obtain accurate sound source classification section information.

[0264] On the other hand, the integration unit 71 can obtain more accurate sound source classification section information by using not only information associated with the target object but also motion information, position information, and sound source classification section information associated with other objects.

[0265] <Description of Data Generation Processing>

[0266] In the recording system Figure 8 In the case of the configuration depicted, each recording device 11 performs reference Figure 6 The described collection processing is performed and the transmission data is sent to the server 12.

[0267] Thereafter, the server 12 executes Figure 9 The data generation process described in . Figure 9 The flowchart in the description is given by Figure 8 The data generation process is performed by the server 12 depicted in FIG.

[0268] Note that the processing in step S71 and step S72 is similar to Figure 7 The processing in step S41 and step S42 in is not repeated here.

[0269] However, in step S71 , the collection signal, motion information, and position information extracted from the transmission data by the receiving unit 31 are also supplied to the integrating unit 71 .

[0270] Furthermore, in step S72 , the generated sound source classification section information is supplied from the section detection unit 32 to the integration unit 71 .

[0271] In step S73 , the integration unit 71 integrates the information supplied from the section detection unit 32 and the reception unit 31 .

[0272] Specifically, the integration unit 71 performs position information comparison processing, time interval integration processing and interval smoothing processing based on the recorded signal, motion information and position information provided by the receiving unit 31, and based on the sound source classification interval information provided by the interval detection unit 32 to obtain the final sound source classification interval information.

[0273] The integration unit 71 supplies the obtained final sound source classification section information to the signal processing unit 41 , generates sound source classification information based on the final sound source classification section information, and supplies the generated sound classification information to the metadata generation unit 42 .

[0274] After obtaining the sound source classification interval information in this manner, the processing in steps S74 to S76 is performed. Thereafter, the data generation processing ends. These processing are similar to Figure 7 The processing in steps S43 to S45 in is described above, so the description is not repeated.

[0275] In the above-described manner, the server 12 receives transmission data from the plurality of recording devices 11 , generates target sound source data using the transmission data, and outputs the generated target sound source data.

[0276] At this time, by generating final sound source classification section information associated with the target object using information associated with other objects, a higher quality target sound can be obtained.

[0277] <Third embodiment>

[0278] <Configuration example of the recording system>

[0279] Furthermore, according to the above description, the recorded signal and position information are used to generate the sound source classification section information. However, image information can also be used for this purpose.

[0280] In this case, for example, Figure 10 The configuration described in the collection system. Note that Figure 10 Zhongyu Figure 8 Corresponding parts in are given the same reference symbols, and descriptions of these parts are omitted where appropriate.

[0281] Figure 10 The collection system depicted in FIG includes a collection device 11 and a server 12 .

[0282] In this example, the recording device 11 includes a microphone 21 , a motion measurement unit 22 , a position measurement unit 23 , an imaging unit 101 , a recording unit 24 , and a transmission unit 25 .

[0283] Figure 10 The configuration of the recording device 11 depicted in Figure 8 The configuration of the recording device 11 depicted in FIG is different in that an imaging unit 101 is provided, and is otherwise the same as Figure 8 The configuration of the recording device 11 is the same as that depicted in .

[0284] The imaging unit 101 includes a small camera and is configured to capture an image containing a portion of an object as a subject, for example, from a viewpoint corresponding to the position of the object, and supply the obtained image information (image signal) to the transmission unit 25. Note that in some cases, an image based on the image information does not contain the object as a subject.

[0285] The transmission unit 25 generates transmission data containing the collection signal, motion information, and position information supplied from the collection unit 24 , and the image information supplied from the imaging unit 101 , and transmits the generated transmission data to the server 12 .

[0286] Meanwhile, the server 12 includes a receiving unit 31 , a section detecting unit 32 , an integrating unit 71 , and an object sound source data generating unit 33 . The object sound source data generating unit 33 includes a signal processing unit 41 and a metadata generating unit 42 .

[0287] in this case, Figure 10 The configuration of the server 12 depicted in Figure 8 The configuration of the server 12 is the same as that depicted in FIG. However, Figure 10 The server 12 depicted in FIG provides the image information extracted from the transmission data by the receiving unit 31 to the section detecting unit 32 and the integrating unit 71 .

[0288] Therefore, the section detection unit 32 generates sound source classification section information based on the recorded signal, motion information, and image information supplied from the reception unit 31 .

[0289] For example, when an image based on image information includes a part of a target object as a subject, the image information is used to detect the motion of the target object.

[0290] Specifically, for example, based on image information, the sound source classification section information is corrected based on the motion of the target object detected at each time point.

[0291] Alternatively, for example, calculation may be performed by assigning image information, motion information, and recorded signals to identifiers such as DNN to obtain the presence or absence of an action sound at each time point in the recorded signal.

[0292] Similarly, the integration unit 71 also performs position information comparison processing, time interval integration processing, and interval smoothing processing based on the recorded signal, motion information, position information, image information, and sound source classification interval information.

[0293] At this time, similar to the case of the interval detection unit 32, the image information can be used to detect the movement of the target object, time interval integration processing, etc., or can be used to detect whether there are other objects around the target object, detect the movement of other objects, etc.

[0294] <Description of the Inclusion Process>

[0295] Next, we will describe Figure 10 The operations of the recording device 11 and the server 12 are depicted in FIG.

[0296] First, refer to Figure 11 The flowchart in describes the ingestion processing performed by the ingestion device 11.

[0297] Note that the processing in steps S101 and S102 is similar to Figure 6 The processing in step S11 and step S12 in is described above, so the description is not repeated.

[0298] In step S103 , the imaging unit 101 captures an image of an object (ie, the surroundings of the recording device 11 ) as a subject, and supplies image information obtained thereby to the transmission unit 25 .

[0299] The transmission unit 25 generates transmission data including the image information supplied from the imaging unit 101 and the recorded signal, motion information, and position information supplied from the recording unit 24 .

[0300] After the transmission data is generated, the processing in step S104 and step S105 is executed. Thereafter, the collection processing ends. These processing are similar to Figure 6 The processing in step S13 and step S14 in is not repeated here.

[0301] In the above manner, the recording device 11 captures images of surrounding subjects, generates transmission data containing the obtained image information, and transmits the generated transmission data to the server 12. In this manner, the server 12 is allowed to use not only motion information and position information but also image information to obtain higher-quality target sounds.

[0302] <Description of Data Generation Processing>

[0303] The following will refer to Figure 12 The flowchart in the description is given by Figure 10 The data generation process is performed by the server 12 depicted in FIG.

[0304] Note that the processing in step S131 is similar to Figure 9 However, in step S131, the receiving unit 31 also extracts image information from the transmission data and provides the image information to the section detecting unit 32 and the integrating unit 71.

[0305] In step S132 , the section detection unit 32 generates sound source classification section information based on the recorded signal, motion information, and image information supplied from the reception unit 31 , and supplies the generated sound source classification section information to the integration unit 71 .

[0306] Note that the same Figure 9 In this case, the image information here is used to detect the motion of the target object or the like to generate the sound source classification section information.

[0307] In step S133 , the integration unit 71 integrates the information supplied from the section detection unit 32 and the reception unit 31 to generate final sound source classification section information.

[0308] In step S133, execute Figure 9 This process is similar to the process in step S73 in [ ]. However, here, not only the sound source classification interval information, the recorded signal, motion information, and position information are used, but also image information is used to perform position information comparison, time interval integration, and interval smoothing. In other words, image information is used, for example, to select a reference object or generate relative orientation information.

[0309] After the final sound source classification interval information is obtained in this manner, the processing in steps S134 to S136 is performed. Thereafter, the data generation processing ends. These processes are similar to Figure 9 The processing in steps S74 to S76 in is described above, so the description is not repeated.

[0310] In the above-described manner, the server 12 receives transmission data from the plurality of recording devices 11 , generates target sound source data using the transmission data, and outputs the generated target sound source data.

[0311] At this time, by also using image information to generate sound source classification section information associated with the target object, a higher quality target sound can be obtained.

[0312] <Fourth embodiment>

[0313] <Configuration example of the recording system>

[0314] exist Figure 10 In the recording system described in

[15] , an example has been described in which image information obtained from a viewpoint corresponding to the position of each object is used. However, image information associated with an image of the entire target space as a subject, in which the objects to which the recording devices 11 are attached, that is, all objects, exist, may be used.

[0315] In this case, for example, Figure 13 The configuration described in the collection system. Note that Figure 13 Zhongyu Figure 8 Corresponding parts in are given the same reference symbols, and descriptions of these parts are omitted where appropriate.

[0316] Figure 13 The collection system described in the embodiment includes a collection device 11, an imaging device 131 and a server 12. The collection device 11 and the server 12 each have Figure 8 The same configuration as the corresponding configuration depicted in .

[0317] For example, the imaging device 131 includes a camera or the like, and is configured to capture an image of the entire target space (in which the object to which the recording device 11 is attached exists) as a subject, and transmit the image information obtained thereby to the server 12. Note that while the recording device 11 is recording, that is, while the microphone 21 is collecting sound, the imaging device 131 continues imaging.

[0318] Furthermore, the receiving unit 31 of the server 12 receives not only the transmission data transmitted from the collection device 11 but also the image information transmitted from the imaging device 131 .

[0319] The receiving unit 31 supplies the received image information to the integrating unit 71. The integrating unit 71 also generates final sound source classification section information based on the recorded signal, motion information, position information, and image information supplied from the receiving unit 31 and the sound source classification section information supplied from the section detecting unit 32.

[0320] In this example, the integration unit 71 detects the motion of each object using image information.

[0321] For example, the integration unit 71 is provided with positional information associated with each object, and can therefore, based on this positional information, specify which object corresponds to each object in an image obtained by performing image recognition, etc. on the image information. Furthermore, the integration unit 71 can specify which action is to be performed by each object by performing image recognition, etc. on the image information. In other words, the integration unit 71 can specify which sound of the object sound source is to be emitted from the corresponding object at each point in time.

[0322] The integration unit 71 generates final sound source classification section information by performing time section integration processing using the actions of each object specified in the above manner. In addition, for example, image information can be input to an identifier such as a DNN for calculation processing performed in the time section integration processing.

[0323] Note that the section detection unit 32 may also detect the motion of each object using image information.

[0324] <Description of Data Generation Processing>

[0325] In the recording system Figure 13 In the case of the configuration depicted, each recording device 11 performs reference Figure 6 The image forming apparatus 131 transmits image information to the server 12.

[0326] Thereafter, the server 12 executes Figure 14 The data generation process described in . Figure 14 The flowchart in the description is given by Figure 13 The data generation process is performed by the server 12 depicted in FIG.

[0327] In step S161 , the receiving unit 31 receives image information transmitted from the imaging device 131 , and supplies the image information to the integrating unit 71 .

[0328] Furthermore, the transmission data is sent from the recording device 11 to the server 12 , and the server 12 performs the processing in steps S162 and S163 to generate sound source classification section information.

[0329] Note that the processing in step S162 and step S163 is similar to Figure 9 The processing in step S71 and step S72 in is described above, so the description is not repeated.

[0330] In step S164 , the integration unit 71 performs information integration.

[0331] Specifically, the integration unit 71 performs position information comparison processing, time interval integration processing, and interval smoothing processing based on the image information, recorded signal, motion information, and position information provided by the receiving unit 31, and the sound source classification interval information provided by the interval detection unit 32, to obtain final sound source classification interval information. At this time, for example, the image information is used to select a reference object.

[0332] The integration unit 71 supplies the obtained final sound source classification section information to the signal processing unit 41 , generates sound source classification information based on the final sound source classification section information, and supplies the generated sound classification information to the metadata generation unit 42 .

[0333] After the sound source classification section information is obtained in this manner, the processing in steps S165 to S167 is performed, and the data generation processing ends. These processes are similar to Figure 9 The processing in steps S74 to S76 in is described above, so the description is not repeated.

[0334] In this manner, the server 12 receives transmission data from the plurality of recording devices 11 and image information from the imaging device 131, uses this transmission data and image information to generate target sound source data, and then outputs the generated target sound source data. This, by using image information, also allows for the acquisition of higher-quality target sounds.

[0335] <Fifth embodiment>

[0336] <Configuration example of the recording system>

[0337] Note that, the example has been described above in which the sound source classification section information is generated using the server 12. However, the sound source classification section information may be generated using the recording apparatus 11.

[0338] In this case, for example, Figure 15 As depicted in FIG, the above-mentioned section detection unit 32 is provided on the recording device 11 side. Note that Figure 15 Zhongyu Figure 8 Corresponding parts in are given the same reference symbols, and descriptions of these parts are omitted where appropriate.

[0339] Figure 15 The collection system depicted in FIG includes a collection device 11 and a server 12 .

[0340] Furthermore, the recording device 11 includes a microphone 21 , a motion measurement unit 22 , a position measurement unit 23 , a recording unit 24 , a section detection unit 32 , and a transmission unit 25 .

[0341] Figure 15 The configuration of the recording device 11 depicted in Figure 8 The configuration of the recording device 11 depicted in FIG is different in that a section detection unit 32 is provided, and is otherwise the same as Figure 8 The configuration of the recording device 11 is the same as that depicted in .

[0342] The section detection unit 32 generates sound source classification section information based on the recorded signal and motion information supplied from the recording unit 24 , and supplies the sound source classification section information obtained thereby and the recorded signal, motion information, and position information supplied from the recording unit 24 to the transmission unit 25 .

[0343] The transmission unit 25 generates transmission data including the recorded signal, motion information, position information, and the sound source classification section information supplied from the section detection unit 32 , and transmits the transmission data to the server 12 .

[0344] Meanwhile, the server 12 includes a receiving unit 31, an integrating unit 71, and an object sound source data generating unit 33. The object sound source data generating unit 33 includes a signal processing unit 41 and a metadata generating unit 42.

[0345] The configuration of server 12 here is the same as Figure 8 The server 12 depicted in FIG. 1 is different in that the interval detection unit 32 is not provided, and is otherwise the same as Figure 8 The configuration of the server 12 is the same as that depicted in FIG.

[0346] exist Figure 15 In the example depicted in , the receiving unit 31 of the server 12 extracts the recorded signal, motion information, position information, and sound source classification section information from the received transmission data.

[0347] Thereafter, the receiving unit 31 provides the recorded signal, motion information, position information and sound source classification interval information to the integration unit 71 , provides the recorded signal, motion information and position information to the signal processing unit 41 , and provides the motion information and position information to the metadata generation unit 42 .

[0348] Furthermore, the integration unit 71 generates final sound source classification interval information based on the recorded signal, motion information, position information, and sound source classification interval information provided by the receiving unit 31, and provides the final sound source classification interval information to the signal processing unit 41. Furthermore, the integration unit 71 generates sound source classification information and provides the sound source classification information to the metadata generation unit 42.

[0349] By generating the sound source classification section information on the recording device 11 side in the above manner, it is possible to reduce the processing load imposed on the server 12 and obtain high-quality target sounds. Figure 10 or Figure 13 In the recording system described in , sound source classification section information can be generated on the recording device 11 side.

[0350] According to the present technology, as described above, in an environment where multiple moving bodies (objects) exist and each of them emits sound, the sound of the target object included in the recorded signal can be distinguished from the sounds of other objects using motion information, position information, and image information.

[0351] In this way, it is possible to detect a time interval containing a sound classified for each object sound source, perform signal processing for each object sound source classification, and recognize behavior of the action state of each object.

[0352] For example, the time interval of the sound classified as each object sound source can accurately detect the time interval of walking sounds, running sounds, soccer kicking sounds, baseball batting sounds or catching sounds, applause, rustling sounds of clothes, or dancing footsteps.

[0353] Typically, action sounds cannot be acquired solely from sensor signals. Furthermore, information related to the direction of speech or speaker personality (personality) is required to distinguish the same type of action sounds generated from the target object and from other objects and included in the recorded signal from the microphone.

[0354] In this regard, compared to the case of using only sensor signals or only recorded signals, the present technology can accurately detect the time period of the sound of the target sound source and obtain a high-quality target sound source signal.

[0355] Specifically, it is assumed that when only the time interval of the action sound is detected from the recorded signal, the target object and another object are close to each other.

[0356] In this case, the direction of the sound source is estimated using multiple microphones, and the estimated direction is used to distinguish the action sound of the target object from the action sounds of other objects, similar to the case of speech.

[0357] However, for example, in the case where the time interval of an action sound such as a walking sound is short or the direction of the sound source changes over time, it is generally difficult to identify which object is making the action sound.

[0358] On the other hand, the motion information includes only body motion information based on the motion of the target object, and does not include information associated with the motion of other objects.

[0359] Therefore, as in the present technology, by combining the recorded signal and the motion information for detecting the time period of the action sound, the time period of the action sound of the target object can be accurately detected.

[0360] For example, when detecting the time interval of walking sounds as motion sounds, the condition of the ground or shoes significantly affects the detection accuracy when using only the recorded signal. However, by combining motion information and the recorded signal, the time interval of walking sounds can be accurately detected.

[0361] Furthermore, this technology can detect time intervals containing valid sounds from target sound sources in audio reproduction of recorded content, such as sports and games, and prevent unnecessary transmission of target sound source signals during these time intervals. This reduces the amount of information associated with the content being transmitted or recorded, particularly the amount of target sound source signals, and the amount of processing required in subsequent stages.

[0362] Furthermore, according to the present technology, an object sound source signal is generated for each object or each object sound source of an object. Therefore, in a subsequent stage, audio image localization can be set for each object sound source, thereby achieving more accurate audio image localization.

[0363] Furthermore, this technology generates a target sound source signal for each target sound source. This allows selective reproduction of only sounds classified under certain target sound sources, such as when reproducing only the action sounds and no speech in a sports broadcast. This improves playback performance.

[0364] Furthermore, according to the present technology, when real-time processing is performed by the server 12 during recording of content such as sports games, for example, when instant replay is used in the current situation, information associated with the action conditions of each athlete can be provided, and this information is effective additional information.

[0365] Specifically, as information associated with the action conditions of the player, for example, information indicating a time interval of a predetermined action sound or a time interval of voice may be provided based on the sound source classification interval information.

[0366] In addition, this technology is applicable not only to the collection of content, but also to various situations, such as when there are multiple vehicles on the road, when multiple flying objects (such as drones) are flying, and when there are multiple robots.

[0367] For example, the recording device 11 may be provided on a vehicle, and the recording signal, motion information, etc. obtained by the recording device 11 and information obtained by a drive recorder equipped on the vehicle may be used to determine contact with other vehicles.

[0368] <Computer Configuration Example>

[0369] Meanwhile, the series of processes described above can be executed by hardware or software. In the case where the series of processes are executed by software, the program constituting the software is installed in a computer. Examples of the computer here include computers incorporated in dedicated hardware and computers capable of executing various functions under various programs installed in the computer, such as general-purpose personal computers.

[0370] Figure 16 It is a block diagram depicting a hardware configuration example of a computer that executes the above-described series of processes under a program.

[0371] The computer includes a CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 , which are connected to one another via a bus 504 .

[0372] An input / output interface 505 is further connected to the bus 504 . An input unit 506 , an output unit 507 , a storage unit 508 , a communication unit 509 , and a drive 510 are connected to the input / output interface 505 .

[0373] Input unit 506 includes a keyboard, a mouse, a microphone, an imaging element, etc. Output unit 507 includes a display, a speaker, etc. Storage unit 508 includes a hard disk, a nonvolatile memory, etc. Communication unit 509 includes a network interface, etc. Drive 510 drives removable recording media 511, such as a magnetic disk, an optical disk, a magneto-optical disk, and a semiconductor memory.

[0374] In the computer configured as described above, for example, the CPU 501 loads the program recorded in the storage unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executes the loaded program to perform the above-described series of processes.

[0375] For example, the program executed by the computer (CPU 501) is allowed to be recorded in a removable recording medium 511 such as a package medium and provided in this form. In addition, the program is allowed to be provided via a wired or wireless transmission medium such as a local area network, the Internet, and digital satellite broadcasting.

[0376] The computer-enabled program is installed in the storage unit 508 from the removable recording medium 511 attached to the drive 510 via the input / output interface 505. Furthermore, the communication unit 509 receives the program via a wired or wireless transmission medium and installs it in the storage unit 508. Conversely, the program is installed in advance in the ROM 502 or the storage unit 508.

[0377] Note that the program executed by the computer may be a program that executes the processes in time series in the order described in this specification, or may be a program that executes the processes in parallel or at necessary timing (for example, upon a call).

[0378] Furthermore, the embodiment according to the present technology is not limited to the above-described embodiment, and can be modified in various ways without departing from the subject matter of the present technology.

[0379] For example, the present technology is allowed to have a configuration of cloud computing in which one function is shared and processed by a plurality of devices in cooperation with each other via a network.

[0380] Furthermore, each step described in the above flowchart is allowed to be executed by one device or to be shared and executed by a plurality of devices.

[0381] Furthermore, in the case where one step includes a plurality of processes, the plurality of processes included in one step are allowed to be executed by one device or to be shared and executed by a plurality of devices.

[0382] Furthermore, the present technology may also have the following configurations. (1)

[0384] A signal processing device, comprising:

[0385] The interval detection unit is configured to detect a time interval containing a sound emitted from a mobile object, where the sound is included in a collected signal obtained by collecting sounds around the mobile object in a state where other mobile objects are present around the mobile object, and the time interval is detected based on the collected signal and a sensor signal output from a sensor attached to the mobile object. (2)

[0387] The signal processing device according to (1), further comprising:

[0388] The data generating unit is configured to generate an audio signal of the voice or movement sound of the moving object from the collected signal based on the detection result of the time interval. (3)

[0390] The signal processing device according to (2), wherein the data generation unit outputs target sound source data including the audio signal and position information indicating the position of the moving object. (4)

[0392] The signal processing device according to (2) or (3), wherein the data generation unit outputs target sound source data including the audio signal and information indicating the direction of the moving object. (5)

[0394] The signal processing device according to any one of (2) to (4), wherein the data generation unit outputs target sound source data including the audio signal and sound source classification information indicating sound classification based on the audio signal. (6)

[0396] A signal processing device according to any one of (1) to (5), wherein the interval detection unit detects the time interval of the sound emitted from the moving body based on the recorded signal and the sensor signal of the moving body, and based on the recorded signal or the sensor signal of the other moving body. (7)

[0398] The signal processing device according to (6), wherein the section detection unit detects the time section of the sound emitted from the moving object based on a distance from the moving object to the other moving object. (8)

[0400] The signal processing device according to (6) or (7), wherein the section detection unit detects the time section of the sound emitted from the moving object based on the direction of the moving object and the position of the other moving object. (9)

[0402] A signal processing device according to any one of (6) to (8), wherein the interval detection unit obtains a final detection result of the time interval based on the detection result of the time interval by connecting multiple time intervals included in the recorded signal and containing sounds of the same classification, and the time intervals are continuously aligned and include the time intervals that are shorter than a predetermined time width. (10)

[0404] The signal processing device according to (9), wherein the section detection unit connects the plurality of time sections by performing smoothing processing on the detection results of the time sections. (11)

[0406] The signal processing device according to any one of (2) to (5), wherein the data generating unit generates the audio signal by performing sound source separation on the recorded signal based on the detection result of the time interval. (12)

[0408] The signal processing device according to any one of (2) to (5), wherein the data generating unit generates the final audio signal by replacing a part of the collected signal or a part of the audio signal with another signal based on the detection result of the time interval. (13)

[0410] A signal processing method, comprising:

[0411] By the signal processing device,

[0412] A time interval containing a sound emitted from a moving object is detected, where the sound is included in a recorded signal obtained by collecting sounds around the moving object in a state where other moving objects are present around the moving object, and the time interval is detected based on the recorded signal and a sensor signal output from a sensor attached to the moving object. (14)

[0414] A program causing a computer to execute a process, comprising:

[0415] A step of detecting a time interval including a sound emitted from a moving object, wherein the sound is included in a collected signal obtained by collecting sounds around the moving object in a state where other moving objects are present around the moving object, and detecting the time interval based on the collected signal and a sensor signal output from a sensor attached to the moving object.

[0416] Reference Symbols List

[0417] 11: Recording equipment

[0418] 12: Server

[0419] 21: Microphone

[0420] 24: Recording Unit

[0421] 25: Transmission unit

[0422] 31: Receiving unit

[0423] 32: Interval detection unit

[0424] 33: Object sound source data generation unit

[0425] 41: Signal processing unit

[0426] 42: Metadata generation unit

[0427] 71: Integrated unit.

Claims

1. A signal processing device, comprising: The interval detection unit is configured to detect a time interval containing sound emitted from the mobile body and included in the collected signal based on a collected signal obtained by collecting sound around the mobile body and a sensor signal output from a sensor attached to the mobile body in a state where another mobile body is present around the mobile body, The signal processing device further comprises: a data generating unit configured to generate an audio signal of the voice or movement sound of the moving object from the collected signal based on the detection result of the time interval, The data generating unit generates the audio signal by performing sound source separation on the recorded signal based on the detection result of the time interval.

2. The signal processing device according to claim 1, wherein The data generating unit outputs target sound source data including the audio signal and position information indicating the position of the moving object.

3. The signal processing device according to claim 1, wherein: The data generating unit outputs target sound source data including the audio signal and information indicating the direction of the moving object.

4. The signal processing device according to claim 1, wherein: The data generating unit outputs target sound source data including the audio signal and sound source classification information indicating a sound classification based on the audio signal. The signal processing device according to claim 1 , wherein: The section detection unit detects the time section of the sound emitted from the moving object based on the recorded signal and the sensor signal of the moving object and based on the recorded signal or the sensor signal of the other moving object. The signal processing device according to claim 5 , wherein: The section detection unit detects the time section of the sound emitted from the moving object based on the distance from the moving object to the other moving object.

7. The signal processing device according to claim 5, wherein: The section detection unit detects the time section of the sound emitted from the moving object based on the direction of the moving object and the position of the other moving object.

8. The signal processing device according to claim 5, wherein: The interval detection unit obtains a final detection result of the time interval by connecting a plurality of the time intervals in which sounds of the same category in the recorded signal are continuously arranged, including the time interval shorter than a predetermined time width, based on the detection result of the time interval.

9. The signal processing device according to claim 8, wherein: The section detection unit connects a plurality of the time sections by performing a smoothing process on the detection results of the time sections.

10. The signal processing device according to claim 1, wherein: The data generating unit generates a final audio signal by replacing a portion of the recorded signal or a portion of the audio signal with another signal based on the detection result of the time interval.

11. A signal processing method, comprising: By the signal processing device, In a state where another moving object exists around a moving object, based on a recorded signal obtained by collecting sounds around the moving object and a sensor signal output from a sensor attached to the moving object, a time interval containing sounds emitted from the moving object and included in the recorded signal is detected; An audio signal of the voice or movement sound of the moving object is generated from the recorded signal based on the detection result of the time interval, wherein the audio signal is generated by performing sound source separation on the recorded signal based on the detection result of the time interval.

12. A program for causing a computer to execute a process comprising the following steps: In a state where another moving object exists around a moving object, based on a recorded signal obtained by collecting sounds around the moving object and a sensor signal output from a sensor attached to the moving object, a time interval containing sounds emitted from the moving object and included in the recorded signal is detected; An audio signal of the voice or movement sound of the moving object is generated from the recorded signal based on the detection result of the time interval, wherein: The audio signal is generated by performing sound source separation on the recorded signal based on the detection result of the time interval.

Citation Information

Patent Citations

  • Information processing device, information processing method and program

    JP2017205213A

  • Endpoint detection apparatus for sound source and method thereof

    US20130238335A1