Program, device and method for adjusting acoustic signals using vibration signals

The audio signal adjustment program and device utilize vibration signals to process sound data efficiently, addressing the challenge of real-time generation of movement sounds by amplifying or reducing specific sound components based on vibration intensity, thus overcoming computational burdens and noise interference.

JP7738968B2Active Publication Date: 2025-09-16KDDI CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021192153
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-09-16
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

Existing technologies face challenges in generating audio signals for movements like rustling clothes or sitting sounds in real-time with low computational burden, as they often capture unwanted impulse sounds and require extensive training data and complex processing.

Method used

An audio signal adjustment program and device that uses vibration signals to identify and amplify or reduce specific sound components based on predetermined actions, employing adjustment operators that modify sound signals based on vibration intensity ranges.

Benefits of technology

Enables real-time generation of targeted audio signals with reduced computational burden, effectively separating desired movement sounds from unwanted noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007738968000001
    Figure 0007738968000001
  • Figure 0007738968000002
    Figure 0007738968000002
  • Figure 0007738968000003
    Figure 0007738968000003
Patent Text Reader

Abstract

To provide an acoustic signal adjustment program capable of generating a signal related to acquisition target sound by calculation processing less in burden.SOLUTION: The present program is the one for identifying or classifying an acquisition target acoustic signal part related to sound generated with a predetermined behavior from an original acoustic signal including the sound generated with a behavior. The program allows a computer to function as adjustment operator determination means and target output signal generation means. The adjustment operator determination means determines an adjustment operator, which amplifies or relatively enlarges a part of an original acoustic signal corresponding to a vibration signal part being the part in an acquired vibration signal including a vibration generated together with sound with the behavior and also being the vibration signal part related to the vibration generated together with the sound on the acquisition target acoustic signal part, based on the vibration signal. The target output signal generation means allows the adjustment operator to work on the original acoustic signal, so as to generate a target output signal where the acquisition target acoustic signal part is amplified / relatively-enlarged or the acoustic signal part other than the acquisition target acoustic signal part is reduced / cut.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for adjusting an acoustic signal to a desired signal. [Background technology]

[0002] In recent years, with the spread of web conferencing applications, remote session applications, and the like, services that use microphones and speakers to provide a suitable or desired acoustic environment in accordance with the video being provided have become widely used. High-presence remote conference systems and online concerts are also being actively utilized. Furthermore, in various fields, including the film industry, video content that creates a desired acoustic space in accordance with the video is being created and provided.

[0003] As a technique for adding sound to video, including sounds resulting from movements and the surrounding environment, so-called Foley sound, for example, Non-Patent Document 1 discloses a method for analyzing a recorded waveform and generating a wide variety of sound effects from a single recording, without having to search for and record the sound corresponding to the video each time.Non-Patent Document 2 also explains research trends in techniques for analyzing various environmental sounds, not limited to recorded voices and musical sounds.

[0004] Several methods have been proposed for analyzing recorded data that utilize simultaneously measured acceleration data. For example, Patent Document 1 discloses an integrated sensor that combines a microphone and a vibration acceleration pickup, which can simultaneously measure sound and vibration acceleration at a measurement point, although it is not a technology related to adding sound effects to video. This sensor is said to make it possible to accurately grasp the operating status of equipment in operation from sound and acceleration vibration.

[0005] Patent Document 2 also discloses shoes equipped with a microphone and an acceleration sensor, which generate acoustic features for sound signals input from the microphone for a period when the acceleration value detected by the acceleration sensor is equal to or greater than a predetermined threshold, store the acoustic features in a storage device, and transmit an identifier for identifying the shoes (the shoes) and the acoustic features stored in the storage device to an external device. The external device to which the acoustic features are transmitted extracts acoustic features related to the sound of the sole of the shoe rubbing against the ground from the acoustic features obtained from the shoes, executes processing related to a predetermined service for the extracted acoustic features, and causes the processing results to be displayed on an information processing device that is the source of the display request.

[0006] Furthermore, Patent Document 3 discloses a sound detection device that detects sound radiated from a specific sound radiating surface of a vibrating object. This device is equipped with a vibration accelerometer that detects the vibration acceleration of the specific sound radiating surface, a vibration acceleration-vibration velocity converter that converts the vibration acceleration into vibration velocity, a microphone that detects sound pressure near the specific sound radiating surface, and a cross spectrum calculation unit that calculates the cross spectrum of the vibration velocity and sound pressure, and detects only sound radiated from the specific sound radiating surface. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Japanese Patent Application Publication No. 63-108235 [Patent Document 2] Japanese Patent Publication No. 2021-094373 [Patent Document 3] Japanese Patent Application Laid-Open No. 2003-222553 [Non-patent literature]

[0008] [Non-Patent Document 1] Hikaru Hatazawa, Yoshinori Dobashi, Tsuyoshi Yamamoto, "Generating Variations of Sound Effects Using Recorded Waveforms," ​​Research Report Graphics and CAD (CG) 2011-CG-143(3), pp.1-6, 2011 [Non-patent document 2] Keisuke Imoto, "Research Trends in Environmental Sound Analysis," Journal of the Acoustical Society of Japan, Vol. 75, No. 9, pp. 512-518, 2019 Summary of the Invention [Problem to be solved by the invention]

[0009] For example, when generating sounds of movements such as the rustling of clothes or the sound of someone sitting down to be added to a video image, it has been extremely difficult in conventional technologies, including the technologies described in Non-Patent Documents 1 and 2 and Patent Documents 1 to 3, to generate an audio signal containing only the desired sounds of movements using low-burden processing that allows for real-time processing, for example.

[0010] In fact, when capturing the sound of clothes rustling, for example, if a person performs some action, such as standing or sitting, near the microphone, the microphone will also capture the larger impulse sound associated with that action. Therefore, it is absolutely necessary to process the desired behavioral sound from the acquired acoustic data. Currently, research into environmental sound recognition technology using machine learning to detect and extract specific environmental sounds is being actively conducted. However, when using machine learning in this way, the burden of model construction using a large amount of training data and the environmental sound extraction process using this model are usually very large.

[0011] Furthermore, without applying machine learning, it is generally necessary to analyze time-related information such as frequency and power for sound data over a certain time period. However, this process is usually very burdensome, and it is currently not possible to extract desired gesture sounds in real time.

[0012] The inventors of the present application have considered that it may be possible to apply the acceleration vibration data described in the above-mentioned Patent Documents 1 to 3 to the analysis of such acoustic data. That is, they have considered that it may be possible to detect and extract, for example, a desired gesture sound by utilizing the acceleration vibration data acquired together with the acoustic data.

[0013] In this regard, for example, Patent Document 1 presents a sensor structure capable of simultaneously acquiring acoustic data and acceleration vibration data, but does not propose a specific method for analyzing acoustic data. Furthermore, for example, the technology described in Patent Document 2 only uses the output of an acceleration sensor to acquire acoustic features related to the sound of a shoe sole rubbing against the ground, and does not anticipate detecting or extracting, for example, the sound of a desired movement. Furthermore, to begin with, it is only possible to identify acoustic data accompanied by vibrations greater than a predetermined level.

[0014] Furthermore, the technology described in Patent Document 3 converts vibration acceleration obtained by a vibration accelerometer into vibration velocity, calculates the cross spectrum between this vibration velocity and the sound pressure obtained by a microphone, and detects only sounds emitted from a specific sound emitting surface of a vibrating object that is vibrating in a specific manner. In other words, because sounds to be suppressed other than the sound to be detected are processed based on the cross spectrum, only sounds in a specific frequency band are detected, making it very difficult to remove or reduce sudden movement sounds such as footsteps, which therefore span a wide frequency band.

[0015] Therefore, an object of the present invention is to provide an audio signal adjustment program, an audio signal adjustment device, and an audio signal adjustment method that can generate a signal related to a target sound by performing calculation processing with a smaller burden. [Means for solving the problem]

[0016] According to the present invention, from an original sound signal containing information on sounds generated in association with a movement, One of a plurality of predetermined actions that have been set in advance An audio signal adjustment program that causes a computer to function to identify or sort an audio signal portion to be acquired that is related to a sound generated in association with a predetermined action, This acoustic signal conditioning program The vibration signal portion in the acquired vibration signal includes information on the vibration generated together with the sound in accordance with the movement, and an adjustment operator that amplifies or relatively increases a portion of the original acoustic signal corresponding to a vibration signal portion related to the vibration generated together with the sound related to the target acoustic signal portion is used to obtain the vibration signal obtained from the acquired vibration signal. A predetermined action is an adjustment operator determination means for determining an adjustment operator based on information relating to whether the adjustment operator falls within a predetermined range; a target output signal generating means for applying the adjustment operator to the original acoustic signal to generate a target output signal in which the target acoustic signal portion is amplified or relatively increased, or in which the acoustic signal portion other than the target acoustic signal portion is reduced or eliminated; to make the computer function 、 A predetermined range of the vibration intensity is set in advance for each of the plurality of predetermined actions. An audio signal conditioning program is provided.

[0017] In the sound signal adjustment program according to the present invention, the adjustment operator determining means determines whether the amplitude of the vibration signal or a representative value of the amplitude is be To the specified operation On the other hand It is also preferable to determine the adjustment operator that amplifies or relatively increases the part of the original acoustic signal when it falls within a preset amplitude range as the part of the acoustic signal to be acquired.

[0018] In addition, in the audio signal adjustment program according to the present invention, ,eye The target output signal generating means is Multiple of the above Predetermined action Each of It is also preferable to generate a plurality of target output signals corresponding to the predetermined motion, the target output signals including information on the sound generated together with the vibrations whose amplitude or a representative value of the amplitude is within a predetermined amplitude range for the predetermined motion.

[0019] Furthermore, as one embodiment of the acoustic signal adjustment program according to the present invention, the adjustment operator determination means (a) determines, as the first adjustment operator, an operator that multiplies the amplitude of the vibration signal or the representative value of the amplitude by an adjustment coefficient that is a monotonically decreasing function when the acoustic signal portion to be acquired is an acoustic signal portion related to a predetermined action accompanied by smaller vibrations, and the amplitude or the representative value of the amplitude of the corresponding vibration signal is within an amplitude range below a predetermined threshold, and / or (b) determines, as the second adjustment operator, an operator that multiplies the amplitude of the vibration signal or the representative value of the amplitude by an adjustment coefficient that is a monotonically increasing function when the acoustic signal portion to be acquired is an acoustic signal portion related to a predetermined action accompanied by larger vibrations, and the amplitude or the representative value of the amplitude is within an amplitude range equal to or greater than a predetermined threshold, The desired output signal generating means (a) generates a first desired output signal by applying a first adjustment operator to the original sound signal, and / or (b) generates a second desired output signal by applying a second adjustment operator to the original sound signal. It is also preferable to

[0020] In another embodiment of the acoustic signal adjustment program according to the present invention, the adjustment operator determination means (a) determines, as the first adjustment operator, an operator that performs a Fourier transform on the original acoustic signal, multiplies the result of the Fourier transform by an adjustment coefficient that is a monotonically decreasing function for the frequency components of the vibration signal, and further performs an inverse Fourier transform on the result of the multiplication, when the acoustic signal portion to be acquired is an acoustic signal portion related to a predetermined action accompanied by smaller vibrations, and the amplitude or a representative value of the amplitude of the corresponding vibration signal is within an amplitude range less than a predetermined threshold; and / or (b) determines, as the second adjustment operator, an operator that performs a Fourier transform on the original acoustic signal, multiplies the result of the Fourier transform by an adjustment coefficient that is a monotonically increasing function for the frequency components of the vibration signal, and further performs an inverse Fourier transform on the result of the multiplication, when the acoustic signal portion to be acquired is an acoustic signal portion related to a predetermined action accompanied by larger vibrations, and the amplitude or a representative value of the amplitude is within an amplitude range equal to or greater than a predetermined threshold; The desired output signal generating means (a) generates a first desired output signal by applying a first adjustment operator to the original sound signal, and / or (b) generates a second desired output signal by applying a second adjustment operator to the original sound signal. It is also preferable to

[0021] Furthermore, in the cases of "(a) and (b)" of the above-mentioned embodiments, it is also preferable that the target output signal generating means generates the first target output signal and the second target output signal as signals to be output to a first speaker and a second speaker having different specifications or positions, respectively, or as signals to be output to a certain speaker and a vibration generating device.

[0022] It is also preferable that the vibration signal according to the present invention is an acceleration vibration signal having an acceleration amplitude generated by an acceleration sensor installed at a position where vibrations occurring due to the movement can be detected.

[0023] According to the present invention, from an original sound signal containing information on sounds generated in association with a movement, One of a plurality of predetermined actions that have been set in advance An audio signal adjustment device for identifying or sorting audio signals to be acquired related to sounds generated in association with a predetermined action, This acoustic signal conditioning device The vibration signal portion in the acquired vibration signal includes information on the vibration generated together with the sound in accordance with the movement, and an adjustment operator that amplifies or relatively increases a portion of the original acoustic signal corresponding to a vibration signal portion related to the vibration generated together with the sound related to the target acoustic signal portion is used to obtain the vibration signal obtained from the acquired vibration signal. A predetermined action is an adjustment operator determination means for determining an adjustment operator based on information relating to whether the adjustment operator falls within a predetermined range; a target output signal generating means for applying the adjustment operator to the original acoustic signal to generate a target output signal in which the target acoustic signal portion is amplified or relatively increased, or in which the acoustic signal portion other than the target acoustic signal portion is reduced or eliminated; With death, A predetermined range of the vibration intensity is set in advance for each of the plurality of predetermined actions. An audio signal conditioning device is provided.

[0024] According to the present invention, the following is further obtained from an original sound signal containing information on sounds generated in association with a movement: One of a plurality of predetermined actions that have been set in advance 1. A computer-implemented method for adjusting an audio signal, the method comprising: identifying or classifying audio signals to be acquired that relate to sounds generated by a predetermined action; This acoustic signal conditioning method includes: The vibration signal portion in the acquired vibration signal includes information on the vibration generated together with the sound in accordance with the movement, and an adjustment operator that amplifies or relatively increases a portion of the original acoustic signal corresponding to a vibration signal portion related to the vibration generated together with the sound related to the target acoustic signal portion is used to obtain the vibration signal obtained from the acquired vibration signal. A predetermined action is determining whether the predetermined range is met based on the information; applying the adjustment operator to the original acoustic signal to generate a target output signal in which the target acoustic signal portion is amplified or relatively increased, or in which the acoustic signal portion other than the target acoustic signal portion is reduced or eliminated; With death, A predetermined range of the vibration intensity is set in advance for each of the plurality of predetermined actions. A method for conditioning an audio signal is provided. [Effects of the Invention]

[0025] According to the sound signal adjustment program, device and method of the present invention, a signal related to a target sound can be generated by a less burdensome calculation process. [Brief explanation of the drawings]

[0026] [Figure 1] 1 is a functional block diagram showing a functional configuration of an embodiment of an acoustic signal adjustment device according to the present invention; [Figure 2] FIG. 10 is a schematic diagram illustrating an embodiment in which two adjustment operators corresponding to two acquisition target acoustic signal portions are determined. [Figure 3] 1 is a schematic diagram illustrating an embodiment for determining two target output signals corresponding to two acoustic signal portions of acquisition targets; FIG. DETAILED DESCRIPTION OF THE INVENTION

[0027] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0028] [Acoustic signal conditioning device] FIG. 1 is a functional block diagram showing the functional configuration of an embodiment of an acoustic signal adjustment device according to the present invention.

[0029] The acoustic signal adjustment device 1 according to one embodiment of the present invention shown in FIG. 1 is a device that can identify or sort a "target acoustic signal portion" related to a sound generated in association with a predetermined action from an "original acoustic signal" that contains information about the sound generated in association with the action, by using a "vibration signal" that also contains information about the vibrations generated in association with the action.

[0030] In this embodiment, the acoustic signal adjustment device 1 acquires an "original acoustic signal" and a "vibration signal" as time-series output signals via a communication network from a microphone 2 installed at a predetermined position in the recording room (sound space), an acceleration sensor 3 installed on the floor (at the same position as the microphone 2), and a microphone 4 with an acceleration sensor attached to a chair or person in the recording room.

[0031] Here, the "original acoustic signal" may be a signal obtained by combining all the acoustic signals output from the microphone 2 and the microphone 4 with an acceleration sensor installed in the recording room, but it is also preferable to use the acoustic signal from the acoustic signal output source selected according to the "audio signal portion to be acquired" as the "original acoustic signal."

[0032] For example, if the "audio signal portion to be acquired" is "(a signal portion related to) the sound of clothes rustling," the audio signal from a microphone 4 with an acceleration sensor attached to the person from whom the sound of clothes rustling is to be acquired can be the "original audio signal." Also, if the sound is "the sound of sitting down on a chair," it is preferable to use an audio signal from a microphone 4 with an acceleration sensor attached to the chair (for example, installed inside the chair) as the "original audio signal." Furthermore, if the sound is "footsteps," the audio signal from a microphone 2 installed near the floor can be used as the "original audio signal."

[0033] On the other hand, the "vibration signal" can be a vibration signal from a vibration signal output source provided at a position where vibrations generated by actions related to the "target sound signal portion" can be detected. For example, if the "target sound signal portion" is "(a signal portion related to) the sound of clothes rustling," the "vibration signal" can be a vibration signal from a microphone 4 with an acceleration sensor attached to a person from whom the sound of clothes rustling is to be acquired. Also, if the sound is "the sound of sitting down on a chair," it is preferable to use a vibration signal from a microphone 4 with an acceleration sensor attached to the chair. Furthermore, if the sound is "footsteps," the "vibration signal" can be a vibration signal from an acceleration sensor 3 installed on the floor.

[0034] In any case, in order to simultaneously detect the sound and vibration generated by the action related to the "target acoustic signal portion," it is preferable that the microphone providing the "original acoustic signal" and the acceleration sensor providing the "vibration signal" are located at the same position or in nearby positions.

[0035] Upon receiving the "original acoustic signal" and the "vibration signal" as described above, the acoustic signal adjustment device 1 generates a "target output signal" as a result of identifying or sorting the "target acoustic signal portion to be acquired." (A) an adjustment operator determination unit 111 that determines, based on the acquired "vibration signal," an "adjustment operator" that amplifies or relatively increases a vibration signal portion in the acquired "vibration signal" that corresponds to a vibration signal portion in the "original acoustic signal" related to vibrations that occurred together with a sound related to the "acoustic signal portion to be acquired"; (B) a target output signal generating unit 112 that applies an “adjustment operator” to the “original acoustic signal” to generate a “target output signal” in which the “target acoustic signal portion” is amplified or relatively increased, or in which acoustic signal portions other than the “target acoustic signal portion” are reduced or eliminated; It is characterized by having:

[0036] Here, a predetermined action related to the "target audio signal portion" typically generates not only the action sound of the action but also a vibration specific to the action. For example, if the predetermined action is "a movement that causes clothes to rustle," a smaller vibration having an (average) amplitude below a predetermined threshold is generated along with the "clothes rustling" action sound, while if the action is sitting down on a chair, a larger vibration having an (average) amplitude above a predetermined threshold is generated along with the "sound of sitting down on the chair" action sound.

[0037] The acoustic signal adjustment device 1 focuses on vibrations specific to such "audio signal portions to be acquired," and regards the acoustic signal portions in the "original acoustic signal" corresponding to the vibrations as the "audio signal portions to be acquired," thereby generating a "target output signal." Specifically, an "adjustment operator" that amplifies or relatively increases the acoustic signal portions corresponding to the vibrations is determined based on the "vibration signal," and the "target output signal" is generated from the "original acoustic signal" using this.

[0038] In this regard, if an output signal related to the target sound is generated based solely on the loudness of the target sound (acoustic signal portion), for example, if a corresponding predetermined motion happens to occur near the sound pickup microphone, the picked-up sound may be louder than it actually is, and the picked-up sound may be determined to be not the target sound. The predetermined motion related to the target sound (acoustic signal portion) typically involves almost no vibration or generates vibrations greater than a predetermined level. Therefore, while the acoustic signal adjustment device 1 can determine the magnitude (presence or absence) of such vibrations and generate an appropriate target output signal, judging based solely on the loudness of the sound may result in an incorrect target output signal being generated.

[0039] Furthermore, as will be described later, the above-mentioned "adjustment operator" can be, for example, an operator that multiplies (the original sound signal) by an adjustment coefficient that is a monotonically decreasing or monotonically increasing function of the (average) amplitude of the "vibration signal," and the calculation process using this is a process with an extremely low burden compared to, for example, sound signal extraction processing using known DNNs (Deep Neural Networks). In other words, with the sound signal adjustment device 1, a signal related to the sound to be acquired (target output signal) can be generated with a calculation process with a lower burden.

[0040] Furthermore, in this embodiment, the acoustic signal adjustment device 1 uses a less burdensome calculation process, and is therefore capable of generating and outputting a "target output signal" in real time, for example, by receiving an "original acoustic signal" and a "vibration signal" and within a predetermined small delay time.

[0041] 1, the acoustic signal adjustment device 1 may transmit in real time a "first target output signal" and a "second target output signal" containing the "sound of clothes rustling" and the "sound of sitting down on a chair" (the sounds to be acquired) collected in a sound collection room to a speaker 5 and a haptic device (vibration generating device) 6 (in the listening room), respectively. This allows the viewer to enjoy the sound of clothes rustling in real time as sound along with an image of the sound collection room (for example, an image from a camera installed in the sound collection room), and also to enjoy the action of sitting down in real time as vibrations (of the chair on which the viewer is sitting).

[0042] Incidentally, the recording room and listening room in Figure 1 may be places connected by communication, for example, in a remote conference system or web conferencing system, and may also represent the stage and audience locations in an online theater or concert, respectively.

[0043] Furthermore, the "target output signal," which is the output (result) of the audio signal adjustment device 1, can be used in various forms other than for the real-time production of sounds and vibrations as described above. For example, it can be used when a robot with hearing detects the source of a sound (the cause of the sound) in real time from collected acoustic information. Furthermore, although real-time performance is not required, it may be used as Foley sounds added to video content such as movies and television dramas. In this case, it is easy to sort the sounds to be used, making it easier to add sounds to the video.

[0044] In this embodiment, the "vibration signal" is a vibration signal (acceleration vibration signal having acceleration amplitude (or velocity amplitude)) generated by an acceleration sensor (3, 4) installed at a position where it can detect vibrations generated in conjunction with a predetermined operation related to the "acoustic signal portion to be acquired," but of course it is not limited to this. In other words, as a sensor for generating the "vibration signal," various types of sensors can be used, such as a vibration sensor such as a piezoelectric sensor, as long as they are capable of outputting a vibration waveform signal having a displacement amplitude, velocity amplitude, or acceleration amplitude in the vibration.

[0045] [Device configuration, acoustic signal adjustment program and method] 1, the acoustic signal conditioning device 1 of this embodiment includes a communication interface 101, an acoustic signal / vibration signal storage unit 102, and a processor / memory. The processor / memory stores an embodiment of an acoustic signal conditioning program according to the present invention, has computer functions, and executes the acoustic signal conditioning program to perform acoustic signal conditioning processing.

[0046] Furthermore, for this reason, the acoustic signal adjustment device 1 may be a device dedicated to acoustic signal adjustment processing, or it may be a general-purpose cloud server or non-cloud server equipped with the acoustic signal adjustment program according to the present invention, or it may even be a personal computer (PC), a notebook or tablet computer, a smartphone, etc.

[0047] The processor memory also has, as functional components, an adjustment operator determination unit 111 including a first operator determination unit 111a and a second operator determination unit 111b, a target output signal generation unit 112 including a first target signal generation unit 112a and a second target signal generation unit 112b, and a communication control unit 121. These functional components can be considered to be functions of an embodiment of an acoustic signal adjustment program according to the present invention stored in the processor memory, and the processing flow shown by connecting the functional components of the acoustic signal adjustment device 1 with arrows in the functional block diagram of Figure 1 can also be understood as an embodiment of an acoustic signal adjustment method according to the present invention.

[0048] 1 , the acoustic signal / vibration signal storage unit 102 of this embodiment acquires and stores original acoustic signals and vibration signals from the microphone 2, acceleration sensor 3, and microphone 4 with acceleration sensor provided in the recording room (sound space) via a communication network, a communication interface 101, and a communication control unit 121. Here, in this embodiment, each original acoustic signal and each vibration signal is associated with an identifier (ID) of the signal output source (microphone or acceleration sensor), and the acoustic signal / vibration signal storage unit 102 associates installation / mounting position information with the original acoustic signal or vibration signal for each signal output source (ID) and manages them.

[0049] In this embodiment, the original sound signal is time-series data of amplitude corresponding to sound intensity, i.e., amplitude x(t) as a function of time, and the vibration signal is time-series data of acceleration amplitude corresponding to the second derivative of the displacement amplitude, i.e., (acceleration) amplitude y(t) as a function of time. The amplitude y(t) of the vibration signal can also be velocity amplitude or displacement amplitude. Furthermore, the amplitude y(t) of the vibration signal can be the sensor output as is, but it is also preferable to use a representative value of the output amplitude, for example, a moving average value of the absolute amplitude value over a predetermined time interval. This makes it possible to more appropriately determine the adjustment operator described below. In the embodiment described below, the amplitude y(t) of the vibration signal can be a representative value of the amplitude (for example, a moving average value).

[0050] <Adjustment operator determination process> Also in the functional block diagram of Figure 1, the adjustment operator determination unit 111 of this embodiment holds a preset amplitude range (e.g., a range of amplitudes below a predetermined threshold) for a predetermined action that produces a sound (e.g., a "rustling sound") related to the "audio signal portion to be acquired" set by user input (from a user interface not shown).

[0051] Here, in this embodiment, when the amplitude of the vibration signal falls within a preset amplitude range for this predetermined motion, the adjustment operator determination unit 111 determines, based on the corresponding vibration signal, an adjustment operator that amplifies or relatively increases the portion of the original sound signal at that time as the "sound signal portion to be acquired." Note that, in this embodiment, this corresponding vibration signal is a vibration signal linked to the ID of a vibration signal output source (the acceleration sensor 3 or the microphone 4 with acceleration sensor in FIG. 1) that can detect vibration corresponding to the sound related to the set "sound signal portion to be acquired."

[0052] For example, if the sound related to the set "acoustic signal portion to be acquired" is "clothes rustling" and this "clothes rustling" is usually accompanied by smaller vibrations, and an amplitude range including amplitudes less than a predetermined threshold th1 is set, the adjustment operator is expressed by the following formula: (1) x'(t) = x(t) × c1(y(t) <th1) =x(t) / c2(y(t)≧th1) In this case, it can be the portion (×c1 or / c2) acting on x(t). Here, x(t) is the amplitude of the acoustic signal, x'(t) is the (amplitude of) the acoustic signal after adjustment as the target output signal, and y(t) is the amplitude of the vibration signal. Also, c1 is an amplification factor of 1 or more, and c2 is a suppression factor of a value greater than 1.

[0053] In addition, in consideration of the fact that the sound related to the set "acoustic signal portion to be acquired" is the "sound of sitting on a chair" and that this "sound of sitting on a chair" is usually accompanied by a larger vibration, when an amplitude range including an amplitude equal to or greater than a predetermined threshold th2 is set, the adjustment operator is expressed by the following formula: (2) x'(t)=x(t) / c2(y(t) <th2) =x(t)×c1(y(t)≧th2) This can be the part acting on x(t) (×c1 or / c2).

[0054] Furthermore, in consideration of the fact that the sound related to the set "acoustic signal portion to be acquired" is a "hand clapping sound" and that this "hand clapping sound" is usually accompanied by a medium-sized vibration, when an amplitude range including an amplitude equal to or greater than a predetermined threshold th3 and less than a predetermined threshold th4 (>th3) is set, the adjustment operator is expressed by the following formula: (3) x'(t)=x(t) / c2(y(t) <th3) =x(t)×c1(th3≦y(t) <th4) =x(t) / c2(y(t)≧th4) This can be the part acting on x(t) (×c1 or / c2).

[0055] Here, when a plurality of sounds (acoustic signal portions) for acquisition purposes are set, it is also preferable to determine a plurality of adjustment operators corresponding to each of them. For example, when "clothing rubbing sound", "sound of clapping hands", and "sound of sitting on a chair" are set, three adjustment operators corresponding to each of the above formulas (1) to (3) may be determined. Thereby, it is also possible to individually generate a plurality of target output signals including each of the set sounds (acoustic signal portions) for acquisition purposes.

[0056] Also, when a plurality of sounds (acoustic signal portions) for acquisition purposes are set, it is also possible to determine an adjustment operator for generating one target output signal including those sounds (acoustic signal portions). For example, when "clothing rubbing sound", "sound of clapping hands", and "sound of sitting on a chair" are set, assuming that the vibration ranges related to each sound do not overlap, that is, th1 < th3 < th4 < th2, the following formula (4) x'(t)=x(t)×c1(y(t)<th1) x(t) / c2(th1≦y(t)<th3) =x(t)×c1(th3≦y(t)<th4) =x(t) / c2(th4≦y(t)<th2) =x(t)×c1(y(t)≧th2) The portions acting on x(t) (×c1 or / c2) in may be used as the adjustment operator.

[0057] The adjustment operator is not limited to the above. For example, various types of operators can be used as the adjustment operator as long as they multiply the amplitude y(t) by a larger coefficient when the amplitude y(t) is within a set amplitude range compared to when the amplitude y(t) is not within the set amplitude range. It is also possible to determine an adjustment operator in which the amplification factor c1 (≧1) and the suppression factor c2 (>1) in the above formulas (1) to (4) are functions of y(t) (for example, in the form of formulas (6), (7), (9), and (10) described below) rather than set fixed values. In any case, such adjustment operators make it possible to amplify or relatively increase the amplitude of the original sound signal portion corresponding to the vibration signal portion associated with the vibration generated together with the sound associated with the set "target sound signal portion" (for example, the sound of rustling clothes, clapping hands, or the sound of sitting down on a chair).

[0058] Next, a typical embodiment of an adjustment operator determination process will be described below in which two sounds (acoustic signal portions) to be acquired are set and two adjustment operators (a first adjustment operator and a second adjustment operator) are determined to generate two target output signals corresponding to the respective sounds.

[0059] In this embodiment, when the "acquisition target sound signal portion" set (by the user) is a sound signal portion related to a predetermined action accompanied by smaller vibrations (for example, an action causing clothes to rub against each other), and the amplitude of the corresponding vibration signal at that time is included in an amplitude range less than a predetermined threshold value Th1, the first operator determination unit 111a of the adjustment operator determination unit 111 determines an operator that multiplies an adjustment coefficient C1 that is a monotonically decreasing function of the amplitude y(t) of the vibration signal, i.e., (5) x'(t)=x(t)×C1 The operator (the ×C1 part) that acts on x(t) as follows is determined as the first adjustment operator.

[0060] The adjustment coefficient C1 can be expressed as follows in a simple form: (6) C1=y0 / y(t), or (7) C1=(y0 / y(t))N (where N is a real number greater than 1, e.g., N=2) Here, y0 is a reference amplitude, specifically, a reference state is defined as a state in which no acceleration (external force) other than gravity is acting on the vibration signal output source (acceleration sensor 3 or microphone 4 with acceleration sensor in FIG. 1 ), and y0 is a representative value of the (acceleration) amplitude of the vibration signal (from the vibration signal output source) in this reference state, for example, a moving average value of the (acceleration) amplitude over a predetermined time interval (for example, 10.7 msec≒512 samples / 48 kHz). In this embodiment, the adjustment operator determination unit 111 has a function of calculating such a reference amplitude y0. Incidentally, from the definition of y0, y0 / y(t) is a value less than or equal to 1 (y0 / y(t)≦1).

[0061] Of course, the adjustment coefficient C1 is not limited to the above formula (6) or (7), and various forms can be adopted as long as it is a monotonically decreasing function of the amplitude y(t). In any case, the first adjustment operator as described above makes it possible to suppress, for example, an audio signal portion related to a movement sound that has a small volume but a large vibration, and to identify or sort the target sound (audio signal portion) that has a smaller vibration.

[0062] Furthermore, when the "acquisition target sound signal portion" set (by the user) is a sound signal portion related to a predetermined action (for example, the action of sitting down on a chair) accompanied by larger vibrations, and the amplitude of the corresponding vibration signal at that time is included in an amplitude range equal to or greater than a predetermined threshold value Th2 (≧Th1), the second operator determination unit 111b of the adjustment operator determination unit 111 determines an operator that multiplies the amplitude y(t) of the vibration signal by an adjustment coefficient C2 that is a monotonically increasing function, i.e., (8) x'(t)=x(t)×C2 The operator (the ×C2 part) that acts on x(t) as follows is determined as the second adjustment operator.

[0063] The adjustment coefficient C2 can be expressed as follows in a simple form: (9) C2=y(t) / y0, or (10) C2=(y(t) / y0) N (where N is a real number greater than 1, e.g., N=2) It is possible to obtain the following. Of course, the adjustment coefficient C2 is not limited to the above formula (9) or (10), and various forms can be adopted as long as it is a monotonically increasing function of the amplitude y(t). In any case, by using the second adjustment operator as described above, it is possible to suppress the sound signal portion related to the movement sound, which is loud but hardly accompanied by vibration, and to identify or sort the target sound (sound signal portion) that is accompanied by greater vibration.

[0064] FIG. 2 is a schematic diagram for explaining the embodiment described above in which two adjustment operators (first and second adjustment operators) corresponding to two sounds (acoustic signal portions) to be acquired are determined.

[0065] 2, the acquired original sound signal contains two "target sound signal portions," specifically, sound signal portions relating to the "sound of clothes rustling" and the "sound of sitting down on a chair." In addition, the vibration signal (acceleration vibration signal) measured and acquired at the same time also contains vibration signal portions relating to the "sound of clothes rustling" and the "sound of sitting down on a chair."

[0066] Of these, the amplitude (acceleration amplitude or its representative value) y(t) of the vibration signal portion relating to the "sound of clothes rustling" falls within a first amplitude range below a predetermined threshold th, while the amplitude (acceleration amplitude or its representative value) y(t) of the vibration signal portion relating to the "sound of sitting down on a chair" falls within a second amplitude range above the predetermined threshold th. Here, in this example, the above-mentioned predetermined thresholds th1 and th2 are equal to each other, and the predetermined threshold th is set to the same value as these thresholds (th=th1=th2).

[0067] Here, when the "clothes rustling sound" is the sound to be acquired (acoustic signal part), the amplitude y(t) of the corresponding vibration signal part falls within the first amplitude range, and therefore, for example, the following equation is used: (5') x'(t)=x(t)×y0 / y(t) The operator (the part of ×y0 / y(t)) that acts on x(t) as follows can be determined as the first adjustment operator.

[0068] On the other hand, when the "sound of sitting on a chair" is the target sound (acoustic signal part), the amplitude y(t) of the corresponding vibration signal part is within the second amplitude range, and therefore, for example, the following equation is used: (8') x'(t)=x(t)×y(t) / y0 The operator that acts on x(t) (the part ×y(t) / y0) can be determined as the second adjustment operator.

[0069] 1, next, an embodiment will be described in which the first operator determination unit 111a and the second operator determination unit 111b determine an adjustment operator including a Fourier transform, rather than determining a multiplication operator as described above as the adjustment operator. Incidentally, such an embodiment that performs a Fourier transform is adopted when it is necessary to process vibration signals in a predetermined time interval as needed, and a predetermined delay in the output of the final target output signal is acceptable.

[0070] In this embodiment, when the "acquisition target audio signal portion" set (by the user) is an audio signal portion related to a predetermined action accompanied by smaller vibrations (for example, an action causing clothes to rub against each other) and the amplitude of the corresponding vibration signal at that time is within an amplitude range below a predetermined threshold Th1, the first operator determination unit 111a performs a Fourier transform on the original audio signal, multiplies the result of this Fourier transform by an adjustment coefficient CF1 that is a monotonically decreasing function for the frequency components of the vibration signal, and determines the operator that performs an inverse Fourier transform on the result of this multiplication as the first adjustment operator Φ1. Specifically, the first adjustment operator Φ1 is expressed by the following equation: (11) Φ1:x(t)-<Fourier transform>→X(p) X(p) - <adjustment coefficient multiplication> → X'(p) = X(p) × CF1 X'(p)-<inverse Fourier transform> → x'(t) Here, p is a discrete frequency, and X(p) is a frequency component of the acoustic signal (its amplitude x(t)).

[0071] Of these, the adjustment coefficient CF1 can be expressed as follows in a simple form: (12) CF1=Y0 / Y(p), or (13) CF1=(Y0 / Y(p)) N (where N is a real number greater than 1, e.g., N=2) Here, Y0 (= Y0(p)) is the reference frequency component, which is the frequency component of the reference amplitude y0(t) obtained by applying a Fourier transform to the reference amplitude y0 (= y0(t)). Also, Y(p) is the frequency component of the vibration signal (its amplitude y(t)). Incidentally, based on this definition, Y0 / Y(p) is a value less than or equal to 1 (Y0 / Y(p)≦1).

[0072] Of course, the adjustment coefficient CF1 is not limited to the above equation (12) or (13), and various forms can be adopted as long as they are a monotonically decreasing function of the frequency component Y(P).

[0073] Furthermore, when the "acquisition target sound signal portion" set (by the user) is a sound signal portion related to a predetermined action (for example, the action of sitting down on a chair) accompanied by larger vibrations, and the amplitude of the corresponding vibration signal at that time is within an amplitude range equal to or greater than a predetermined threshold value Th2 (≧Th1), the second operator determination unit 111b of the adjustment operator determination unit 111 performs a Fourier transform on the original sound signal, multiplies the result of this Fourier transform by an adjustment coefficient CF2 which is a monotonically increasing function of the frequency components of the vibration signal, and determines the operator that performs an inverse Fourier transform on the result of this multiplication as the second adjustment operator Φ2. Specifically, the second adjustment operator Φ2 is expressed by the following equation: (14) Φ2:x(t)-<Fourier transform>→X(p) X(p) - <adjustment coefficient multiplication> → X'(p) = X(p) × CF2 X'(p)-<inverse Fourier transform> → x'(t) It can be an operator expressed as:

[0074] Of these, the adjustment coefficient CF2 can be expressed as follows in a simple form: (15) CF1=Y(p) / Y0, or (16) CF1=(Y(p) / Y0) N (where N is a real number greater than 1, e.g., N=2) Of course, the adjustment coefficient CF2 is not limited to the above equation (15) or (16), and various forms can be adopted as long as they are a monotonically increasing function of the frequency component Y(p).

[0075] The above describes a typical embodiment in which the adjustment operator determination unit 111 determines the first adjustment operator and the second adjustment operator for each of two target sounds (e.g., the "sound of clothes rustling" and the "sound of sitting down on a chair"). In this embodiment, the target output signal generation unit 112 then uses these two adjustment operators to generate the first target output signal and the second target output signal, respectively.

[0076] However, it is also possible that the adjustment operator determination unit 111 determines only one of the first adjustment operator and the second adjustment operator, and in response to this, the target output signal generation unit 112 generates only one of the first target output signal and the second target output signal. For example, it is also possible to adopt an embodiment in which only the first target output signal whose main component is the "target sound signal portion" relating to "clothes rustling" is generated from the determined first adjustment operator.

[0077] In any case, by using the adjustment operator as described above, it is possible to amplify or relatively increase the volume of the original sound signal portion that corresponds to the vibration signal portion related to vibrations that occur together with sounds related to the set "sound signal portion to be acquired" (for example, "rustling clothes," "clapsing hands," or "sound of sitting down on a chair").

[0078] <Target output signal adjustment processing> Also in the functional block diagram of Figure 1, the target output signal generation unit 112 receives the adjustment operator determined by the adjustment operator determination unit 111 and applies the adjustment operator to the original acoustic signal to generate a target output signal in which the "target acoustic signal portion" has been amplified or made relatively larger, or in which acoustic signal portions other than the "target acoustic signal portion" have been reduced or eliminated.

[0079] Specifically, the target output signal generation unit 112 may receive a corresponding adjustment operator determined according to the "target acoustic signal portion to be acquired" and generate a target output signal having an amplitude x'(t) using, for example, any one of the above equations (1), (2), (3), (4), (5), (8), (11) and (14).

[0080] Furthermore, when a plurality of sounds (sound signal portions) to be acquired are set, the target output signal generating unit 112 (a) receiving the plurality of adjustment operators determined for each sound (sound signal portion) and generating a plurality of target output signals for each sound (sound signal portion), for example, three adjustment operators using the above equations (1) to (3); or (b) It is also possible to receive an adjustment operator that generates a desired output signal that includes those sounds (acoustic signal portions), and generate the desired output signal using, for example, the above equation (4).

[0081] Hereinafter, an exemplary embodiment will be described in which the target output signal generator 112 generates two target output signals corresponding to two "acoustic signal portions to be acquired", respectively.

[0082] Specifically, in this embodiment, the first target signal generator 112a of the target output signal generator 112 receives the first adjustment operator determined by the first operator determiner 111a, and generates a first target output signal corresponding to one of the two "acoustic signal portions to be acquired" using, for example, the above formula (5) or (11). Meanwhile, the second target signal generator 112b of the target output signal generator 112 receives the second adjustment operator determined by the second operator determiner 111b, and generates a second target output signal corresponding to the other of the two "acoustic signal portions to be acquired" using, for example, the above formula (8) or (14).

[0083] FIG. 3 is a schematic diagram for explaining this embodiment in which two target output signals (first target output signal and second target output signal) corresponding to two sounds (acoustic signal portions) to be acquired are determined.

[0084] According to Fig. 3, the acquired original sound signal includes two "acoustic signal portions to be acquired", specifically, audio signal portions relating to the "clothes rustling sound" and the "sound of sitting down on a chair". As explained using Fig. 2, a first adjustment operator (for example, the ×y0 / y(t) part in the above formula (5')) has been determined for the "clothes rustling sound", which is one of the acquisition objectives, and further, a second adjustment operator (for example, the ×y(t) / y0 part in the above formula (8')) has been determined for the "sound of sitting down on a chair", which is the other acquisition objective.

[0085] Here, the first target signal generator 112a applies a first adjustment operator (e.g., ×y0 / y(t)) to the original sound signal to generate a first target output signal. This first target output signal is a signal obtained by extracting only the sound signal portion related to the "clothes rustling sound" that is the target of acquisition from the original sound signal, as also shown in FIG. 3.

[0086] On the other hand, the second target signal generator 112b applies a second adjustment operator (for example, ×y(t) / y0) to the original sound signal to generate a second target output signal. As shown in Fig. 3, this second target output signal is also a signal obtained by extracting only the sound signal portion related to the target sound, "the sound of sitting down on a chair," from the original sound signal.

[0087] In this embodiment, the target output signal generator 112 generates the first target output signal and the second target output signal as signals to be output to a speaker (speaker 5 in the listening room in FIG. 3) and a vibration generating device (haptic device 6 attached to a chair placed in the listening room in FIG. 3), respectively. Specifically, the first target output signal is generated as an acoustic signal (for example, capable of being input to speaker 5), and the second target output signal can also be generated as an acoustic signal (for example, if haptic device 6 is a device that receives an acoustic signal as input). On the other hand, for example, if haptic device 6 receives a signal equivalent to a vibration signal as input, it is also preferable to generate the second target output signal as a signal equivalent to the vibration signal.

[0088] In this embodiment, the generated first target output signal and second target output signal are transmitted from the communication control unit 121 and the communication interface 101 to the speaker 5 and the haptic device 6 in the listening room, respectively, via the communication network.

[0089] This allows the viewer to enjoy, for example, an image from inside a sound recording room (for example, an image from a camera installed inside the sound recording room) on a display (not shown), while (a) The sound of clothes rustling is enjoyed in real time as sound through the speaker 5, and (b) The user can experience the action of sitting down in a chair in real time as vibrations (of the chair they are sitting on) via the haptic device 6. In other words, it is possible to present the viewer with appropriate sounds of movements and physical sensations, creating a more realistic experience.

[0090] The output destinations (transmission destinations) of the first and second target output signals are not limited to those described above. For example, the first and second target output signals may be output (transmitted) to a first speaker and a second speaker, respectively, that have different specifications or locations. For example, the first speaker may be a tweeter for high frequencies (or low volume) installed near the display, and the second speaker may be a woofer for low frequencies (or high volume) installed in a chair. Furthermore, the second target output signal may be output (transmitted) to a haptic actuator that generates a virtual force sensation, together with or instead of a speaker.

[0091] In a further preferred embodiment, the acoustic signal adjustment device 1 may be a communication terminal such as a smartphone or tablet computer, which receives video content and vibration signal information created to match the video from a distribution server, outputs a sound equivalent to a first target output signal from a built-in speaker (together with the audio of the video content) in accordance with the video content displayed on the display, and further outputs a vibration equivalent to a second target output signal from a built-in vibrator (in accordance with the display of the video).

[0092] As described above in detail, in the present invention, attention is paid to vibrations specific to the target sound signal portion, and the target output signal is generated by regarding the sound signal portion in the original sound signal corresponding to the vibration as the target sound signal portion. Specifically, an adjustment operator that amplifies or relatively increases the sound signal portion corresponding to the vibration is determined based on the vibration signal, and the target output signal is generated from the original sound signal using this adjustment operator.

[0093] Here, the computational processing using such an adjustment operator is a process with a much lighter burden than, for example, acoustic signal extraction processing using known machine learning such as DNN, etc. Therefore, according to the present invention, it is possible to generate a signal related to the target sound (target output signal) through computational processing with a lighter burden.

[0094] Furthermore, according to the present invention, as one example of application, it is possible to reproduce in real time at the receiving end a virtual auditory and tactile space that corresponds to the acoustic and vibration environment of the source in a remote location in web conferences, remote sessions, remote conferences, and even online concerts, which have become widely used in recent years.

[0095] Recreating such virtual spaces is highly effective for online music classes and seminars that demonstrate playing musical instruments, online science classes and seminars that demonstrate experiments using laboratory equipment, online cooking seminars that demonstrate cooking using kitchen utensils, and even online physical education classes and sports seminars that demonstrate various movements. This makes it possible to provide high-quality, immersive classes and seminars, especially for participants in areas lacking in infrastructure. In other words, this invention can contribute to Goal 4 of the United Nations' Sustainable Development Goals (SDGs), which states, "Ensure inclusive and equitable quality education and promote lifelong learning opportunities for all."

[0096] With respect to the various embodiments of the present invention described above, various changes, modifications, and omissions within the scope of the technical spirit and perspective of the present invention may be easily made by those skilled in the art. The above description is merely an example and is not intended to be limiting in any way. The present invention is limited only by the claims and their equivalents. [Explanation of symbols]

[0097] 1 Sound signal conditioning device 101 Communication Interface 102 Acoustic signal / vibration signal storage unit 111 Adjustment operator determination part 111a First operator determination part 111b Second operator determination part 112 target output signal generation unit 112a First target signal generation section 112b Second objective signal generation section 121 Communication control unit 2. Microphone 3. Accelerometer 4. Microphone with accelerometer 5 speakers 6 Haptic Devices

Claims

1. An audio signal adjustment program that causes a computer to function to identify or sort an audio signal portion to be acquired, the audio signal portion relating to a sound generated in association with a predetermined action among a plurality of predetermined actions, from an original audio signal including information on a sound generated in association with an action, the audio signal adjustment program comprising: an adjustment operator determination means for determining an adjustment operator that amplifies or relatively increases a vibration signal portion in the acquired vibration signal, the vibration signal portion including information on vibrations generated together with the sound in accordance with the action, the vibration signal portion corresponding to the vibration signal portion related to the vibration generated together with the sound related to the target sound signal portion, based on information on whether the vibration intensity of the generated vibration, obtained from the acquired vibration signal, falls within a predetermined range set in advance for the certain action; a target output signal generating means for applying the adjustment operator to the original acoustic signal to generate a target output signal in which the target acoustic signal portion is amplified or relatively increased, or in which the acoustic signal portion other than the target acoustic signal portion is reduced or eliminated; to make the computer function, A predetermined range of the intensity of the vibration is set in advance for each of the plurality of predetermined actions. An acoustic signal adjustment program comprising:

2. The acoustic signal adjustment program according to claim 1, characterized in that the adjustment operator determination means determines an adjustment operator that, when the amplitude of the vibration signal or a representative value of the amplitude is within a predetermined amplitude range for the certain specified motion, amplifies or relatively increases the part of the original acoustic signal at that time as the acoustic signal part to be acquired.

3. The target output signal generating means generates a plurality of target output signals corresponding to each of the plurality of predetermined motions, the target output signals including information on sounds generated together with vibrations whose amplitudes or representative values ​​of the amplitudes are within a predetermined amplitude range for the predetermined motions.

3. The acoustic signal adjustment program according to claim 2.

4. the adjustment operator determination means determines, as the first adjustment operator, an operator that multiplies the amplitude of the vibration signal or the representative value of the amplitude by an adjustment coefficient that is a monotonically decreasing function when the target sound signal portion is a sound signal portion related to a predetermined action accompanied by smaller vibrations and the amplitude or the representative value of the amplitude of the corresponding vibration signal is included in an amplitude range less than a predetermined threshold; and / or determines, as the second adjustment operator, an operator that multiplies the amplitude of the vibration signal by an adjustment coefficient that is a monotonically increasing function when the target sound signal portion is a sound signal portion related to a predetermined action accompanied by larger vibrations and the amplitude or the representative value of the amplitude is included in an amplitude range equal to or greater than a predetermined threshold; The desired output signal generating means generates a first desired output signal by applying a first adjustment operator to the original sound signal, and / or generates a second desired output signal by applying a second adjustment operator to the original sound signal.

4. The acoustic signal adjustment program according to claim 1, wherein:

5. the adjustment operator determination means determines, as a first adjustment operator, an operator that performs a Fourier transform on the original sound signal, multiplies the result of the Fourier transform by an adjustment coefficient that is a monotonically decreasing function for the frequency components of the vibration signal, and further performs an inverse Fourier transform on the result of the multiplication, when the target sound signal portion is a sound signal portion related to a predetermined movement accompanied by smaller vibrations and the amplitude or a representative value of the amplitude of the corresponding vibration signal is within an amplitude range less than a predetermined threshold; and / or determines, as a second adjustment operator, an operator that performs a Fourier transform on the original sound signal, multiplies the result of the Fourier transform by an adjustment coefficient that is a monotonically increasing function for the frequency components of the vibration signal, and further performs an inverse Fourier transform on the result of the multiplication, when the target sound signal portion is a sound signal portion related to a predetermined movement accompanied by larger vibrations and the amplitude or a representative value of the amplitude is within an amplitude range equal to or greater than a predetermined threshold; The desired output signal generating means generates a first desired output signal by applying a first adjustment operator to the original sound signal, and / or generates a second desired output signal by applying a second adjustment operator to the original sound signal.

4. The acoustic signal adjustment program according to claim 1, wherein:

6. The acoustic signal adjustment program according to claim 4 or 5, characterized in that the target output signal generation means generates the first target output signal and the second target output signal as signals to be output to a first speaker and a second speaker having different specifications or positions, respectively, or generates the first target output signal and the second target output signal as signals to be output to a certain speaker and a vibration generating device.

7. The acoustic signal adjustment program according to any one of claims 1 to 6, characterized in that the vibration signal is an acceleration vibration signal having an acceleration amplitude generated by an acceleration sensor installed at a position capable of detecting vibrations generated in association with the movement.

8. An audio signal adjustment device that identifies or sorts an audio signal to be acquired, the audio signal relating to a sound generated in association with a predetermined action among a plurality of predetermined actions set in advance, from an original audio signal including information on a sound generated in association with an action, the audio signal adjustment device comprising: an adjustment operator determination means for determining an adjustment operator that amplifies or relatively increases a vibration signal portion in the acquired vibration signal, the vibration signal portion including information on vibrations generated together with the sound in accordance with the action, the vibration signal portion corresponding to the vibration signal portion related to the vibration generated together with the sound related to the target sound signal portion, based on information on whether the vibration intensity of the generated vibration, obtained from the acquired vibration signal, falls within a predetermined range set in advance for the certain action; a target output signal generating means for applying the adjustment operator to the original acoustic signal to generate a target output signal in which the target acoustic signal portion is amplified or relatively increased, or in which the acoustic signal portion other than the target acoustic signal portion is reduced or eliminated; and A predetermined range of the intensity of the vibration is set in advance for each of the plurality of predetermined actions. An acoustic signal conditioning device comprising:

9. 1. An acoustic signal adjustment method implemented by a computer for identifying or sorting an acoustic signal to be acquired, the acoustic signal relating to a sound generated in association with a predetermined action among a plurality of predetermined actions, from an original acoustic signal including information on a sound generated in association with the action, the acoustic signal adjustment method comprising: determining an adjustment operator for amplifying or relatively increasing a vibration signal portion in the acquired vibration signal, the vibration signal portion including information on vibrations generated together with the sound in accordance with the action, the portion of the original acoustic signal corresponding to the vibration signal portion related to vibrations generated together with the sound related to the target acoustic signal portion, based on information on whether the vibration intensity of the generated vibrations, obtained from the acquired vibration signal, falls within a predetermined range set in advance for the certain action; applying the adjustment operator to the original acoustic signal to generate a target output signal in which the target acoustic signal portion is amplified or relatively increased, or in which the acoustic signal portion other than the target acoustic signal portion is reduced or eliminated; and A predetermined range of the intensity of the vibration is set in advance for each of the plurality of predetermined actions.

10. A method for adjusting an acoustic signal, comprising:

Citation Information

Patent Citations

  • Vibration and acoustic measurement integral type sensor

    JP1988108235A

  • Sound detecting method and device using the same

    JP2003222553A

  • Voice processing apparatus and voice processing method

    JP2003264883A

  • Shoe and information processing system cooperating with the shoe

    JP2021094373A

  • Signal processing device, method, and program

    WO2021054152A1