State estimation device, state estimation method, and computer program
The state estimation system addresses the challenge of positional changes by using adversarial learning to isolate state estimation components, improving accuracy in estimating object states.
Patent Information
- Application Number
- JP2024122907
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Conventional methods for estimating the state of an object using acoustic signals struggle with accuracy when the position of the object changes, as fluctuations in sound waves make it difficult to distinguish between position and state changes.
A state estimation system that utilizes adversarial learning to separate position and state estimation components, using a state estimator trained to accurately estimate the object's state by removing position-dependent components from acoustic features.
The system achieves higher accuracy in estimating the state of an object regardless of positional fluctuations, enhancing the precision of state estimation.
Smart Images

Figure 2026021206000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a state estimation device, a state estimation method, and a computer program. [Background technology]
[0002] There has been a demand for a method for estimating the state (e.g., posture) of an object without using a visible light camera. As one example, a technique has been proposed that uses sound waves (acoustic signals) instead of visible light. For example, a technique has been proposed in which sound waves are emitted from a speaker into a space in which an object (e.g., a person) exists, and the sound waves are acquired by a microphone (hereinafter referred to as a "microphone") and used in an estimation process (see, for example, Patent Document 1 and Non-Patent Document 1). The sound waves used in such estimation processes are sound waves emitted from a speaker and acquired by the microphone after being affected by the object.
[0003] In these conventional techniques, the position of the measurement object is implicitly fixed, and if the position of the measurement object changes, the accuracy of the pose estimation significantly deteriorates.
[0004] More specifically, conventional technology assumes that the speaker, object, and microphone are aligned in a straight line. When the position of the object is fixed, the state of the sound waves emitted from the speaker and recorded by the microphone due to the influence of the object (for example, blocking, reflection, and diffraction) is expected to be roughly similar if the state of the object is the same. In other words, fluctuations in the sound waves recorded by the microphone can be considered to be caused by fluctuations in the state of the object. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2023-104109 [Non-patent literature]
[0006] [Non-Patent Document 1] Shibata, Kawashima, Isogawa, Irie, Kimura, Aoki, “Listening human behavior: 3D human pose estimation with acoustic signals,” Proc. IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. Summary of the Invention [Problem to be solved by the invention]
[0007] However, when the position of the object being measured changes, the effects (blocking, reflection, diffraction) of the object also change. As a result, with conventional technology, it is difficult to distinguish whether the fluctuations in the sound waves recorded by the microphone are due to changes in the state of the object or changes in the object's position. This makes it difficult to estimate the state of the object with high accuracy.
[0008] The present invention has been made in consideration of the above-mentioned circumstances, and provides a technology that makes it possible to estimate the state of an object with higher accuracy regardless of fluctuations in the position of the object. [Means for solving the problem]
[0009] One aspect of the present invention is a state estimation device that includes a control unit that estimates a state of an object by using an acoustic signal acquired in a space in which the object is located and a state estimator that is a model that represents the correspondence between the acoustic signal and the state of the object and that is obtained in advance, where the state estimator is an estimator that is obtained by a learning process that is performed so that a position estimator that estimates the position of the object using the acoustic signal becomes unable to estimate the position of the object.
[0010] One aspect of the present invention is a state estimation method that estimates a state of an object by using an acoustic signal acquired in a space in which the object is located and a state estimator that is a model that represents the correspondence between the acoustic signal and the state of the object and that is obtained in advance, where the state estimator is an estimator obtained by a learning process that is performed so that a position estimator that estimates the position of the object using the acoustic signal becomes unable to estimate the position of the object.
[0011] One aspect of the present invention is a computer program for causing a computer to function as a state estimation device, which includes a control unit that estimates the state of an object by using an acoustic signal acquired in a space in which the object is located and a state estimator that is a model that represents the correspondence between the acoustic signal and the state of the object and that is obtained in advance, wherein the state estimator is an estimator that is obtained by a learning process that is performed so that a position estimator that estimates the position of the object using the acoustic signal becomes unable to estimate the position of the object. [Effects of the Invention]
[0012] According to the present invention, it is possible to estimate the state of an object with higher accuracy regardless of fluctuations in the position of the object. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a schematic block diagram showing a system configuration of a state estimation system 100 according to the present invention. [Figure 2] 2 is a schematic block diagram showing a specific example of the functional configuration of a learning device 30. FIG. [Figure 3] 10 is a flowchart showing a specific example of the flow of processing by the learning device 30. [Figure 4] FIG. 2 is a diagram showing an outline of the flow of data in the learning device 30. [Figure 5] FIG. 2 is a diagram showing an outline of the flow of data in the learning device 30. [Figure 6] 2 is a schematic block diagram showing a specific example of the functional configuration of a state estimation device 40. FIG. [Figure 7]10 is a flowchart showing a specific example of the flow of processing by the state estimation device 40. [Figure 8] FIG. 10 shows the results of an experiment. [Figure 9] FIG. 2 is a diagram illustrating an outline of an example of the hardware configuration of an information processing device 90 applied to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] In the following explanation, subscripts (characters written in relatively small letters to the lower right of another character) may be indicated by adding an underscore to the other character. For example, if a subscript B is written in relatively small letters to the lower right of an A, it may be indicated as "A_B." Superscripts (characters written in relatively small letters to the upper right of another character) may be indicated by adding a sharp to the other character. For example, if a superscript B is written in relatively small letters to the upper right of an A, it may be indicated as "A#B." Characters may also be indicated with a hat "^" above them. For example, an A with a hat above it may be indicated as "A^." An A with a hat above it, with a subscript B and a superscript C, may be indicated as "A^_B#C."
[0015] [System Overview] FIG. 1 is a schematic block diagram showing the system configuration of a state estimation system 100 of the present invention. The state estimation system 100 is a system used to estimate the state of an object 80. The object 80 is an object of a predetermined type. The object 80 may be, for example, a living thing such as an animal or a plant, or an object such as a robot. The object 80 is an object whose state, such as its outer shape or surface structure, can change over time or due to changes in the environment (e.g., humidity, temperature, etc.). The estimated state refers to a shape that can physically change, such as its outer shape or surface structure. For example, if the object 80 is an animal (e.g., a person) or a robot, the estimated state may be the posture of the animal (e.g., a person) or robot. For example, if the object 80 is a plant or an object, the estimated state may be its shape.
[0016] For example, the state estimation system 100 includes one or more speakers 10, a microphone 20, and a state estimation device 40. The state estimation device 40 estimates the state of the object 80 using a trained model generated by the learning device 30. The state estimation system 100 may further include the learning device 30.
[0017] The speaker 10 is a device capable of outputting an acoustic signal. Any type of speaker may be used as the speaker 10. The speaker 10 may be a device that generates any type of acoustic signal. A specific example of an acoustic signal generated by the speaker 10 is a time stretched pulse (TSP) signal. The TSP signal is a signal that is generated by continuously increasing the frequency from low to high frequencies.
[0018] The microphone 20 acquires an acoustic signal generated from the speaker 10. A part or all of the acoustic signal acquired by the microphone 20 has been affected by the object 80. Examples of the influence of the object 80 include blocking, reflection, and diffraction. In the following description, the acoustic signal acquired by the microphone 20 may be represented as S_T. For example, a general microphone may be applied to the microphone 20, or an Ambisonics microphone that employs the Ambisonics method to record the entire sound in a space in a 360-degree circle may be applied.
[0019] The relative positional relationship between the speaker 10 and the microphone 20 may be defined in any way. For example, as shown in Fig. 1, the relative positional relationship may be defined so that the microphone 20 is located on the center line of the multiple speakers 10, which is the direction in which sound emitted from the multiple speakers 10 travels.
[0020] The learning device 30 performs a learning process using a plurality of pieces of learning data, each of which includes a plurality of acoustic signals acquired by the microphone 20, position information indicating the position of the object 80 when the acoustic signals were acquired, and state information indicating the state of the object 80 when the acoustic signals were acquired. The state estimation device 40 uses a trained model obtained as a result of the learning process by the learning device 30 to estimate the state of the object 80 based on a new acoustic signal acquired by the microphone 20.
[0021] [Processing Overview] In the learning process, a position estimator is used to estimate the position of the object 80 based on acoustic features obtained from the acoustic signal. Adversarial learning is performed using the position estimator and a state estimator that estimates the state of the object 80 based on the acoustic features. In the adversarial learning, the following two steps of processing are executed alternately. Position estimator learning process: A process of learning the position estimator so that the position estimator can correctly estimate the position of the object 80. State estimator learning process: The acoustic feature and position estimator are learned so that the state estimator can correctly estimate the state of the object 80 and the position estimator can accurately estimate the position of the object 80.
[0022] By using this type of adversarial learning, it becomes possible to remove components that contribute to position estimation from the acoustic features and selectively extract only components that contribute to state estimation, thereby performing state estimation. Note that since the information that we actually want to estimate is the state (e.g., posture) of the object 80, once the learning process is complete, the position estimator is no longer necessary.
[0023] A predetermined pre-processing may be performed on the acoustic signal included in the data used for the learning process (the acoustic signal acquired by the microphone 20). For example, as a pre-processing step, an acoustic signal may be generated that has been processed so that the leading time is shifted by a certain time, and the pre-processed acoustic signal may also be used to perform the learning process.
[0024] [Processing details] 2 is a schematic block diagram showing a specific example of the functional configuration of learning device 30. Learning device 30 is configured using an information processing device such as a personal computer or a server device. Learning device 30 includes a learning data acquisition unit 31, a storage unit 32, a control unit 33, and an information output unit 34.
[0025] The training data acquisition unit 31 is an interface that acquires data from another device. The training data acquisition unit 31 may read acoustic signal data recorded on a recording medium such as a DVD-ROM or a USB memory (Universal Serial Bus Memory). The training data acquisition unit 31 may receive acoustic signal data from the microphone 20. In this case, data exchange between the training data acquisition unit 31 and the microphone 20 may be performed using wired communication or wireless communication. If the training device 30 is built into an information processing device equipped with the microphone 20, the training data acquisition unit 31 may receive acoustic signals from a bus. Alternatively, the training data acquisition unit 31 may receive acoustic signal data from another information processing device via a network. The training data acquisition unit 31 may be configured in any manner as long as it is capable of receiving input of acoustic signals acquired by the microphone 20.
[0026] The storage unit 32 is configured using a storage device such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 32 stores data used by the control unit 33. The storage unit 32 functions as, for example, a training data storage unit 321, a preprocessed training acoustic signal storage unit 322, and a training result storage unit 323.
[0027] The training data storage unit 321 stores training data. The training data includes an acoustic signal used in the training process (hereinafter referred to as a "training acoustic signal"), position data indicating the position of the object 80 when the training acoustic signal was obtained, and state data indicating the state of the object 80 when the training acoustic signal was obtained. In the following description, the training acoustic signal is represented as S_L.
[0028] The preprocessed training acoustic signal storage unit 322 stores a preprocessed training acoustic signal SL obtained by performing preprocessing on the training acoustic signal S_L.
[0029] The learning result storage unit 323 stores data indicating the state estimator obtained by the learning process (data of the learned model).
[0030] The control unit 33 is configured using a processor such as a CPU (Central Processing Unit) and a memory (main storage device). The control unit 33 functions as an information control unit 331, a preprocessing unit 332, a feature acquisition unit 333, and a learning unit 334 when the processor executes a program. Note that all or part of the functions of the control unit 33 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The above program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, and a semiconductor storage device (e.g., a solid-state drive (SSD)), as well as storage devices such as a hard disk or semiconductor storage device built into a computer system. The above program may be transmitted via a telecommunications line.
[0031] The information control unit 331 controls the input and output of data in the learning device 30. For example, the information control unit 331 controls the learning data acquisition unit 31 to acquire learning data from outside the learning device 30. The information control unit 331 records the acquired learning data in the learning data storage unit 321. The information control unit 331 outputs the trained model obtained by the learning process to an external device. For example, the information control unit 331 may output the trained model via the information output unit 34.
[0032] The preprocessing unit 332 performs predetermined preprocessing on the training acoustic signal to obtain a preprocessed training acoustic signal. For example, the predetermined preprocessing may be a process of deleting a portion of the training acoustic signal. Specific examples of such preprocessing are described below.
[0033] For example, when a TSP signal is used as a training acoustic signal, a process of deleting a part of the signal in a region including the beginning in the time axis direction may be performed as preprocessing. The length of the deleted signal may be set to be shorter than the length of one period of the TSP signal.
[0034] The feature acquisition unit 333 acquires features (acoustic features) of the acoustic signal used in the learning process. For example, when both a training acoustic signal and a preprocessed training acoustic signal are used in the learning process, the feature acquisition unit 333 acquires acoustic features for each of the training acoustic signal and the preprocessed training acoustic signal. For example, when a preprocessed training acoustic signal is used in the learning process, the feature acquisition unit 333 acquires acoustic features for the preprocessed training acoustic signal.
[0035] A specific example of the processing of the feature acquisition unit 333 will be described below. In the following description, a feature acquired from an acoustic signal (hereinafter referred to as an "acoustic feature") will be represented as F_T. An acoustic feature is typically expressed as a vector or a set of vectors. Any method for extracting an acoustic feature may be used. For example, a Mel spectrogram may be used as the acoustic feature, or an Intensity vector may be used as the acoustic feature. A Mel spectrogram is obtained by performing a short-time Fourier transform on an acoustic signal to convert it into a Mel scale. An Intensity vector is an acoustic feature that is widely used when predicting the direction from which a sound is coming. A specific example of a method for acquiring each acoustic feature will be described below.
[0036] First, a method for acquiring the Mel spectrogram feature amount will be described. The Mel spectrogram feature amount may be acquired using, for example, the following mathematical formula 1.
[0037]
number
[0038] In Equation 1, f is the frequency, t is the time, c is the channel, k is the Mel frequency bin index, H_mel(k) is the coefficient of the Mel filter bank at Mel frequency bin k, h(k) is the frequency interval included in Mel frequency bin k, and F(F,S) is the real part of the frequency f component in the Fourier transform of the time series signal S.
[0039] That is, the Mel spectrogram feature can be expressed as a third-order tensor with dimensions of the number of Mel frequency bins, the number of time points, and the number of channels.
[0040] Note that the "channel" in this description is a value that depends on the microphone 20. For example, if the microphone 20 is configured using a stereo microphone, it corresponds to the left-right directions X and Y. That is, the number of channels in this case is two. Also, if the microphone 20 is configured using an Ambisonics microphone, it corresponds to the non-directional component W and the directional components X, Y, and Z. That is, the number of channels in this case is four.
[0041] Next, a method for acquiring the intensity vector feature amount will be described. The intensity vector feature amount may be acquired using the technique disclosed in Reference 1 below, for example. Reference 1: Yasuda, Koizumi, Saito, Uematsu, Imoto, “Sound event localization based on sound intensity vector refined by DNN-based denoising and source separation,” In Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020.
[0042] For example, the intensity vector feature amount may be obtained using the following mathematical expressions 2 and 3. The value obtained from Expression 3 corresponds to the intensity vector feature amount.
[0043]
number
[0044]
number
[0045] W, X, Y, and Z represent the signals obtained by short-time Fourier transform (STFT) of the omnidirectional component calculated from the Ambisonics microphone and the signals of each axis (vertical Y, horizontal X, height Z). R(·) represents the real part of a complex number, and * represents the complex conjugate. In this way, the intensity vector feature can be expressed as a third-order tensor with dimensions of number of Mel frequency bins × number of times × (number of channels - 1).
[0046] Next, a description will be given of the processing of the learning unit 334. The learning unit 334 performs a learning process to generate a state estimator used when estimating the state of the object 80. In the learning process of the state estimator, training acoustic features are used as input.
[0047] The state estimator is a model that represents the correspondence between state data that expresses the state of the object 80 and acoustic features. For example, the state estimator may be configured using a neural network that receives acoustic features as input and outputs state data. By training this state estimator, it becomes possible to estimate state data from acoustic features with higher accuracy.
[0048] The state data may be represented using any data that can represent the state of the object 80. For example, the state data may be configured using data that indicates the position of each feature point included in the object 80. For example, the state data may be configured using data that represents the state of the object 80 relative to a predetermined reference state. For example, the state data may be configured using data that represents the position or area of each part that makes up the object 80.
[0049] In the training process, state training data consisting of a combination of training acoustic features and correct state data is used. A plurality of such state training data are prepared in advance, and the training process is executed. The values of the network parameters of the state estimator are updated so that the training estimated state data, which is the output when the training acoustic features are input to the state estimator, approaches the correct state data.
[0050] A specific example of how to obtain a state estimator will be described below. More specifically, the values of the network parameters of the state estimator are successively updated so as to reduce the loss function L_pose expressed by the following equation 4.
[0051]
number
[0052] p^=(p^_1,p^_2,...,P^_T)#T=P_Θ_P(F_L) indicates the estimated state data. p=(p_1,p_2,...,P_T)#T indicates the correct state data. P_Θ() indicates the state estimator with network parameters Θ. Θ_P indicates the network parameters of the state estimator. F_L indicates the training acoustic features. T indicates the length of the state data output from the state estimator.
[0053] The learning process further uses position learning data composed of a combination of training acoustic features and ground truth position data. A plurality of pieces of training data are prepared in advance, each of which combines such position learning data with state learning data acquired under the same circumstances (same object, same position, same state). The learning process is performed using the position learning data. The values of the network parameters of the state estimator are updated so that the training estimated position data, which is the output when the training acoustic features are input to the position estimator, does not approach the ground truth position data.
[0054] A specific example of processing using a position estimator will be described below. More specifically, the values of the network parameters of the state estimator are successively updated so as to reduce the weighted sum of either or both of the loss function L_std(Θ_P) expressed by the following formula 5 and L_pos(Θ_P) expressed by the following formula 6.
[0055]
number
[0056]
number
[0057] l^=(l^_1,l^_2,...,l^_T)#T=G_(Θ_G,Θ_P)(F_L) indicates the estimated location data. l=(l_1,l_2,...,l_T)#T indicates the correct location data. G_Θ() indicates the location estimator with network parameters Θ. Θ_G indicates the network parameters of the location estimator. std() indicates the standard deviation of the vector elements.
[0058] By performing such processing, it becomes possible to remove components that contribute to position estimation from the acoustic feature amount and to further emphasize components that contribute to state estimation.
[0059] It is also possible to adopt a method of training the state estimator so that the state data does not fluctuate suddenly. For example, the network parameters of the state estimator are successively updated so as to reduce the loss function L_smooth(Θ_P) expressed by the following equation 7.
[0060]
number
[0061] The state estimator may be trained based on all of the loss functions described above. For example, the values of the network parameters of the state estimator may be successively updated so as to reduce the value of the loss function expressed by the following Equation 8.
[0062]
number
[0063] w_pose, w_smooth, and w_pos indicate the weights of each loss function. Next, the learning process of the position estimator will be described. In the learning process of the position estimator, a position estimator that estimates the position of the object 80 using acoustic features is obtained. The learned acoustic features are used as input, and the position estimator is obtained as output. The position estimator is a model that represents the correspondence between position data that expresses the position of the object 80 and acoustic features.
[0064] For example, a neural network that receives acoustic features as input and outputs position data may be used as the position estimator. By training such a position estimator, it becomes possible to estimate position data from acoustic features with higher accuracy.
[0065] The position estimator may be configured in any manner. The position data may be configured in any manner as long as it can represent the position of the object 80. For example, the position data may be defined as information indicating the absolute three-dimensional position of the object 80. For example, the position data may be defined as a relative three-dimensional position using any one of the speakers 10 as a reference point. For example, the position data may be defined as a relative three-dimensional position using the microphone 20 as a reference point. For example, the position data may be defined as information indicating the distance from a line on a horizontal plane connecting the speaker 10 and the microphone 20.
[0066] In the training process of the position estimator, position training data consisting of a combination of training acoustic features and ground truth position data is used. The values of the network parameters of the position estimator are updated so that the estimated position data, which is the output when the training acoustic features are input to the position estimator, approaches the ground truth position data. More specifically, the values of the network parameters of the position estimator are successively updated so as to reduce the weighted sum of either or both of the loss function L_std(Θ_G) expressed by the following equation 9 or the loss function L_pos(Θ_G) expressed by the following equation 10.
[0067]
number
[0068]
number
[0069] The information output unit 34 is an interface that outputs data to another device. The information output unit 34 may record data on a recording medium such as a DVD-ROM or a USB memory. The information output unit 34 may transmit data of the trained model of the state estimator to an information processing device such as a state estimation device. In this case, data may be exchanged between the information output unit 34 and the information processing device using wired communication or wireless communication.
[0070] FIG. 3 is a flowchart showing a specific example of the processing flow of the learning device 30. First, the information control unit 331 acquires learning data (step S101). The information control unit 331 records the acquired learning data in the learning data storage unit 321. Next, the preprocessing unit 332 generates a preprocessed learning acoustic signal by performing preprocessing (e.g., data augmentation) on the acoustic signal included in the learning data (step S102). The preprocessing unit 332 records the preprocessed learning acoustic signal in the preprocessed learning acoustic signal storage unit 322. Next, the feature acquisition unit 333 acquires acoustic features from the acoustic signal of the learning data or the preprocessed learning acoustic signal (step S103). The learning unit 334 performs learning processing using the acoustic features and the position data and state data corresponding to the acoustic signal from which the acoustic features were obtained (step S104). At this time, the learning unit 334 acquires a learned model of the state estimator by alternately and repeatedly executing the learning processing of the position estimator and the learning processing of the state estimator (step S105). The information control unit 331 outputs the acquired trained model (trained state estimator) to another device such as the state estimation device 40 (step S106).
[0071] 4 and 5 are diagrams showing an outline of the data flow in the learning device 30. First, a multi-channel acoustic signal is acquired as a training acoustic signal. The multi-channel acoustic signals are used as a single acoustic signal. The acoustic signals may be used as they are, or a preprocessed training acoustic signal obtained by performing preprocessing such as shifting the signals by a predetermined time (e.g., time α) may be used, or both may be used. Acoustic features are acquired from each acoustic signal. For example, a log-Mel spectrum and an intensity vector may be used as the acoustic features. Then, a state estimator (Pose Estimation Module) and a position estimator (Position Discriminator Module) are used to acquire loss functions L_pose, L_smooth, and L_std, respectively.
[0072] 6 is a schematic block diagram showing a specific example of the functional configuration of the state estimation device 40. The state estimation device 40 is configured using an information processing device such as a personal computer or a server device. The state estimation device 40 includes an acoustic signal acquisition unit 41, a storage unit 42, and a control unit 43.
[0073] The acoustic signal acquisition unit 41 is an interface that acquires data from another device. The acoustic signal acquisition unit 41 may read acoustic signal data recorded on a recording medium such as a DVD-ROM or a USB memory. The acoustic signal acquisition unit 41 may receive acoustic signal data from the microphone 20. In this case, data exchange between the acoustic signal acquisition unit 41 and the microphone 20 may be performed using wired communication or wireless communication. When the state estimation device 40 is built into an information processing device equipped with the microphone 20, the acoustic signal acquisition unit 41 may receive acoustic signals from a bus. Alternatively, the acoustic signal acquisition unit 41 may receive acoustic signal data from another information processing device via a network. The acoustic signal acquisition unit 41 may be configured in any manner as long as it is capable of receiving input of an acoustic signal acquired by the microphone 20.
[0074] The storage unit 42 is configured using a storage device such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 42 stores data used by the control unit 43. The storage unit 42 functions as, for example, a trained model storage unit 421 and an acoustic signal storage unit 422.
[0075] The trained model storage unit 421 stores trained model data. The trained model data is data generated and output by the learning device 30.
[0076] The acoustic signal storage unit 422 stores acoustic signals used to estimate the state of the object 80. The acoustic signal storage unit 422 stores acoustic signals acquired by the acoustic signal acquisition unit 41, for example.
[0077] The control unit 43 is configured using a processor such as a CPU and a memory (main storage device). The control unit 43 functions as an information control unit 431, a feature acquisition unit 432, and a state estimation unit 433 by the processor executing a program. All or part of the functions of the control unit 43 may be realized using hardware such as an ASIC, PLD, or FPGA. The above program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and semiconductor storage devices (e.g., SSDs), as well as storage devices such as hard disks and semiconductor storage devices built into a computer system. The above program may be transmitted via a telecommunications line.
[0078] The information control unit 431 controls the input and output of data in the state estimation device 40. For example, the information control unit 431 controls the acoustic signal acquisition unit 41 to acquire an acoustic signal to be processed from outside the state estimation device 40. The acoustic signal to be processed is an acoustic signal obtained regarding the state of the object 80 to be estimated. The information control unit 431 records the acquired acoustic signal in the acoustic signal storage unit 422. The information control unit 431 may further acquire a trained model obtained by the training process from the training device 30. Such training data may be acquired by wireless communication or wired communication with the training device 30, or may be acquired via a recording medium. The information control unit 431 may output information indicating the estimation result of the state of the object 80 to an external information processing device or output device.
[0079] The feature acquisition unit 432 acquires acoustic features by performing the same processing as that performed by the feature acquisition unit 333 of the learning device 30 on the acoustic signal to be processed.
[0080] The state estimation unit 433 estimates the state of the object 80 using the acoustic features acquired by the feature acquisition unit 432 and the trained model (trained state estimator) stored in the trained model storage unit 421. The state estimation unit 433 performs estimation processing to acquire, for example, state data indicating the state of the object 80.
[0081] 7 is a flowchart showing a specific example of the processing flow of the state estimation device 40. First, the information control unit 431 acquires an acoustic signal to be processed (step S201). The information control unit 431 records the acquired acoustic signal in the acoustic signal storage unit 422. Next, the feature acquisition unit 432 acquires acoustic features from the acoustic signal to be processed (step S202). The state estimation unit 433 executes a state estimation process using the acoustic features and a trained model stored in the trained model storage unit 421 (step S203). The information control unit 431 outputs the acquired estimation result (state data) (step S204).
[0082] Next, we conducted experiments (data collection and computer experiments) to compare this system with other methods. An Ambisonics microphone was used as microphone 20. A motion capture system was used to obtain correct posture data. Data collection was conducted in a typical classroom-like environment. Therefore, external noise and room reverberations could be picked up by the microphone. A human was used as object 80, and the subject was instructed to stand at a designated position on a horizontal line perpendicular to the line connecting speaker 10 and microphone 20. Three distances from the line connecting speaker 10 and microphone 20 were used: 0 cm, 50 cm, and 100 m. The subject, representing object 80, was instructed to change posture as a concrete example of a state. Specifically, the subject was instructed to assume one of five postures (walking in place, crouching, bowing, raising both arms until they were parallel to the horizontal plane, and simply standing with hands down) in random order. Twenty-one commonly used human feature points were used to capture human features using the motion capture system. As a result, we obtained a total of about three hours of training audio signals from six people, along with the corresponding correct posture and position data. Part of this data was used as evaluation data, and the rest was used as training data.
[0083] The effectiveness of this system was verified using the training data and evaluation data described above, and the following three methods were compared.
[0084] Comparative technology 1: Technology described in the following document Jiang, Xue, Miao, Wang, Lin, Tian, Murali, Hu, Sun, Su, “Towards 3D human pose construction using Wi-Fi,” Proc. International Conference on Mobile Computing and Networking (MobiCom), 2020. This comparative technology 1 is a method for estimating a person's three-dimensional posture using radio waves. In this computer experiment, the neural network configuration of this technology was used as is, but the input was changed to a Mel spectrogram and intensity vector.
[0085] Comparative technology 2: Technology described in the following document Ginosar, Bar, Kohavi, Chan, Owens, Malik, “Learning individual styles of conversational gesture,” Proc. IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. This comparative technique 2 uses only the Mel spectrogram as input.
[0086] Comparative Technology 3: Technology described in Patent Document 1
[0087] The following three indicators are used for evaluation. Metric 1: Root mean square error (RMSE) We calculated the root mean square error between the estimated pose data and the ground truth pose data. The smaller this value, the more accurate the estimation. Metric 2: Mean Absolute Error (MAE) We calculated the mean absolute error between the estimated pose data and the ground truth pose data. The smaller this value, the more accurate the estimation. Metric 3: Percentage of correctly estimated feature points (PCK h@0.5) Feature points that could be estimated with an error smaller than 0.5 times the distance between the head and neck feature points were considered to be "correctly estimated," and the percentage of such feature points was calculated. A higher value indicates a more accurate estimation.
[0088] Figure 8 shows the results of the above-mentioned experiment. "Single Subject" indicates a condition where the person of object 80 included in the evaluation data is always included in the training data (i.e., a condition where the person's movements can be learned to some extent). "Cross Subject" indicates a condition where the person of object 80 included in the evaluation data is not included in the training data (i.e., a condition where the person's movements cannot be learned in advance). As shown in Figure 8, our system (Ours) showed the best performance under all conditions and in all indicators.
[0089] FIG. 9 is a diagram illustrating an outline of an example hardware configuration of an information processing device 90 applied to the present embodiment. The information processing device 90 includes a processor 91, a main storage device 92, a communication interface 93, an auxiliary storage device 94, an input / output interface 95, and an internal bus 96. The processor 91, the main storage device 92, the communication interface 93, the auxiliary storage device 94, and the input / output interface 95 are communicably connected to each other via the internal bus 96. The information processing device 90 may be applied to, for example, the learning device 30 and the state estimation device 40. In this case, for example, the learning data acquisition unit 31, the information output unit 34, and the acoustic signal acquisition unit 41 may be configured using the communication interface 93 or the input / output interface 95. For example, the memory unit 32 and the memory unit 42 may be configured using the auxiliary storage device 94. Furthermore, the control unit 33 and the control unit 43 may be configured using the processor 91 and the main storage device 92.
[0090] (Variation) In the learning process and the state estimation process, the acoustic signal may be used as is without using acoustic features. In this configuration, the learning device 30 is configured not to include the feature acquisition unit 333, and the state estimation device 40 is configured not to include the feature acquisition unit 432.
[0091] In this embodiment, the learning device 30 and the state estimation device 40 are configured as separate devices, but they may also be configured as an integrated device.
[0092] The learning device 30 may be implemented using multiple information processing devices. For example, the learning device 30 may be implemented using a device such as a cloud. For example, in the learning device 30, the memory unit 32 and the control unit 33 may be implemented in different information processing devices. For example, the memory unit 32 of the learning device 30 may be distributed and implemented in multiple information processing devices.
[0093] The state estimation device 40 may be implemented using a plurality of information processing devices. For example, the state estimation device 40 may be implemented using a device such as a cloud. For example, in the state estimation device 40, the storage unit 42 and the control unit 43 may be implemented in different information processing devices. For example, the storage unit 42 of the state estimation device 40 may be implemented in a distributed manner in a plurality of information processing devices.
[0094] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Explanation of symbols]
[0095] 100...state estimation system, 10...speaker, 20...microphone, 30...learning device, 31...learning data acquisition unit, 32...storage unit, 321...learning data storage unit, 322...preprocessed training acoustic signal storage unit, 323...learning result storage unit, 33...control unit
Claims
1. a control unit that estimates a state of the object by using an acoustic signal acquired in a space in which the object is located and a state estimator that is a model that represents a correspondence between the acoustic signal and a state of the object and that is obtained in advance; A state estimation device, wherein the state estimator is an estimator obtained by a learning process that is performed so that a position estimator that estimates the position of the object using the acoustic signal becomes unable to estimate the position of the object.
2. the state estimator is an estimator obtained using a feature of the acoustic signal, The state estimation device according to claim 1 , wherein the control unit acquires a feature amount based on the acoustic signal, and estimates the state of the object by using the feature amount and the state estimator.
3. the control unit performs a learning process for the position estimator so that the position estimator can more accurately estimate the position of the object based on the acoustic signal, 2. The state estimation device according to claim 1, wherein the control unit performs, as a learning process for the state estimator, a learning process so that the state estimator can more accurately estimate the state of the object based on the acoustic signal, and a learning process so that the position estimator can more inaccurately estimate the position of the object based on the acoustic signal.
4. The state estimation device according to claim 1 , wherein the acoustic signal is obtained by collecting a sound emitted from a speaker in a space where the object is located with a microphone.
5. The state estimation device according to claim 4 , wherein the microphone is an Ambisonics microphone that employs an Ambisonics method for recording the entire sound in a space in a 360-degree circle.
6. Estimating a state of the object by using an acoustic signal acquired in a space in which the object is located and a state estimator that is a model representing a correspondence between the acoustic signal and a state of the object and that is obtained in advance; A state estimation method, wherein the state estimator is an estimator obtained by a learning process that is performed so that a position estimator that estimates the position of the object using the acoustic signal becomes unable to estimate the position of the object.
7. a control unit that estimates a state of the object by using an acoustic signal acquired in a space in which the object is located and a state estimator that is a model that represents a correspondence between the acoustic signal and a state of the object and that is obtained in advance; A computer program for causing a computer to function as a state estimation device, wherein the state estimator is an estimator obtained by a learning process that is performed so that a position estimator that estimates the position of the object using the acoustic signal becomes unable to estimate the position of the object.
Citation Information
Patent Citations
Posture estimation method, posture estimation device and program
JP2023104109A