Posture estimation method, posture estimation device, and program

The acoustic signal-based posture estimation method addresses the challenges of dark environments and restricted radio wave usage by using acoustic signals to estimate object posture, achieving effective and non-invasive posture estimation.

JP7694910B2Active Publication Date: 2025-06-18NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2022004911
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-17
Publication Date
2025-06-18
Estimated Expiration
2042-01-17

Smart Images

  • Figure 0007694910000001
    Figure 0007694910000001
  • Figure 0007694910000002
    Figure 0007694910000002
  • Figure 0007694910000003
    Figure 0007694910000003
Patent Text Reader

Abstract

To estimate noninvasively the posture of an object even in a darkroom environment or in an environment where the use of equipment that emits radio waves is restricted.SOLUTION: A posture estimation device includes an acoustic signal acquisition unit, a feature extraction unit, and a posture estimation unit. The acoustic signal acquisition unit acquires an acoustic signal. The feature extraction unit extracts an acoustic feature from the acoustic signal acquired by the acoustic signal acquisition unit. The posture estimation unit inputs the acoustic feature extracted by the feature extraction unit to a posture estimator, which is a model representing the correspondence between the acoustic features and posture data representing the posture of the object, to obtain an estimated posture of the object.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a posture estimation method, a posture estimation device, and a program.

Background Art

[0002] There is a great need in various fields such as healthcare, nursing care, and sports for non-invasive human posture estimation that does not require a person to wear a device to be measured. Conventionally, many methods that use camera images / videos as input have been proposed as non-invasive human posture estimation methods (see, for example, Non-Patent Document 1). However, when using camera images / videos as input, since visible light wavelength signals are used, there is a problem that the accuracy significantly decreases in a darkroom environment.

[0003] On the other hand, a method has been proposed that enables human posture estimation even in a darkroom environment by using radio wave signals as input (see, for example, Non-Patent Document 2). Also, a method has been proposed that can restore scene information even in a darkroom environment by utilizing acoustic information (see, for example, Non-Patent Document 3).

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

[0005] When radio wave signals are used as input as in the technology of Non-Patent Document 2, it cannot be used in an environment where the use of devices that emit radio waves, such as in a hospital room or an aircraft, is restricted. The technology of Non-Patent Document 3, although it does not use radio, is premised on visualizing objects with a non-complex shape, and it was difficult to estimate the posture of the object.

[0006] In view of the above circumstances, an object of the present invention is to provide a posture estimation method, a posture estimation device, and a program that can non-invasively estimate the posture of an object even in a dark room environment or an environment where the use of devices that emit radio waves is restricted. [Means for Solving the Problems]

[0007] One aspect of the present invention includes an acoustic signal acquisition step of acquiring an acoustic signal, a feature quantity extraction step of extracting an acoustic feature quantity from the acoustic signal acquired in the acoustic signal acquisition step, and inputting the acoustic feature quantity extracted in the feature quantity extraction step into a posture estimator, which is a model representing the correspondence between the acoustic feature quantity and posture data representing the posture of an object based on the positions of respective feature points included in the object, to obtain posture data representing the estimated posture of the object. In the acoustic signal acquisition step, an acoustic signal emitted from a speaker and partially blocked by the object is acquired by a microphone placed at a distance between 50 centimeters and 2 meters from the object. It is a posture estimation method.

[0008] One aspect of the present invention includes an acoustic signal acquisition unit that acquires an acoustic signal, a feature quantity extraction unit that extracts an acoustic feature quantity from the acoustic signal acquired by the acoustic signal acquisition unit, and a posture estimation unit that inputs the acoustic feature quantity extracted by the feature quantity extraction unit into a posture estimator, which is a model representing the correspondence between the acoustic feature quantity and posture data representing the posture of an object based on the positions of respective feature points included in the object, to obtain posture data representing the estimated posture of the object. The acoustic signal acquisition unit acquires an acoustic signal emitted from a speaker and partially blocked by the object, which is the acoustic signal picked up by a microphone placed at a distance between 50 centimeters and 2 meters from the object. It is a posture estimation device.

[0009] One aspect of the present invention is a program for causing a computer to execute the above-described posture estimation method.

Advantages of the Invention

[0010] According to the present invention, it is possible to non-invasively estimate the posture of an object even in a darkroom environment or an environment where the use of devices that emit radio waves is restricted.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Best Mode for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. FIG. 1 is a functional block diagram showing the configuration of a posture estimation device 1 according to an embodiment of the present invention. In FIG. 1, only the functional blocks related to the present embodiment are extracted and shown. The posture estimation device 1 includes an acoustic signal acquisition unit 2, a feature amount extraction unit 3, a posture estimator learning unit 4, and a posture estimation unit 5.

[0013] The acoustic signal acquisition unit 2 acquires the acoustic signal picked up by the microphone and outputs it to the feature amount extraction unit 3. The feature amount extraction unit 3 inputs the acoustic signal from the acoustic signal acquisition unit 2 and extracts the feature amount from the input acoustic signal. The feature amount extracted from the acoustic signal is referred to as an acoustic feature amount.

[0014] The posture estimator learning unit 4 learns the posture estimator. The posture estimator is a model representing the correspondence between the acoustic feature amount and the posture data representing the posture of the object. The posture estimator is, for example, a neural network that takes the acoustic feature amount as an input and outputs the posture data. Any data can be used as the posture data as long as it can represent the posture of the object. For example, the posture data may be data indicating the positions of each feature point included in the object, data representing the relative posture with respect to the reference posture, or data representing the positions or regions of each part constituting the object. In the following, the case where the object is a person will be described as an example, but the object is not limited to a person. The posture estimator learning unit 4 uses the set of the acoustic feature amount extracted by the feature amount extraction unit 3 and the correct posture data as learning data to estimate the value of the network parameter of the posture estimator. The value of the network parameter represents the coupling weight between the nodes constituting the neural network. The network parameter is also referred to as a weight parameter.

[0015] The posture estimation unit 5 inputs the acoustic signal acquired by the acoustic signal acquisition unit 2 into a posture estimator using the values of the network parameters estimated in the posture estimator learning unit 4, and obtains posture data representing the estimated posture of the object as a posture estimation result.

[0016] The posture estimation device 1 is realized by, for example, a computer device. The posture estimation device 1 may be realized by a plurality of computer devices connected to a network. In this case, it can be arbitrary which of these plurality of computer devices realizes each functional unit of the posture estimation device 1. Also, the same functional unit of the posture estimation device 1 may be realized by a plurality of computer devices. For example, the posture estimator learning unit 4 and the posture estimation unit 5 may be realized by different computer devices. In this case, the acoustic signal acquisition unit 2 and the feature amount extraction unit 3 may be realized in either or both of the computer device realizing the posture estimator learning unit 4 and the computer device realizing the posture estimation unit 5, and may further be realized by another computer device.

[0017] FIG. 2 is a device configuration diagram showing an example of the hardware configuration of the posture estimation device 1. The posture estimation device 1 includes a processor 71, a storage unit 72, a communication interface 73, and a user interface 74. The processor 71 is a central processing unit that performs calculations and controls. The processor 71 is, for example, a CPU (central processing unit). The processor 71 realizes the functions of the acoustic signal acquisition unit 2, the feature amount extraction unit 3, the posture estimator learning unit 4, and the posture estimation unit 5 by reading and executing a program from the storage unit 72. The storage unit 72 further has a work area and the like when the processor 71 executes various programs. The communication interface 73 is connected to be communicable with other devices. The user interface 74 is an input device such as a keyboard, a pointing device (mouse, tablet, etc.), a button, a touch panel, etc., and a display device such as a display. An artificial operation is input by the user interface 74.

[0018] Note that all or part of the functions of the acoustic signal acquisition unit 2, the feature quantity extraction unit 3, the posture estimator learning unit 4, and the posture estimation unit 5 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).

[0019] FIG. 3 is a diagram showing an equipment installation environment for acquiring an acoustic signal. The acoustic signal acquisition unit 2 acquires the acoustic signal picked up by the microphone 6. In the present embodiment, the type and position of the microphone are not limited. However, the microphone 6 is preferably an ambisonics microphone capable of acquiring the arrival direction of the acoustic signal. Also, in consideration of the reflection and attenuation of the acoustic signal, the position of the microphone 6 is preferably set at a position about 50 cm to 2 m away from the person 7.

[0020] Further, the present embodiment does not limit active sensing or passive sensing depending on the presence or absence of a speaker. However, for the purpose of ensuring a more measurable intensity of the acoustic signal, as shown in FIG. 3, emitting some acoustic signal from the speaker 8 and picking up the acoustic signal at least partially blocked by the person 7 with the microphone 6 can be cited as an effective equipment installation method.

[0021] FIG. 4 is a flowchart showing the operation of the posture estimation device 1. The acoustic signal acquisition unit 2 acquires the acoustic signal picked up by the microphone 6 for learning of the posture estimator, and outputs it to the feature amount extraction unit 3 (step S1). The feature amount extraction unit 3 inputs the learning acoustic signal acquired by the acoustic signal acquisition unit 2 in step S1. The feature amount extraction unit 3 vectorizes the acoustic feature amount extracted from the acoustic signal in order to input the acoustic signal to the posture estimator learning unit 4, and generates a feature amount vector f (step S2). In the present embodiment, the algorithm for vectorizing the feature amount and the dimension of the feature amount vector f are not limited. However, for example, a method such as using a log mel spectrogram, which is signal information obtained by performing a short-time Fourier transform on an acoustic signal and converting it to a mel scale, as an acoustic feature amount can be considered. Thereby, it is considered that the difference between the actual sound and the human pitch perception can be absorbed, and more effective learning becomes possible. The feature amount extraction unit 3 outputs the generated acoustic feature amount to the posture estimator learning unit 4.

[0022] The posture estimator learning unit 4 acquires the feature amount vector f generated by the feature amount extraction unit 3 in step S2 and the posture data of the value representing the correct posture of the object as learning data. The posture data of the value representing the correct posture may be received from an external device, for example, may be input by the user using an input device (not shown), or may be read from a recording medium. The posture estimator learning unit 4 learns the posture estimator so that the value of the posture data obtained by inputting the feature amount vector f to the posture estimator approaches the value of the correct posture data (step S3). By this learning, the value of the weight parameter of the neural network used as the posture estimator is obtained. The present embodiment does not limit the network structure of the posture estimator. However, for example, it is possible to use the existing network structure described in Non-Patent Document 1. The posture estimator learning unit 4 outputs the value of the weight parameter obtained by learning to the posture estimation unit 5.

[0023] After the learning of the posture estimator, the acoustic signal acquisition unit 2 acquires the acoustic signal picked up by the microphone 6 for posture estimation and outputs it to the feature amount extraction unit 3 (step S4). The feature amount extraction unit 3 generates a feature amount vector f from the acoustic signal acquired by the acoustic signal acquisition unit 2 in step S4 by the same process as in step S2 (step S5). The feature amount extraction unit 3 outputs the generated feature amount vector f to the posture estimation unit 5.

[0024] The posture estimation unit 5 sets the value of the weight parameter calculated by the posture estimator learning unit 4 in step S3 in the posture estimator. The posture estimation unit 5 inputs the feature amount vector f generated by the feature amount extraction unit 3 in step S5 to the posture estimator and obtains posture data representing the estimated posture of the person as an estimation result (step S6). The posture estimation unit 5 outputs the obtained estimation result. The output may be display on a screen, printing by a printing device, transmission to another device connected via a network, or writing to a recording medium.

[0025] According to the present embodiment, it is possible to non-invasively estimate the posture of an object having a complex shape such as a person even in a dark room environment or an environment where the use of devices that emit radio waves is restricted.

[0026] According to the above-described embodiment, the posture estimation device includes an acoustic signal acquisition unit, a feature amount extraction unit, and a posture estimation unit. The acoustic signal acquisition unit acquires an acoustic signal. The feature amount extraction unit extracts an acoustic feature amount from the acoustic signal acquired by the acoustic signal acquisition unit. The acoustic feature amount is, for example, a log mel spectrogram. The posture estimation unit inputs the acoustic feature amount extracted by the acoustic signal acquisition unit to a posture estimator that is a model representing the correspondence between the acoustic feature amount and posture data representing the posture of the object, and obtains posture data representing the estimated posture of the object.

[0027] The posture estimation device may further include a learning acoustic signal acquisition unit, a learning feature quantity extraction unit, and a learning unit. For example, the learning acoustic signal acquisition unit corresponds to the acoustic signal acquisition unit 2 of the embodiment, and the learning feature quantity extraction unit corresponds to the feature quantity extraction unit 3 of the embodiment. The learning acoustic signal acquisition unit acquires a learning acoustic signal. The learning feature quantity extraction unit extracts an acoustic feature quantity from the acoustic signal acquired by the learning acoustic signal acquisition unit. The learning unit learns a posture estimator using the acoustic feature quantity extracted by the learning feature quantity extraction unit and the posture data representing the correct posture of the object. The posture estimation unit inputs the acoustic feature quantity extracted by the feature quantity extraction unit to the posture estimator learned in the learning unit to obtain posture data representing the estimated posture of the object.

[0028] In the acoustic signal acquisition unit and the learning acoustic signal acquisition unit, an acoustic signal may be acquired from a microphone that picks up an acoustic signal emitted from a speaker and at least partially blocked by an object. The microphone is, for example, an ambisonics microphone.

[0029] As described above, the embodiments of the present invention have been described with reference to the drawings. However, it is obvious that the above embodiments are merely examples of the present invention, and the present invention is not limited to the above embodiments. Therefore, additions, omissions, substitutions, and other changes of components may be made without departing from the technical idea and scope of the present invention.

Description of Reference Numerals

[0030] 1... Posture estimation device, 2... Acoustic signal acquisition unit, 3... Feature quantity extraction unit, 4... Posture estimator learning unit, 5... Posture estimation unit, 6... Microphone, 7... Person, 8... Speaker

Claims

1. An acoustic signal acquisition step of acquiring an acoustic signal; A feature amount extraction step of extracting an acoustic feature amount from the acoustic signal acquired in the acoustic signal acquisition step; An estimation step of inputting the acoustic feature amount extracted in the feature amount extraction step into a pose estimator, which is a model representing the correspondence between the acoustic feature amount and the pose data representing the pose of each feature point included in the object, to obtain pose data representing the estimated pose of the object; having; In the acoustic signal acquisition step, an acoustic signal emitted from a speaker and partially blocked by the object, and the acoustic signal is acquired by a microphone placed at a distance between 50 centimeters and 2 meters from the object. A pose estimation method.

2. A learning acoustic signal acquisition step of acquiring a learning acoustic signal; A learning feature amount extraction step of extracting an acoustic feature amount from the acoustic signal acquired in the learning acoustic signal acquisition step; Further having a learning step of learning the pose estimator using the acoustic feature amount extracted in the learning feature amount extraction step and the pose data representing the correct pose of the object; In the estimation step, the acoustic feature amount extracted in the feature amount extraction step is input into the pose estimator learned in the learning step to obtain pose data representing the estimated pose of the object. The pose estimation method according to claim 1.

3. The microphone is an ambisonics microphone. The pose estimation method according to claim 2.

4. The acoustic feature amount is a log mel spectrogram. The pose estimation method according to any one of claims 1 to 3.

5. An acoustic signal acquisition unit that acquires an acoustic signal, A feature amount extraction unit that extracts an acoustic feature amount from the acoustic signal acquired by the acoustic signal acquisition unit, An attitude estimation unit that inputs the acoustic feature amount extracted by the feature amount extraction unit to an attitude estimator, which is a model representing the correspondence between the acoustic feature amount and the attitude data representing the attitude of each feature point included in the object, to obtain attitude data representing the estimated attitude of the object, Comprising, The acoustic signal acquisition unit acquires an acoustic signal emitted from a speaker and partially blocked by the object, and the microphone placed at a distance between 50 centimeters and 2 meters from the object picks up the acoustic signal, An attitude estimation device.

6. A learning acoustic signal acquisition unit that acquires an acoustic signal for learning, A learning feature amount extraction unit that extracts an acoustic feature amount from the acoustic signal acquired by the learning acoustic signal acquisition unit, Further comprising a learning unit that learns the attitude estimator using the acoustic feature amount extracted by the learning feature amount extraction unit and the attitude data representing the correct attitude of the object, The attitude estimation unit inputs the acoustic feature amount extracted by the feature amount extraction unit to the attitude estimator learned in the learning unit to obtain attitude data representing the estimated attitude of the object, The attitude estimation device according to claim 5.

7. A program for causing a computer to execute the attitude estimation method according to any one of claims 1 to 4. ​

Citation Information

Patent Citations

  • Recognizing system of three-dimensional object

    JP1991188391A

  • Sensing system and height measuring system

    JP2005337954A

  • Method of manufacturing fluttering robot using method of preparing fluid-structure interactive numerical model

    JP2008276807A

  • Sound processing device, sound processing method, and sound processing program

    JP2015070321A

  • Anomaly detection system, anomaly detection device, and anomaly detection method

    JP2021185352A