Learning device, estimation device, learning method, estimation method, and program
The learning device enhances 3D pose estimation by using a weight estimation process to adapt to varying sound waves and environments, improving the accuracy and convenience of posture estimation using acoustic signals.
Patent Information
- Application Number
- JP2024082150
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-12-03
AI Technical Summary
Conventional 3D pose estimation techniques using acoustic signals are vulnerable to occlusion and darkness, and the accuracy may be affected by the lighting environment, and the accuracy may be poor in environments where precision equipment is present, and the accuracy may be poor unless the sound waves used to estimate the posture are the same regardless of the object or the space in which the object exists.
A learning device that includes a control unit that performs learning of a learning object model, which is a mathematical model that estimates the posture of the posture of the object to be estimated based on information about a first sound wave, which is a sound wave, which is a sound wave that has propagated through a specified space in a state where the object to be estimated exists, and in the learning object model, a weight estimation process that estimates a weight for each frequency component or each period in a time domain of a sound wave to be processed, which is a sound wave related to the first sound wave, and a posture estimation process that estimates the posture based on the weights estimated by the weight estimation process, and in the learning, the content of the weight estimation process is updated based on the estimation result of the posture estimation process.
Improves the convenience of posture estimation by enabling accurate estimation using any sound wave, reducing the need for consistent sound waves and environmental conditions.
Smart Images

Figure 2025175847000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a learning device, an estimation device, a learning method, an estimation method, and a program. [Background technology]
[0002] Technology for estimating the 3D pose of objects, such as the human body, is expected to be applicable to a wide range of applications, including rehabilitation support, elderly care, and disaster relief. Previously, techniques using visible light signals, including RGB video, and wireless signals have been proposed. However, methods using RGB video are vulnerable to occlusion and darkness, and the relatively high-resolution measurement information they acquire poses challenges in terms of protecting personal information. Furthermore, methods using wireless signals may be limited in environments where precision equipment is present, such as in medical settings or on passenger aircraft. One method that can address these challenges is the use of acoustic signals. Acoustic signals have wavelengths (m-scale) much longer than visible light signals (nm-scale) or wireless signals (cm-scale), making them prone to diffraction and less susceptible to occlusion. Furthermore, acoustic signals have the advantage that their estimation accuracy is not affected by the lighting environment, making them unrestricted in situations where precision equipment is present. Patent Document 1 and Non-Patent Document 1, listed below, are known for 3D pose estimation using acoustic signals. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-104109 [Non-patent literature]
[0004] [Non-Patent Document 1] Shibata, Kawashima, Isogawa, Irie, Kimura, Aoki, “Listening human behavior: 3D human pose estimation with acoustic signals,” Proc. IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the conventional techniques described in Patent Document 1 and Non-Patent Document 1, unless the sound waves used to estimate the posture are the same regardless of the object whose posture is to be estimated or the space in which the object exists, the estimation accuracy may be poor.
[0006] In view of the above circumstances, the present invention aims to improve the convenience of posture estimation by providing a technology that enables posture estimation using any sound wave. [Means for solving the problem]
[0007] One aspect of the present invention is a learning device that includes a control unit that performs learning of a learning object model, which is a mathematical model that estimates the posture of an object to be estimated based on information about a first sound wave, which is a sound wave that has propagated through a specified space in a state where the object to be estimated exists, and in the learning object model, a weight estimation process that estimates a weight for each frequency component or each period in a time domain of a sound wave to be processed, which is a sound wave related to the first sound wave, and a posture estimation process that estimates the posture based on the weights estimated by the weight estimation process are performed, and in the learning, the content of the weight estimation process is updated based on the estimation result of the posture estimation process.
[0008] One aspect of the present invention is an estimation device comprising: a control unit that performs learning of a learning object model, which is a mathematical model that estimates the posture of an estimation object based on information about a first sound wave, which is a sound wave that has propagated through a specified space in a state where an estimation object is present; in the learning object model, a weight estimation process that estimates a weight for each frequency component or each period in a time domain of a processing object sound wave, which is a sound wave related to the first sound wave, and a posture estimation process that estimates the posture based on the weights estimated by the weight estimation process, and in the learning, the content of the weight estimation process is updated based on the estimation result of the posture estimation process, and an estimation unit that performs estimation using a trained learning object model obtained by a learning device.
[0009] One aspect of the present invention is a learning method performed by a learning device, which includes a control unit that learns a learning object model, which is a mathematical model that estimates the posture of an object to be estimated based on information about a first sound wave, which is a sound wave that has propagated through a specified space in a state where the object to be estimated exists, and in which the learning object model executes a weight estimation process that estimates a weight for each frequency component or each period in a time domain of a sound wave to be processed, which is a sound wave related to the first sound wave, and a posture estimation process that estimates the posture based on the weights estimated by the weight estimation process, and in which the learning updates the content of the weight estimation process based on the estimation result of the posture estimation process, and the learning method includes a control step that performs the learning.
[0010] One aspect of the present invention is an estimation method executed by an estimation device comprising an estimation unit that performs estimation using a trained learning object model obtained by a learning device, the estimation method including an estimation step of performing the estimation, the estimation method including a control unit that performs learning of a learning object model, which is a mathematical model that estimates the posture of an estimation object based on information about a first sound wave, which is a sound wave that has propagated through a specified space in a state where an estimation object is present, and the learning object model performs a weight estimation process that estimates a weight for each frequency component or each period in a time domain of a processing object sound wave, which is a sound wave related to the first sound wave, and a posture estimation process that estimates the posture based on the weights estimated by the weight estimation process, and the learning process updates the content of the weight estimation process based on the estimation result of the posture estimation process.
[0011] One aspect of the present invention is a program for causing a computer to function as either the learning device or the estimation device described above. [Effects of the Invention]
[0012] The present invention makes it possible to improve the convenience of technology for estimating posture using sound waves. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is an explanatory diagram illustrating an information processing system according to an embodiment. [Figure 2] FIG. 4 is an explanatory diagram illustrating an example of how first sound wave information is acquired in the embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a hardware configuration of a learning device according to an embodiment. [Figure 4] 10 is a flowchart showing an example of a flow of processing executed by a learning device according to an embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of a hardware configuration of an estimation apparatus according to an embodiment. [Figure 6] 1 is a flowchart showing an example of a flow of processing executed by an estimation device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] (Embodiment) FIG. 1 is an explanatory diagram illustrating an information processing system 100 according to an embodiment. The information processing system 100 includes a learning device 1, an estimation device 2, a sound wave information acquisition device 3, and a sound wave emitting device 4. However, the information processing system 100 is only required to include at least the learning device 1 and the estimation device 2, and this does not apply to the sound wave information acquisition device 3 and the sound wave emitting device 4. Therefore, the information processing system 100 may include, for example, the learning device 1, the estimation device 2, and the sound wave information acquisition device 3, but not the sound wave emitting device 4. The information processing system 100 may include, for example, the learning device 1, the estimation device 2, and the sound wave information acquisition device 3, but not the sound wave information acquisition device 3. The information processing system 100 may include, for example, the learning device 1 and the estimation device 2, but not the sound wave information acquisition device 3.
[0015] The sound wave emitting device 4 emits sound waves used for posture estimation. The sound wave emitted from the sound wave emitting device 4 is, for example, a zeroth sound wave, which will be described later. The sound wave emitting device 4 is typically a speaker, but may be any type of device.
[0016] The ultrasonic information acquisition device 3 converts the propagating ultrasonic waves into information indicating the ultrasonic waves (hereinafter referred to as "sonic information"). The ultrasonic information may be, for example, an acoustic signal (i.e., a time series), in which case the ultrasonic information acquisition device 3 is, for example, a microphone. The ultrasonic wave propagating to the ultrasonic information acquisition device 3 is, for example, a first ultrasonic wave, which will be described later. The ultrasonic wave propagating to the ultrasonic information acquisition device 3 is, for example, a zeroth ultrasonic wave, which will be described later. The first ultrasonic information, which will be described later, is an example of ultrasonic information.
[0017] The learning device 1 includes a control unit 11 including a processor 91, such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit) or an NPU (Neural Network Processing Unit), and a memory 92, which are connected via a bus.
[0018] The control unit 11 executes, for example, a learning process. The learning process is a process of learning a learning object model, which is a mathematical model to be learned. The learning process is a process of learning the learning object model until a predetermined condition related to the end of learning (hereinafter referred to as a "learning end condition") is met. The learning end condition may be any condition related to the end of learning, and may be, for example, a condition that the learning object model has been updated a predetermined number of times, or a condition that the change in the learning object model due to the update is smaller than a predetermined change. The learning object model at the time when the learning end condition is met is a trained learning object model.
[0019] The learning object model is a mathematical model that estimates the posture of the estimation object based on information indicating a first sound wave (hereinafter referred to as "first sound wave information"), which is a sound wave that propagates through a specified space in a state where the estimation object exists.
[0020] Since the first sound wave is a sound wave that propagates through a specified space in the presence of the object to be estimated, the first sound wave may include sound waves that are scattered or reflected by the object to be estimated, sound waves that have passed through the object to be estimated, and sound waves that are not scattered or reflected by the object to be estimated and that have not passed through the object to be estimated.
[0021] In the learning model, at least a weight estimation process and a posture estimation process are executed. The weight estimation process is a process of estimating weights for each frequency component or each period in the time domain of a sound wave related to a first sound wave (hereinafter referred to as a "sound wave to be processed") from information about the sound wave itself to be processed or information about the sound wave to be processed (hereinafter referred to as "sound wave information to be processed"). A period in the time domain of a sound wave is a section in the time direction when the sound wave information is expressed in a time series. Therefore, estimating a weight for each period in the time domain of a sound wave indicated by sound wave information means estimating a weight for one or more sections that divide the period from the start time t0 of the sound wave indicated by the sound wave information to the end time tf of the sound wave.
[0022] The first sound wave is, for example, a sound emitted from a speaker in a space where the estimation target exists and picked up by a microphone. The sound wave to be processed may be, for example, this first sound wave. For example, if the first sound wave is a sound emitted from a speaker in a space where the estimation target exists and picked up by a microphone, the sound wave to be processed may be a pair of the first sound wave and the sound emitted from the speaker and picked up by the microphone when the estimation target does not exist in that space, or the sound emitted from the speaker itself (hereinafter, such a sound wave emitted for posture estimation will be referred to as the "zeroth sound wave"). The sound wave to be processed may also be the difference between the zeroth sound wave and the first sound wave (hereinafter referred to as the "second sound wave"). In this definition of the second sound wave, the composite wave of the zeroth sound wave and the second sound wave is the first sound wave.
[0023] The posture estimation process is a process of estimating the posture of an object to be estimated based on the weights estimated by the weight estimation process. More specifically, the posture estimation process is a process of estimating the posture of an object to be estimated based on the weights estimated by the weight estimation process and the sound waves to be processed or the sound wave information to be processed that was used when the weights were estimated in the weight estimation process.
[0024] The first sound wave information may represent information about the first sound wave, for example, by representing the first sound wave in a time series (i.e., waveform), or by acoustic features acquired by executing an acoustic feature extraction process that extracts acoustic features, such as a spectrogram or an intensity vector shown in Reference 1 below, from an acoustic signal corresponding to the first sound wave. Furthermore, the first sound wave information may represent information about the first sound wave by an embedding vector acquired by executing an acoustic signal embedding process that is executed by an acoustic signal embedding model that is separately prepared from the acoustic signal corresponding to the first sound wave.
[0025] Reference 1: Yasuda, Koizumi, Saito, Uematsu, Imoto, "Sound event localization based on sound intensity vector refined by DNN-based denoising an source separation," arXiv:2002.05994.
[0026] Similarly, with regard to the sound wave information to be processed, the information on the sound wave to be processed may be represented by an acoustic signal time series corresponding to the sound wave to be processed. The information on the sound wave to be processed may be represented by one or more sets of features acquired by executing an acoustic feature extraction process on the acoustic signal corresponding to the sound wave to be processed. The information on the sound wave to be processed may be represented by an embedding vector acquired by executing an acoustic signal embedding process on the sound wave to be processed.
[0027] The same applies to the second sound wave information, but the second sound wave information may be represented by an acoustic feature or an embedding vector obtained by explicitly calculating the second sound wave, which is the difference between the zeroth sound wave and the first sound wave, and then performing an acoustic feature extraction process or an acoustic signal embedding process on the second sound wave.The second sound wave information may be represented by calculating the first sound wave information and the zeroth sound wave information separately and then using the difference between them.The second sound wave information is information representing the second sound wave.
[0028] The weight estimation process may be, for example, a self-attention mechanism. More specifically, the weight estimation process may be, for example, a mechanism that calculates a weight for the input sound wave information to be processed from the input itself and weights the input, and may be a self-attention mechanism that uses the zeroth sound wave information as a key and the second sound wave information as a query. This allows a weight to be estimated for each frequency component of the second sound wave or each period in the time domain.
[0029] Note that the model to be trained may undergo pre-processing prior to the weight estimation process. The pre-processing is a process of acquiring information indicating the sound wave to be processed in a format to be input to the weight estimation process. The pre-processing includes, for example, a process of acquiring information indicating the second sound wave (i.e., second sound wave information) based on information indicating the first sound wave (i.e., first sound wave information) and information indicating the zeroth sound wave (i.e., zeroth sound wave information) (hereinafter referred to as "second sound wave information acquisition process"). Alternatively, a process of acquiring information indicating the second sound wave based on the first sound wave and the zeroth sound wave may be executed as pre-processing as second sound wave information acquisition process.
[0030] In this case, the information indicating the first sound wave may take any form as long as the form in which the information indicating the first sound wave and the form in which the information indicating the zeroth sound wave take is the same, and may be a time series, an acoustic feature such as a spectrogram, or an embedded vector.
[0031] In the pre-processing, a process of converting the format of the information indicating sound waves (hereinafter referred to as "format conversion process") may be performed. For example, in the pre-processing, a process of converting the format from a time series to a spectrogram may be performed. This process is typically a discrete Fourier transform or a short-time discrete Fourier transform. In addition, in the pre-processing, a process of converting the format from a time series to an acoustic feature such as an intensity vector or an embedding vector may be performed. In addition, in the pre-processing, a process of converting the format from a spectrogram to another acoustic feature or embedding vector may be performed.
[0032] The second sound wave acquisition process may be executed before or after the format conversion process. Alternatively, the second sound wave acquisition process may be executed after the format conversion process, and the format conversion process may be executed after the second sound wave acquisition process.
[0033] Note that each process that can be executed in the weight estimation process does not necessarily have to be executed by the learning target model, and may be executed by, for example, a device other than the learning device.
[0034] 2 is an explanatory diagram illustrating an example of how first sound wave information is acquired in an embodiment. An object 901 shown in FIG. 2 is an example of an estimation object. Therefore, the space shown in FIG. 2 is an example of a predetermined space in which an estimation object exists.
[0035] FIG. 2 shows that two speakers 902, speaker 902-1 and speaker 902-2, are positioned in front of target 901. As mentioned above, the example in FIG. 2 is merely an example, and the number of speakers does not necessarily have to be two, but may be one, or three or more. FIG. 2 shows that sound waves W101 are output from speaker 902. More specifically, sound waves W101-1 are output from speaker 902-1, and sound waves W101-2 are output from speaker 902-2. Both sound waves W101 are output from speakers 902 toward target 901.
[0036] FIG. 2 shows that microphone 903 is located behind target 901. As described above, speaker 902 is located in front of target 901 and outputs sound waves W101 toward target 901. Microphone 903 is located behind target 901. Therefore, at least a portion of sound waves W101 output by speaker 902 acts on target 901, and at least a portion of the sound waves that propagate to microphone 903, both those that acted on target 901 and those that did not, are picked up by microphone 903. In this case, the number of microphones does not necessarily have to be one, and multiple microphones may be located in different positions. Note that microphone 903 is an example of a device that collects sound, and sound collection does not necessarily have to be performed by microphone 903. Sound collection may be performed using a device that integrates multiple microphones into one, such as an Ambisonics microphone or a microphone array.
[0037] The sound waves collected by the microphone 903 or the like in this way are an example of the first sound waves, and information indicating the collected sound waves is an example of the first sound wave information. Note that when a sound wave acts, it means that the sound wave is scattered or reflected by an acting object (object 901 in the example of FIG. 2), or passes through the acting object.
[0038] 2 from which only the object 901 has disappeared is an example of a predetermined space in a state in which an estimation object does not exist. If the sound wave output from the speaker 902 and picked up by the microphone 903 in the space shown in FIG. 2 where the object 901 exists is an example of the first sound wave in the second sound wave acquisition process, an example of the zeroth sound wave in that process is the sound wave output from the speaker 902 and picked up by the microphone 903 in the space from which only the object 901 has disappeared. Alternatively, the zeroth sound wave may be the sound itself output from the speaker 902.
[0039] <About updates during learning process> In the learning process, the training object model is updated based on the results of the posture estimation process. In updating the training object model, the content of the weight estimation process is updated. Updating the content of the weight estimation process means updating the weighting rules. Learning by the learning process may be, for example, unsupervised learning or supervised learning, and the content of the weight estimation process is updated to improve the accuracy of estimation by the posture estimation process. In updating the training object model, the content of the posture estimation process may also be updated. In the case of supervised learning, labeled data is input to the training object model, and the label of the labeled data is information that indicates the posture of the estimation object. In the case of supervised learning, the training object model is updated to reduce the difference between the posture indicated by the label and the posture estimated by the posture estimation process.
[0040] The estimation device 2 performs estimation processing. The estimation processing is processing for performing estimation using a trained learning object model obtained in a learning processing. In the estimation processing, for example, the trained learning object model executes a weight estimation processing and a posture estimation processing based on the sound wave information of the processing object, thereby estimating the posture of the estimation object. Furthermore, in the estimation processing, pre-processing may be performed before executing the weight estimation processing, as in the learning processing.
[0041] <Effects of learning processing> In the learning process, learning of the weight estimation process is performed. Therefore, the trained weight estimation process can estimate which parts of the sound waves to be processed are important for posture estimation according to the sound waves to be processed. Since the posture estimation process performs estimation using the results of the weight estimation process, degradation of estimation accuracy due to changes in the space in which the estimation target exists or changes in the sound waves is smaller than when the results of the weight estimation process are not used. Therefore, the trained training target model obtained in the learning process reduces the need to make the sound waves used for posture estimation the same regardless of the posture estimation target or the environment in which the estimation target exists. In other words, the trained training target model obtained in the learning process improves the convenience of the technology for posture estimation using sound waves. In this way, the learning process can improve the convenience of the technology for posture estimation using sound waves.
[0042] <Other effects> As can be seen from the above definition, the second sound wave is the difference between the first sound wave and the zeroth sound wave, and the zeroth sound wave is background sound. Therefore, the second sound wave can be said to be a sound wave that extracts the influence of the estimation target. Therefore, when weight estimation processing is performed on the second sound wave, the estimation accuracy is higher than when weight estimation processing is performed on the first sound wave.
[0043] <Example of hardware configuration> 3 is a diagram showing an example of the hardware configuration of the learning device 1 according to the embodiment. The learning device 1 includes a control unit 11 and executes a program. By executing the program, the learning device 1 functions as a device including the control unit 11, an interface unit 12, and a storage unit 13.
[0044] More specifically, the processor 91 reads out a program stored in the storage unit 13 and stores the read out program in the memory 92. When the processor 91 executes the program stored in the memory 92, the learning device 1 functions as a device including the control unit 11, the interface unit 12, and the storage unit 13.
[0045] The control unit 11 controls the operation of each functional unit of the learning device 1. The control unit 11 acquires, for example, sound wave information acquired by the sound wave information acquisition device 3. The control unit 11 executes, for example, a learning process. The learning process involves, for example, learning using the first sound wave information acquired by the control unit 11. The control unit 11 acquires, for example, information stored in the memory unit 13. The process of acquiring the information stored in the memory unit 13 is specifically reading.
[0046] The interface unit 12 includes a communication interface for connecting the learning device 1 to an external device. The interface unit 12 communicates with the external device via wired or wireless communication.
[0047] The external device is, for example, a device that transmits information used in the execution of the control unit 11. In such a case, the interface unit 12 acquires the information used in the execution of the control unit 11 from the device that transmits the information used in the execution of the control unit 11. The device that transmits the information used in the execution of the control unit 11 is, for example, a device that transmits first sound wave information. The sound wave information acquisition device 3 is an example of such an external device. The device that transmits the information used in the execution of the control unit 11 may be, for example, a device that transmits information indicating the zeroth sound wave. The sound wave transmission device 4 is an example of a device that transmits the zeroth sound wave. The device that transmits the information used in the execution of the control unit 11 may be, for example, a device that transmits information indicating the second sound wave. A combination of the sound wave information acquisition device 3 and the sound wave transmission device 4 is an example of such an external device. The device that transmits the information used in the execution of the control unit 11 may be, for example, a device that transmits a label, i.e., information indicating posture.
[0048] The external device may be, for example, the estimation device 2. In this case, the estimation device 2 can execute the trained learning object model obtained by executing the learning process through communication via the interface unit 12.
[0049] Interface unit 12 may be configured to include input devices such as a mouse, keyboard, or touch panel. Interface unit 12 may be configured as an interface that connects these input devices to learning device 1. In this way, the input devices of interface unit 12 accept input of various information to learning device 1 via wired or wireless connections. Note that information does not necessarily have to be input to the communication interface of interface unit 12, but may also be input to the input devices of interface unit 12.
[0050] Interface unit 12 outputs, for example, various types of information. Interface unit 12 is configured to include a display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, or an organic EL (Electro-Luminescence) display, and a speaker. Interface unit 12 may be configured as an interface that connects these display devices or speakers to learning device 1. Therefore, the display devices and speakers included in interface unit 12 output, for example, information acquired by the communication interface of interface unit 12 or information input to the input device of interface unit 12 as images or sounds.
[0051] The storage unit 13 is configured using a computer-readable storage medium (non-transitory computer-readable recording medium) such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 13 stores various information related to the learning device 1. The storage unit 13 stores, for example, various information generated by the operation of the control unit 11. The storage unit 13 stores, for example, information used in the learning process and information generated by the execution of the learning process. The storage unit 13 may exist on a cloud, for example.
[0052] 4 is a flowchart showing an example of the flow of processing executed by the learning device 1 in the embodiment. The control unit 11 executes the learning processing (step S101).
[0053] <Example of hardware configuration of estimation device 2> 5 is a diagram illustrating an example of a hardware configuration of the estimation device 2 according to an embodiment. The estimation device 2 includes a control unit 21 including a processor 93 such as a CPU, GPU, or NPU, and a memory 94, which are connected via a bus, and executes a program. By executing the program, the estimation device 2 functions as a device including the control unit 21, an interface unit 22, and a storage unit 23.
[0054] More specifically, the processor 93 reads out a program stored in the storage unit 23 and stores the read program in the memory 94. The processor 93 executes the program stored in the memory 94, causing the estimation device 2 to function as a device including the control unit 21, the interface unit 22, and the storage unit 23.
[0055] The control unit 21 controls the operation of each functional unit included in the estimation device 2. The control unit 21 acquires, for example, sound wave information acquired by the sound wave information acquisition device 3. The control unit 21 performs estimation using, for example, a trained learning object model. Here, the trained learning object model is executed on, for example, the first sound wave information acquired by the control unit 21. As a result, the posture of the estimation object is estimated. The control unit 21 acquires, for example, information stored in the memory unit 23. Specifically, the process of acquiring the information stored in the memory unit 23 is reading.
[0056] The interface unit 22 includes a communication interface for connecting the estimation device 2 to an external device. The interface unit 22 communicates with the external device via wired or wireless communication. The external device is, for example, a device that transmits information to be input to the trained training object model. The interface unit 22 acquires the information to be input to the trained training object model by communicating with the device that transmits the information to be input to the trained training object model. The ultrasonic information acquisition device 3 is an example of such an external device.
[0057] The information to be input to the trained learning model is, for example, first sound wave information. The information to be input to the trained learning model is, for example, information indicating the second sound wave. The information to be input to the trained learning model is, for example, information indicating the zeroth sound wave.
[0058] The external device may be, for example, the learning device 1. In this case, the estimation device 2 can execute the learned learning object model obtained by the learning device 1 through communication via the interface unit 22.
[0059] The interface unit 22 may be configured to include input devices such as a mouse, a keyboard, a touch panel, etc. The interface unit 22 may be configured as an interface that connects these input devices to the estimation device 2. In this way, the input devices of the interface unit 22 accept input of various information to the estimation device 2 via wired or wireless connections. Note that information does not necessarily have to be input to the communication interface of the interface unit 22, and may also be input to the input devices of the interface unit 22.
[0060] The interface unit 22 outputs, for example, various types of information. The interface unit 22 includes, for example, a display device such as a CRT display, a liquid crystal display, or an organic EL display, and a speaker. The interface unit 22 may be configured as an interface that connects these display devices or speakers to the estimation device 2. Therefore, the interface unit 22 may output, for example, information input to an input device of the interface unit 22 as an image or sound.
[0061] The storage unit 23 is configured using a computer-readable storage medium device (non-transitory computer-readable recording medium) such as a magnetic hard disk device or a semiconductor storage device. The storage unit 23 stores various information related to the estimation device 2. The storage unit 23 stores various information generated by the operation of the control unit 21, for example. The storage unit 23 may exist on a cloud, for example.
[0062] 6 is a flowchart showing an example of the flow of processing executed by the estimation device 2 in the embodiment. The control unit 21 executes the estimation processing (step S201). By executing the estimation processing, the posture of the estimation target is estimated based on, for example, first sound wave information. By executing the estimation processing, the posture of the estimation target may be estimated based on, for example, second sound wave information.
[0063] The learning device 1 configured in this way executes the learning process, which can improve the convenience of the technique for estimating posture using sound waves, as described in <Effects of the learning process>.
[0064] Furthermore, the estimation device 2 configured in this manner performs estimation using the learned learning object model obtained by the learning device 1. This improves the convenience of the technique for estimating posture using sound waves.
[0065] Furthermore, the information processing system 100 configured in this manner includes the learning device 1. This makes it possible to improve the convenience of the technique for estimating posture using sound waves.
[0066] (Variation) The learning device 1 may be implemented using multiple information processing devices connected to each other via a network, in which case the functional units of the learning device 1 may be distributed and implemented across the multiple information processing devices.
[0067] The estimation device 2 may be implemented using a plurality of information processing devices communicably connected via a network, in which case the respective functional units of the estimation device 2 may be distributed and implemented among the plurality of information processing devices.
[0068] All or part of the functions of the information processing system 100, the learning device 1, and the estimation device 2 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. The program may be transmitted via a telecommunications line.
[0069] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Explanation of symbols]
[0070] 100...information processing system, 1...learning device, 11...control unit, 12...interface unit, 13...memory unit, 2...estimation device, 21...control unit, 22...interface unit, 23...memory unit, 3...sound wave information acquisition device, 4...sound wave transmission device, 91...processor, 92...memory, 93...processor, 94...memory, 901...object, 902...speaker, 903...microphone
Claims
1. a control unit that learns a learning object model, which is a mathematical model that estimates the posture of the estimation object based on information about a first sound wave that is a sound wave propagated through a predetermined space in a state where the estimation object is present; Equipped with In the learning model, a weight estimation process is executed to estimate a weight for each frequency component or each period in a time domain of a processing target sound wave, which is a sound wave related to the first sound wave, and a posture estimation process is executed to estimate the posture based on the weight estimated by the weight estimation process, In the learning, the content of the weight estimation process is updated based on the estimation result of the posture estimation process. Learning device.
2. The sound wave to be processed is the first sound wave itself. The learning device according to claim 1 .
3. The sound wave to be processed is a zeroth sound wave, which is a sound wave output from a sound source of the first sound wave or a sound wave propagated through the space in a state where the estimation target is not present. The learning device according to claim 1 .
4. The sound wave to be processed is information about a pair of the first sound wave and a zeroth sound wave, which is a sound wave propagated in the space in a state where the estimation target is not present, or information about a sound wave that represents the difference between the first sound wave and the zeroth sound wave. The learning device according to claim 1 .
5. an estimation unit that performs estimation using a trained learning object model obtained by a learning device, the estimation unit including: a control unit that performs learning of a learning object model, which is a mathematical model that estimates the posture of the estimation object based on information about a first sound wave, which is a sound wave that has propagated through a predetermined space in a state where the estimation object exists; wherein the learning object model executes a weight estimation process that estimates a weight for each frequency component or each period in a time domain of a processing object sound wave, which is a sound wave related to the first sound wave, and a posture estimation process that estimates the posture based on the weight estimated by the weight estimation process; and An estimation device comprising:
6. a control unit that performs learning of a learning object model, which is a mathematical model that estimates the posture of an estimation object based on information about a first sound wave that is a sound wave propagated through a predetermined space in a state where an estimation object exists; in the learning object model, a weight estimation process that estimates a weight for each frequency component or each period in a time domain of a processing object sound wave that is a sound wave related to the first sound wave, and a posture estimation process that estimates the posture based on the weight estimated by the weight estimation process are performed; and in the learning, content of the weight estimation process is updated based on an estimation result of the posture estimation process, a control step for performing the learning; A learning method that has
7. an estimation method executed by an estimation device including an estimation unit that performs estimation using a trained learning object model obtained by a learning device, the estimation method including: a control unit that trains a learning object model, which is a mathematical model that estimates the posture of an estimation object based on information about a first sound wave, which is a sound wave that has propagated through a predetermined space in a state where an estimation object exists; in the learning object model, a weight estimation process that estimates a weight for each frequency component or each period in a time domain of a processing object sound wave, which is a sound wave related to the first sound wave, and a posture estimation process that estimates the posture based on the weight estimated by the weight estimation process are executed; and in the learning, content of the weight estimation process is updated based on an estimation result of the posture estimation process, an estimation step of performing the estimation; An estimation method having:
8. A program for causing a computer to function as either the learning device according to any one of claims 1 to 4 or the estimation device according to claim 5.
Citation Information
Patent Citations
Posture estimation method, posture estimation device and program
JP2023104109A