Behavior estimation device, behavior estimation method, and program
Patent Information
- Application Number
- JP2025134621
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-09-08
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-05
Smart Images

Figure 2025166142000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a behavior estimation device, a behavior estimation method, and a program. [Background technology]
[0002] In recent years, there has been a demand to provide users with various services suited to their lifestyles by estimating their behavior (hereinafter also referred to as behavioral information) based on the everyday sounds generated in the homes in which they live.
[0003] For example, Patent Document 1 discloses a technology for identifying the source of sounds detected in a user's home that are classified as real-world sounds based on the learning results of a database that has learned the features and direction of television audio and the features of real-world sounds, and the analysis results of an analysis unit that analyzes the features and sound source direction of the detected sounds, and for estimating the user's behavior in the home based on the identified sound source. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-095517 Summary of the Invention [Problem to be solved by the invention]
[0005] However, with the technology described in Patent Document 1, the captured real-world sounds and television audio are both audible sounds that can be perceived by the human ear, and are susceptible to noise from various everyday sounds, making it difficult to say that human behavior can be estimated with high accuracy.
[0006] Therefore, the present disclosure provides a behavior estimation device, a behavior estimation method, and a program that can accurately estimate human behavior. [Means for solving the problem]
[0007] A behavior estimation device according to one aspect of the present disclosure includes an acquisition unit that acquires sound information related to inaudible sound, which is sound in the ultrasonic band, collected by a sound collection unit; an estimation unit that inputs the sound information acquired by the acquisition unit into a trained model that indicates the relationship between the sound information and behavioral information related to human behavior and estimates the output result as the behavioral information of the person; a date and time information recording unit that records date and time information related to the date and time when the inaudible sound was collected by the sound collection unit; an adjustment unit that adjusts the sound collection frequency of the sound collection unit based on the number of times the behavioral information of the person is estimated by the estimation unit and the date and time information recorded by the date and time information recording unit; and an output unit that outputs information related to the sound collection frequency adjusted by the adjustment unit to the sound collection unit. [Effects of the Invention]
[0008] According to the present disclosure, it is possible to provide a behavior estimation device, a behavior estimation method, and a program that can accurately estimate human behavior. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of a behavior estimation system to which a behavior estimation device according to the first embodiment is applied. [Figure 2] FIG. 2 is a flowchart illustrating an example of the operation of the behavior inference device according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of sound information relating to inaudible sounds generated when a person puts on or takes off clothes. [Figure 4] FIG. 4 is a diagram showing an example of sound information regarding inaudible sounds that occur when a person walks down a corridor. [Figure 5] FIG. 5 is a diagram showing an example of sound information relating to inaudible sounds generated when water trickles from a water faucet. [Figure 6] FIG. 6 is a diagram showing an example of sound information relating to inaudible sounds generated when the skin is lightly scratched. [Figure 7] FIG. 7 is a diagram showing an example of sound information relating to inaudible sounds generated when combing hair. [Figure 8]FIG. 8 is a diagram showing an example of sound information relating to inaudible sounds generated when sniffing. [Figure 9] FIG. 9 is a diagram showing an example of sound information relating to inaudible sounds generated when a belt is put on pants. [Figure 10] FIG. 10 is a block diagram illustrating an example of a configuration of a behavior inference device according to the second embodiment. [Figure 11] FIG. 11 is a flowchart illustrating an example of the operation of the behavior inference device according to the second embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of the database. [Figure 13] FIG. 13 is a flowchart illustrating an example of the operation of the behavior inference device according to the first modification of the second embodiment. [Figure 14] FIG. 14 is a block diagram illustrating an example of a configuration of a behavior inference device according to the third embodiment. [Figure 15] FIG. 15 is a diagram illustrating an example of adjustment of the sound collection frequency by the behavior estimation device according to the third embodiment. [Figure 16] FIG. 16 is a diagram illustrating another example of adjustment of the sound collection frequency by the behavior estimation device according to the third embodiment. [Figure 17] FIG. 17 is a flowchart illustrating an example of the operation of the behavior inference device according to the third embodiment. [Figure 18] FIG. 18 is a block diagram illustrating an example of a configuration of a behavior inference device according to the fourth embodiment. [Figure 19] FIG. 19 is a diagram illustrating an example of display information. DETAILED DESCRIPTION OF THE INVENTION
[0010] (Findings that led to this disclosure) In recent years, there has been a demand for providing users with various services tailored to their lifestyles by estimating user behavior based on audible sounds collected in the user's home. For example, Patent Document 1 discloses a technology for distinguishing whether sounds collected by a microphone in an environment where television audio is detected are real-world sounds generated in response to the user's behavior in the home or television audio, and estimating the user's behavior in the home from learning results regarding the acoustic characteristics of the real-world sounds. However, the technology described in Patent Document 1 collects audible sounds perceptible to human hearing, such as television audio and real-world sounds, and estimates the user's behavior based on the collected audible sounds. Therefore, the technology is susceptible to noise from various everyday sounds and is difficult to accurately estimate human behavior. Furthermore, the technology described in Patent Document 1 collects everyday sounds in the audible range, such as the user's conversation, and transmits and receives the collected audio data, which is therefore undesirable from the perspective of protecting the user's privacy.
[0011] Therefore, the inventors of the present application have conducted extensive research in consideration of the above-mentioned problems and have found that by using inaudible sounds generated in association with user actions, it is possible to accurately estimate user actions. As a result, even when it is difficult to collect audible sounds generated in association with user actions, it is possible to efficiently collect inaudible sounds generated in association with the user actions. Furthermore, the inventors of the present application have found that it is possible to estimate user actions based on the collected inaudible sounds.
[0012] Therefore, according to the present disclosure, it is possible to provide a behavior estimation device, a behavior estimation method, and a program that can accurately estimate a user's behavior.
[0013] (Summary of the Disclosure) An outline of one aspect of the present disclosure is as follows.
[0014] A behavior estimation device according to one aspect of the present disclosure includes an acquisition unit that acquires sound information related to inaudible sounds, which are sounds in the ultrasonic band collected by a sound collection unit, and an estimation unit that inputs the sound information acquired by the acquisition unit into a trained model that indicates the relationship between the sound information and behavioral information related to human behavior, and estimates the output result as the behavioral information of the person.
[0015] As a result, even if it is difficult for the activity estimation device to collect audible sounds generated in association with human activity and estimate activity information based on the audible sounds due to the influence of various audible sounds generated around a person, i.e., sounds that become noise, by collecting inaudible sounds, the influence of sounds that become noise is reduced, thereby increasing the accuracy of sound collection. Furthermore, the activity estimation device can estimate human activity information even for actions that only generate inaudible sounds, making it possible to estimate a wider variety of activities. Therefore, the activity estimation device can accurately estimate human activity.
[0016] For example, in the behavior inference device according to an aspect of the present disclosure, the sound information input to the trained model may include at least one of a frequency band of the inaudible sound, a duration of the inaudible sound, a sound pressure of the inaudible sound, and a waveform of the inaudible sound. Furthermore, for example, the form of the sound information input to the trained model may be time-series numerical data of the inaudible sound, a spectrogram image, or a frequency characteristic image.
[0017] For example, a behavior estimation device according to one aspect of the present disclosure may further include a date and time information recording unit that records date and time information regarding the date and time when the inaudible sound was picked up by the sound collection unit, an adjustment unit that adjusts the sound collection frequency of the sound collection unit by weighting the sound collection frequency of the sound collection unit based on the number of times the behavior information of the person was estimated by the estimation unit and the date and time information recorded in the date and time information recording unit, and an output unit that outputs information regarding the sound collection frequency adjusted by the adjustment unit to the sound collection unit.
[0018] As a result, the behavior estimation device adjusts the sound collection frequency based on date and time information when inaudible sounds are collected by the sound collection unit and the estimated number of times that human behavior information is estimated by the estimation unit. Therefore, for example, rather than collecting sound at a fixed frequency, sound can be collected in accordance with a person's activity time period and activity pattern. Therefore, it is possible to efficiently collect sound and estimate human behavior while suppressing unnecessary power consumption. Furthermore, by optimizing the sound collection frequency, it is possible to suppress overheating of the sound collection unit and the behavior estimation device, thereby extending the life of the device. Furthermore, by appropriately adjusting the sound collection frequency, the load is reduced, thereby realizing faster processing.
[0019] For example, a behavior estimation device according to one aspect of the present disclosure may further include a location information acquisition unit that acquires location information regarding the location of a sound source emitting the inaudible sound, and the estimation unit may input the location information acquired by the location information acquisition unit into the trained model in addition to the sound information, and estimate the output result as the behavior information of the person.
[0020] This allows the behavior estimation device to estimate more detailed behavior that a person may take depending on the location where the sound is generated, even if the sound information has the same characteristics, thereby enabling more accurate estimation of a person's behavior.
[0021] For example, in a behavior estimation device according to one aspect of the present disclosure, the location information acquisition unit may acquire, as the location information, the location of the sound source derived based on the installation location of the sound collection unit that collected the inaudible sound.
[0022] This allows the behavior inference device to derive the installation position of the sound collection unit that collected the inaudible sound as the position of the sound source, and therefore makes it possible to easily obtain position information of the sound source.
[0023] For example, in a behavior estimation device according to one aspect of the present disclosure, the location information acquisition unit may further acquire, as the location information, the location of the sound source derived based on sound information regarding inaudible sounds emitted from an object whose installation position does not change and which is acquired by the acquisition unit.
[0024] This allows inaudible sounds emitted by objects whose installation positions do not change to be used to derive the position of the sound source, making it possible to obtain more accurate position information of the sound source.
[0025] For example, in a behavior estimation device according to one aspect of the present disclosure, the location information acquisition unit may acquire, as the location information, the position of the sound source derived from the direction of the sound source identified based on the directivity of the inaudible sound collected by two or more of the sound collection units.
[0026] This allows the behavior estimation device to identify the direction of the sound source based on the directivity of the inaudible sound collected by two or more sound collection units, thereby making it possible to acquire more detailed location information.
[0027] For example, a behavior estimation device according to one aspect of the present disclosure may further include a database in which the location information of the sound source, the sound information regarding the inaudible sound emitted from the sound source, and the behavior information of the person are stored in association with each other, and the estimation unit may further estimate the behavior information of the person by determining whether the output result of the trained model is likely based on the database.
[0028] This allows the behavior estimation device to determine whether the output result of the trained model is likely based on the database, thereby enabling more accurate estimation of human behavior.
[0029] For example, a behavior estimation device according to one aspect of the present disclosure may further include a display information generation unit that generates display information by superimposing, on layout information that indicates the arrangement of a plurality of rooms in a building in which the sound collection unit is installed and indicates in which of the plurality of rooms the sound collection unit is installed, at least one of operation information regarding the operation of the sound collection unit and the behavior information of the person estimated based on the sound information regarding the inaudible sound collected by the sound collection unit, and the output unit may further output the display information generated by the display information generation unit to an external terminal.
[0030] As a result, the behavior estimation device outputs display information to be displayed on the external terminal, so that once the behavior information is estimated, the user can check the information via the external terminal.
[0031] Furthermore, a behavior estimation method according to one aspect of the present disclosure includes an acquisition step of acquiring sound information relating to inaudible sounds, which are sounds in the ultrasonic band collected by a sound collection unit, and an estimation step of inputting the sound information acquired by the acquisition step into a trained model showing the relationship between the sound information and behavioral information relating to human behavior, and estimating the output result as the behavioral information of the person.
[0032] As a result, even when it is difficult to collect audible sounds generated in association with human behavior and estimate behavioral information based on the audible sounds due to the influence of various audible sounds generated around a person, i.e., sounds that become noise, the behavior estimation method collects inaudible sounds, thereby reducing the influence of sounds that become noise and increasing sound collection accuracy. Furthermore, the behavior estimation method can estimate human behavior information even for behaviors that only generate inaudible sounds, making it possible to estimate a wider variety of behaviors. Therefore, the behavior estimation method can accurately estimate human behavior.
[0033] A program according to an aspect of the present disclosure is a program for causing a computer to execute the behavior estimation method.
[0034] This makes it possible to achieve the same effects as the above-described behavior estimation method using a computer.
[0035] These comprehensive or specific aspects may be realized as a system, a method, an apparatus, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, a method, an apparatus, an integrated circuit, a computer program, and a recording medium.
[0036] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step sequences shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concepts are described as optional components. Furthermore, each drawing is not necessarily an exact illustration. In each drawing, substantially identical components are designated by the same reference numerals, and redundant descriptions may be omitted or simplified.
[0037] Furthermore, in this disclosure, terms indicating the relationship between elements, such as parallel and perpendicular, terms indicating the shape of elements, such as rectangle, and numerical values do not only represent the strict meaning, but also include a substantially equivalent range, for example, a difference of about a few percent.
[0038] (Embodiment 1) Hereinafter, the first embodiment will be specifically described with reference to the drawings.
[0039] [Behavior estimation system] First, the behavior estimation system will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of a behavior estimation system 400 to which a behavior estimation device 100 according to the first embodiment is applied.
[0040] The behavior estimation system 400 acquires sound information regarding inaudible sounds picked up by at least one sound pickup unit 200 installed in a specified space, inputs the acquired sound information into a trained model 130, and estimates the output result as human behavior information. The system outputs display information including the estimated behavior information to an external terminal 300.
[0041] 1, the behavior estimation system 400 includes, for example, a behavior estimation device 100, at least one sound collection unit 200, and an external terminal 300. The behavior estimation device 100 is connected to the sound collection unit 200 and the external terminal 300 via a wide area communication network 50 such as the Internet.
[0042] The behavior estimation device 100 is a device that executes a behavior estimation method including, for example, an acquisition step of acquiring sound information related to inaudible sound collected by a sound collection unit 200, and an estimation step of estimating human behavior information based on an output result obtained by inputting the sound information acquired in the acquisition step into a trained model 130 that indicates the relationship between the sound information and human behavior information. Inaudible sound is sound of a frequency that cannot be perceived by human hearing, such as sound in the ultrasonic band. Sound in the ultrasonic band is sound in a frequency band of 20 kHz or higher, for example. The trained model 130 will be described later.
[0043] The sound collection unit 200 collects inaudible sounds, which are sounds in the ultrasonic band. More specifically, the sound collection unit 200 collects inaudible sounds generated in a space in which the sound collection unit 200 is installed. For example, the sound collection unit 200 collects inaudible sounds generated by the actions of people in the space and inaudible sounds emitted from objects in the space. Examples of objects in the space include household appliances such as water faucets, showers, stoves, windows, or doors; home appliances such as washing machines, dishwashers, vacuum cleaners, air conditioners, fans, lighting, and televisions; furniture such as desks, chairs, beds, or shelves; and household items such as trash cans, storage boxes, umbrella stands, and pet supplies.
[0044] The sound collection unit 200 may be, for example, a microphone, as long as it can collect inaudible sounds. Although not shown, the sound collection unit 200 includes a communication interface, such as an adapter or a communication circuit, for wired or wireless communication, and is connected to the behavior estimation device 100 and the external terminal 300 via a wide-area communication network 50, such as the Internet. In this case, the sound collection unit 200 converts the collected inaudible sounds into electrical signals and outputs the converted electrical signals to the behavior estimation device 100. The sound collection unit 200 may be installed in any space within a building, such as a house where people live, or in a predetermined space. A space is a space partitioned by walls, windows, doors, steps, etc., such as an entrance, a hallway, a dressing room, a kitchen, a closet, and a room. One or more sound collection units 200 may be installed in one space. Note that the multiple rooms within a building may be multiple spaces within the building.
[0045] The external terminal 300 is, for example, a smartphone, a tablet terminal, a personal computer, or a home display, and has a display unit for displaying display information output from the behavior estimation device 100. The display information is generated, for example, by superimposing at least one of operation information related to the operation of the sound collection unit 200 and human behavior information estimated based on sound information related to inaudible sounds collected by the sound collection unit 200 on layout information that indicates the arrangement of multiple rooms in a building in which the sound collection unit 200 is installed and indicates in which of the multiple rooms the sound collection unit 200 is installed. Note that the behavior estimation device 100 may acquire instruction information related to an instruction input by a user to an input unit (not shown) of the external terminal 300, and generate the display information based on the acquired instruction information. The input unit is, for example, a touch panel, a keyboard, a mouse, or a microphone.
[0046] [Behavior estimation device] [1. Configuration] Next, an example of the configuration of the behavior estimation device 100 will be described with reference to Fig. 1. Here, the description of the contents described in the behavior estimation system 400 will be omitted or simplified.
[0047] 1, the behavior estimation device 100 includes, for example, an acquisition unit 110, a learning unit 120, a trained model 130, an estimation unit 140, an output unit 150, and a storage unit 160. Each component will be described below.
[0048] [Acquisition Department] The acquisition unit 110 acquires sound information related to inaudible sounds collected by the sound collection unit 200. The sound information is, for example, time-series numerical data of the inaudible sounds collected by the sound collection unit 200, and includes the frequency band, sound pressure, waveform, duration of the inaudible sounds, or the date and time when the inaudible sounds were collected. The acquisition unit 110 is, for example, an adapter for performing wired or wireless communication, or a communication interface such as a communication circuit.
[0049] [Study Department] The learning unit 120 constructs the trained model 130 through learning (e.g., machine learning). The learning unit 120 performs machine learning (so-called supervised learning) using, for example, one or more pairs of sound information about inaudible sounds collected in the past and behavioral information about human behavior corresponding to the sound information as training data. The sound information may include the frequency band of the inaudible sound and at least one of the duration, frequency, sound pressure, and waveform of the inaudible sound. The sound information may further include the time when the inaudible sound was collected. The sound information may be, for example, image data in a format such as JPEG (Joint Photographic Experts Group) or BMP (Basic Multilingual Plane), or numerical data in a format such as WAV (Waveform Audio File Format). Note that the learning by the learning unit 120 is not limited to the above-mentioned supervised learning, but may also be unsupervised learning or reinforcement learning. Reinforcement learning is a method of learning behavior that maximizes value through trial and error. For example, when performing reinforcement learning, the learning unit 120 performs learning to estimate a type of human behavior based on the context of characteristic parts (e.g., frequency distribution and signal intensity) in sound information, the time of occurrence, the duration, and the relationship with other behaviors (e.g., turning a light switch on / off). The reward is the proximity to an already estimated type of human behavior. By performing reinforcement learning, the learning unit 120 can build a trained model 130 that can estimate behaviors not present in the training data.
[0050] [Pre-trained model] The trained model 130 is obtained by learning (e.g., machine learning) by the learning unit 120. The trained model 130 is constructed by learning the relationship between sound information related to inaudible sounds and behavioral information related to human behavior. As described above, the learning method is not particularly limited and may be supervised learning, unsupervised learning, or reinforcement learning. The trained model 130 is, for example, a neural network model, more specifically, a convolutional neural network model (CNN) or a recurrent neural network (RNN). For example, when the trained model 130 is a CNN, it inputs a spectrogram image and outputs estimated human behavior information. Furthermore, when the trained model 130 is an RNN, it inputs frequency characteristics or time-series numerical data of a spectrogram and estimates user behavior.
[0051] The sound information input to the trained model 130 includes the frequency band of the inaudible sound and at least one of the duration, sound pressure, and waveform of the inaudible sound. The sound information input to the trained model 130 is in the form of time-series numerical data of the inaudible sound, a spectrogram image, or a frequency characteristic image. These data formats have been described above, so a description thereof will be omitted here.
[0052] [Estimation part] The estimation unit 140 inputs the sound information acquired by the acquisition unit 110 into a trained model 130 that indicates the relationship between the sound information and behavioral information related to human behavior, and estimates the output result as human behavioral information. In the example of FIG. 1 , the estimation unit 140 does not include the trained model 130, but the estimation unit 140 may include the trained model 130. The estimation unit 140 may store the estimated human behavioral information in the storage unit 160 or output it to the output unit 150. Note that the estimation unit 140 may store the output result when the estimation unit 140 can estimate human behavior and the input (sound information) at that time in the storage unit 160 as training data. In this case, when a predetermined amount of training data is stored, the estimation unit 140 may read the training data from the storage unit 160 and output it to the learning unit 120. The learning unit 120 may retrain the trained model 130 using the training data.
[0053] The estimation unit 140 is realized by, for example, a microcomputer or a processor.
[0054] [Output section] The output unit 150 outputs, for example, the behavior information of the person estimated by the estimation unit 140 to the external terminal 300. The output unit 150 may output the behavior information of the person to the external terminal 300, for example, based on a user instruction input to the external terminal 300. The output unit 150 is connected, for example, to the behavior estimation device 100, the sound collection unit 200, and the external terminal 300 via communication. The output unit 150 is, for example, a communication module, and may be a wireless communication circuit that performs wireless communication or a wired communication circuit that performs wired communication. There are no particular limitations on the communication standard of the communication performed by the output unit 150.
[0055] [Storage] The storage unit 160 is a storage device that stores computer programs and the like executed by the estimation unit 140. The storage unit 160 is realized by a semiconductor memory, a hard disk drive (HDD), or the like.
[0056] [2. Operation] Next, the operation of the behavior inference device 100 according to the first embodiment will be described with reference to Fig. 1 and Fig. 2. Fig. 2 is a flowchart showing an example of the operation of the behavior inference device 100 according to the first embodiment.
[0057] The acquiring unit 110 acquires sound information related to inaudible sound collected by the sound collecting unit 200 (S101). The sound collecting unit 200 is, for example, a microphone, converts the collected inaudible sound into an electrical signal, and outputs the converted electrical signal to the behavior estimation device 100. The acquiring unit 110 acquires the electrical signal of the inaudible sound collected by the sound collecting unit 200 and converts it into a digital signal using PCM (Pulse Code Modulation) or the like. Such a digital signal of the inaudible sound is also simply referred to as sound information. The digital signal of the inaudible sound is, for example, time-series numerical data of the inaudible sound. Note that the method used by the acquiring unit 110 is not limited to the above as long as it can acquire the digital signal of the inaudible sound. For example, the acquiring unit 110 may acquire the electrical signal of the sound (e.g., sound including audible sound and inaudible sound) collected by the sound collecting unit 200, convert the electrical signal into a digital signal, and then acquire the digital signal of the inaudible sound.
[0058] Next, the estimation unit 140 inputs the sound information acquired in the acquisition step S101 into the trained model 130, which indicates the relationship between sound information and behavioral information related to human behavior, and estimates the output result as human behavioral information (S102). For example, when sound information is acquired by the acquisition unit 110, the behavior estimation device 100 inputs the acquired sound information into the trained model 130. The sound information input to the trained model 130 includes, for example, the frequency band of the picked-up inaudible sound and at least one of the duration, sound pressure, and waveform of the inaudible sound. The form of the sound information input to the trained model 130, i.e., the data format of the sound information, may be time-series numerical data of the picked-up inaudible sound, a spectrogram image, or a frequency characteristic image.
[0059] Although not shown, the estimation unit 140 may output the estimated behavioral information of the person to the output unit 150. At this time, the estimation unit 140 may store the estimated behavioral information of the person in the storage unit 160. For example, the estimation unit 140 may associate the estimated behavioral information with the sound information acquired by the acquisition unit 110 and store the information in the storage unit 160.
[0060] The behavior inference device 100 repeatedly executes the above processing flow every time the acquisition unit 110 acquires sound information.
[0061] [3. Specific examples of behavioral inference] Hereinafter, human behavior information estimated by the behavior estimation device 100 according to the first embodiment will be described with reference to Fig. 3 to Fig. 9. Each of Fig. 3 to Fig. 9 shows an example of sound information input to the trained model 130. In each of Figs. 3 to 9, (a) shows an image of a spectrogram, and (b) shows an image of frequency characteristics.
[0062] The spectrogram shown in (a) is an image in which the horizontal axis represents time (seconds) and the vertical axis represents frequency (Hz), and the time variation of the signal strength of the frequency characteristics is displayed in shades of gray. Here, the higher the whiteness in (a), the stronger the signal strength of the frequency characteristics.
[0063] The frequency characteristics shown in (b) are obtained by Fourier transforming the time series numerical data of the inaudible sound.
[0064] 3 to 9 also include audible sounds collected by the sound collection unit 200. However, as described above, it is difficult to collect audible sounds and they are easily affected by noise. Therefore, even if the collected sounds include audible sounds, the behavior estimation device 100 estimates a person's behavior based on sound information related to inaudible sounds of 20 kHz or higher.
[0065] [Example 1] 3 is a diagram showing an example of sound information relating to inaudible sounds generated when a person puts on or takes off clothes, where the material of the clothes is cotton.
[0066] In the spectrogram image of Figure 3(a), five characteristic signals (1) to (5) are detected in the frequency band above 20 kHz. Signals (1) and (2) exceed 80 kHz, signals (3) and (4) are just under 80 kHz, and signal (5) is just under 70 kHz. The signal strength is particularly strong below 50 kHz. These signals correspond to the sound of clothes rubbing together when a person puts on or takes off clothes.
[0067] In the image of frequency characteristics in FIG. 3(b), the signal strength of frequency components in the frequency band between 20 kHz and 50 kHz is large among frequency components of 20 kHz or more.
[0068] When (a) or (b) in FIG. 3 is input to the trained model 130, the behavioral information of the person that is output is, for example, "undressing" or "changing clothes."
[0069] [Example 2] Fig. 4 shows an example of sound information related to inaudible sounds that occur when a person walks down a corridor. Specifically, the inaudible sounds are generated when a person walks down a corridor in slippers after walking down the corridor barefoot.
[0070] In the spectrogram image of Figure 4(a), a signal corresponding to the sound of feet rubbing against the hallway when a person walks barefoot down the hallway is detected between 0 and 8 seconds, and a signal corresponding to the sound of slippers rubbing against the hallway when a person walks down the hallway in slippers is detected between 8 and 10 seconds. For example, when a person walks down the hallway barefoot, multiple characteristic signals are detected in the frequency band from 20 kHz to 50 kHz, in particular from 20 kHz to 35 kHz. Furthermore, when a person walks down the hallway in slippers, multiple characteristic signals are detected in the frequency band from 20 kHz to 70 kHz, in particular from 20 kHz to 40 kHz.
[0071] In the image of frequency characteristics in FIG. 4(b), the signal strength of frequency components in the frequency band from 20 kHz to 40 kHz is large among frequency components of 20 kHz or higher.
[0072] When (a) or (b) in FIG. 4 is input to the trained model 130, the human behavior information that is output is, for example, "walking."
[0073] [Example 3] FIG. 5 is a diagram showing an example of sound information relating to inaudible sounds generated when water trickles from a water faucet.
[0074] In the spectrogram image of Figure 5(a), a signal corresponding to the sound of running water is detected between 0 and 6 seconds. Continuous signals are detected from around 20 kHz to 35 kHz, and multiple signals above 40 kHz are detected between the continuous signals.
[0075] In the image of frequency characteristics in FIG. 5(b), the signal strength of frequency components in the frequency band from around 20 kHz to 35 kHz is also large for frequency components above 20 kHz.
[0076] When (a) or (b) in FIG. 5 is input to the trained model 130, the human behavior information that is output is, for example, "washing hands."
[0077] [Example 4] FIG. 6 is a diagram showing an example of sound information relating to inaudible sounds generated when the skin is lightly scratched.
[0078] In the spectrogram image of FIG. 6(a), characteristic signals are detected in a wide frequency band from 20 kHz to 90 kHz. The detected signals correspond to the sound of skin rubbing against skin when a person lightly scratches their skin, and characteristic signals are detected particularly between 3 and 4 seconds and around 8 seconds. The multiple signals detected between 3 and 4 seconds are signals in the frequency band from 20 kHz to 40 kHz, and the multiple signals detected around 8 seconds are signals in the frequency band from 20 kHz to 90 kHz.
[0079] In the image of frequency characteristics in FIG. 6(b), among frequency components above 20 kHz, the signal strength of frequency components in the frequency band from 20 kHz to 40 kHz and frequency components from 40 kHz to 90 kHz is large.
[0080] When (a) or (b) in FIG. 6 is input to the trained model 130, the human behavior information that is output is, for example, "scratching an itch."
[0081] [Example 5] FIG. 7 is a diagram showing an example of sound information relating to inaudible sounds generated when combing hair.
[0082] In the spectrogram image of Figure 7(a), characteristic signals are detected in the frequency band from 20 kHz to 60 kHz.
[0083] In the image of frequency characteristics in FIG. 7(b), the signal strength of frequency components in the frequency band from 20 kHz to 50 kHz is large among frequency components of 20 kHz or higher.
[0084] When (a) or (b) in FIG. 7 is input to the trained model 130, the output human behavior information is, for example, "combing hair."
[0085] Next, the sixth and seventh examples will be described.
[0086] [Example 6] FIG. 8 is a diagram showing an example of sound information relating to inaudible sounds generated when sniffing.
[0087] In the spectrogram image of Figure 8(a), a signal corresponding to the sound of sniffing through the nose when sniffing with a tissue stuffed up the nose is detected at around 2 seconds, and a signal corresponding to the sound of sniffing through the nose when sniffing with nothing stuffed up the nose is detected at around 4 seconds. When sniffing with a tissue stuffed up the nose, characteristic signals are detected in the frequency band from around 20 kHz to around 35 kHz, particularly in the frequency band from 30 kHz to 35 kHz. When sniffing with nothing stuffed up the nose, characteristic signals are detected in a wide frequency band from 20 kHz to over 90 kHz. The signal strength is particularly strong in the frequency band from 20 kHz to 35 kHz.
[0088] In the image of frequency characteristics in FIG. 8(b), the signal strength of frequency components in the frequency band from around 30 kHz to 35 kHz is also large in frequency components above 20 kHz.
[0089] When (a) or (b) in FIG. 8 is input to the trained model 130, the human behavior information that is output is, for example, "sniffle" or "blow your nose."
[0090] [Example 7] FIG. 9 is a diagram showing an example of sound information relating to inaudible sounds generated when a belt is put on pants.
[0091] In the spectrogram image of Figure 9(a), many small signals corresponding to the sound of the belt rubbing against clothing when threading the belt through pants are detected in the frequency band between 20 kHz and 60 kHz.
[0092] In the image of frequency characteristics in FIG. 9(b), among frequency components above 20 kHz, the signal strength of frequency components in the frequency band from about 20 kHz to 40 kHz and the frequency band from 40 kHz to 60 kHz is large.
[0093] When (a) or (b) in FIG. 9 is input to the trained model 130, the behavioral information of the person that is output is, for example, "changing clothes."
[0094] Note that the above first to seventh examples are examples in which sounds are barely perceptible to the human ear, or are only slightly perceptible, but it is difficult to collect audible sounds and estimate behavior based on the audible sounds, but it is possible to collect inaudible sounds and estimate behavior based on the inaudible sounds. As described above, according to the behavior estimation device 100, even in cases in which it is difficult to collect audible sounds generated in conjunction with human behavior and estimate behavior based on the audible sounds, it is possible to estimate human behavior based on inaudible sounds generated in conjunction with human behavior.
[0095] [Other examples] Human behavior estimated based on collected inaudible sounds is not limited to the above examples. For example, (1) human behavior information estimated based on the sound of toilet paper rubbing against paper when pulling out the toilet paper and the sound of the toilet paper roll hitting the toilet paper holder core is, for example, "using the toilet." Also, (2) human behavior information estimated based on inaudible sounds generated by opening and closing a window is, for example, "ventilation." Also, (3) human behavior information estimated based on inaudible sounds generated by opening and closing a sliding door is, for example, "entering and leaving the room." Also, (4) human behavior information estimated based on inaudible sounds generated when opening and closing a shelf or desk drawer or a small door with a magnet is, for example, "taking out and putting away dishes" if the sound is emitted from a cupboard, and "learning" if the sound is emitted from a desk. (5) Human behavior information estimated based on inaudible sounds generated when adjusting lighting dimming is, for example, "falling asleep," "waking up," or "entering or leaving the room." (6) Human behavior information estimated based on inaudible sounds generated when moving a comforter or the sound of the comforter rubbing against clothing is, for example, "going to bed," "sleeping," "waking up," "nap," or "turning over." (7) Human behavior information estimated based on inaudible sounds generated when pouring liquid into a glass is, for example, "drinking a beverage."
[0096] [4. Effects, etc.] As described above, the behavior estimation device 100 includes an acquisition unit 110 that acquires sound information related to inaudible sounds, which are sounds in the ultrasonic band collected by the sound collection unit 200, and an estimation unit 140 that inputs the sound information acquired by the acquisition unit 110 into a trained model 130 that indicates the relationship between sound information and behavior information related to human behavior and estimates the output result as human behavior information.
[0097] Even when it is difficult to collect audible sounds generated in association with human behavior and estimate behavior information based on the audible sounds due to the influence of various audible sounds generated around a person, i.e., sounds that become noise, the behavior estimation device 100 collects inaudible sounds, thereby reducing the influence of sounds that become noise and increasing sound collection accuracy. Furthermore, the behavior estimation device 100 can estimate human behavior information even for behaviors that only generate inaudible sounds, making it possible to estimate a wider variety of behaviors. Therefore, the behavior estimation device 100 can accurately estimate human behavior.
[0098] Furthermore, in conventional technologies, audible sounds in a user's home are collected to estimate the user's behavior, and therefore, for example, voice data such as conversations are also collected, which may result in a loss of user privacy. However, the user behavior estimation device 100 collects inaudible sounds to estimate human behavior, and therefore can protect human privacy.
[0099] Therefore, the behavior estimation device 100 can accurately and appropriately estimate a person's behavior.
[0100] The behavior estimation device 100 does not use an active method of irradiating a person with ultrasound and estimating the person's behavior based on the reflected waves, but uses a passive method of estimating the person's behavior based on ultrasound generated in association with the person's behavior, and therefore does not need to include an ultrasound irradiation unit. Therefore, it is possible to accurately estimate the person's behavior with a smaller configuration than a configuration including an ultrasound irradiation unit.
[0101] (Embodiment 2) Next, a behavior estimation device according to embodiment 2 will be described. In embodiment 1, sound information of inaudible sounds collected by the sound collection unit 200 is input to the trained model 130, and the output result obtained is estimated as human behavior information. In embodiment 2, in addition to the above sound information, position information regarding the position of a sound source emitting inaudible sounds is input to the trained model 130, and the output result obtained is estimated as human behavior information. This embodiment differs from embodiment 1 in that the following mainly describes the differences from embodiment 1. Note that explanations of content that overlaps with embodiment 1 will be omitted or simplified.
[0102] [1. Configuration] 10 is a block diagram showing an example of the configuration of a behavior inference device 100a according to Embodiment 2. The behavior inference device 100a according to Embodiment 2 includes, for example, an acquisition unit 110, a learning unit 120a, a trained model 130a, an estimation unit 140a, an output unit 150, a storage unit 160a, and a position information acquisition unit 170.
[0103] In the second embodiment, the learning unit 120a, the trained model 130a, and the estimation unit 140a differ from the first embodiment in that they use position information in addition to sound information of inaudible sounds, and the storage unit 160a differs from the first embodiment in that it stores position information acquired by the position information acquisition unit 170. In particular, the second embodiment differs from the first embodiment in that it includes the position information acquisition unit 170.
[0104] [Location information acquisition unit] The position information acquisition unit 170 acquires position information relating to the position of the sound source that emits the inaudible sound collected by the sound collection unit 200. Acquiring the position information of the sound source does not simply involve acquiring transmitted position information, but also involves deriving (or specifying) the position of the sound source. The sound source refers to the source of the inaudible sound that is generated in conjunction with human activity.
[0105] For example, the position information acquiring unit 170 may acquire, as position information, the position of a sound source derived based on the installation position of the sound collection unit 200 that collected the inaudible sound. In this case, the position information acquiring unit 170, for example, identifies the space in which the sound collection unit 200 that collected the inaudible sound is installed as the position of the sound source, that is, the location where the sound source exists, and acquires the space as position information related to the position of the sound source. As described above, a space is a space partitioned by walls, windows, doors, steps, or the like, and is, for example, a hallway, an entrance, a dressing room, a kitchen, a room, or a closet. For example, if the sound collection unit 200 that collected the inaudible sound generated by a person's behavior is a dressing room, the position information acquiring unit 170 acquires the dressing room as position information related to the position of the sound source. In this case, for example, if the sound collection unit 200 collected the inaudible sound generated by putting on or taking off clothes, the person's behavior information estimated based on the sound information and position information is "bathing." Furthermore, for example, when the sound collection unit 200 that collected the inaudible sound generated by a person's actions is a closet, the position information acquisition unit 170 acquires the closet as position information relating to the position of the sound source. In this case, for example, when the sound collection unit 200 collects the inaudible sound generated by putting on or taking off clothes, the person's behavior information estimated based on the sound information and the position information is "changing clothes."
[0106] Furthermore, for example, the position information acquisition unit 170 may further acquire, as position information, the position of a sound source derived based on sound information regarding inaudible sound emitted from an object whose installation position does not change, acquired by the acquisition unit 110. In this case, for example, when the position information acquisition unit 170 determines that the sound information acquired by the acquisition unit 110 includes sound information regarding inaudible sound emitted from an object whose installation position does not change, the position information acquisition unit 170 acquires the space in which the object is installed as the location of the sound source, i.e., the location of the sound source. The fact that the installation position of an object does not change may mean that the installation position of the object in a predetermined space does not change, or may mean that the space in which the object is installed does not change. For example, if a dishwasher is installed in a kitchen, and the installation position of the dishwasher in the kitchen changes, the installation position of the dishwasher does not change to spaces other than the kitchen. In this way, an object whose installation position does not change is not limited to a dishwasher, but may also be a washing machine, a shower, a water faucet, a television, or the like. For example, if the object whose installation position does not change is a washing machine, and the position information acquisition unit 170 determines that the sound information related to the inaudible sound collected by the sound collection unit 200 includes sound information related to the inaudible sound emitted by the washing machine, the position information acquisition unit 170 acquires the space in which the washing machine is installed, i.e., the changing room, as the position information related to the position of the sound source. In this case, for example, if the sound collection unit 200 collects inaudible sound generated by putting on or taking off clothes, the human behavior information estimated based on the sound information and the position information is "bathing." Furthermore, for example, if the object whose installation position does not change is a television, and the position information acquisition unit 170 determines that the sound information related to the inaudible sound collected by the sound collection unit 200 includes sound information related to the inaudible sound emitted by the television, the position information acquisition unit 170 acquires the space in which the television is installed, i.e., the living room, as the position information related to the position of the sound source. In this case, for example, if the sound collection unit 200 collects inaudible sound generated by putting on or taking off clothes, the human behavior information estimated based on the sound information and the position information is "changing clothes." Changing clothes also includes taking off a coat or other outer garment, or putting on a coat or other outer garment.
[0107] Furthermore, for example, the position information acquisition unit 170 may acquire, as position information, the position of a sound source derived from the direction of the sound source identified based on the directivity of inaudible sounds collected by two or more sound collection units 200. Two or more sound collection units 200 may be installed in a single space, or two or more sound collection units 200 may be installed individually in different spaces. For example, when two or more sound collection units 200 are installed in a single space, the position of the sound source in the space can be identified based on the directivity of inaudible sounds collected by these sound collection units 200. For example, when two or more sound collection units 200 are installed in a room equipped with a closet, when a sound collection unit 200 collects inaudible sounds corresponding to putting on or taking off clothes, the position information acquisition unit 170 identifies the direction of the sound source as the location of the closet based on the directivity of the collected inaudible sounds. In other words, the position information acquisition unit 170 acquires the closet as position information of the sound source based on the directivity of the collected inaudible sounds. In this case, the behavior information of the person estimated based on the sound information and the position information is "changing clothes." Furthermore, for example, if two sound collection units 200 are installed in a changing room and a hallway, respectively, and these sound collection units 200 collect inaudible sounds corresponding to putting on or taking off clothes, the position information acquisition unit 170 identifies the direction of the sound source as the changing room based on the directivity of the collected inaudible sounds. In other words, the position information of the sound source acquired by the position information acquisition unit 170 is the changing room. In this case, the behavior information of the person estimated based on the sound information and the position information is "taking a bath."
[0108] As described above, the behavior estimation device 100a according to the second embodiment can estimate a person's behavior based on sound information of inaudible sounds generated in association with the person's behavior and location information of the sound source of the inaudible sounds, and therefore can estimate a person's behavior with higher accuracy.
[0109] [2. Operation] Next, the operation of the behavior inference device 100a will be described with reference to Fig. 10 and Fig. 11. Fig. 11 is a flowchart showing an example of the operation of the behavior inference device 100a according to the second embodiment.
[0110] The acquisition unit 110 acquires sound information relating to inaudible sound collected by the sound collection unit 200 (see FIG. 1) (S201). Step S201 is the same as step S101 in FIG.
[0111] Next, the position information acquisition unit 170 acquires position information relating to the position of the sound source emitting the inaudible sound collected by the sound collection unit 200 (S202). As described above, the position information acquisition unit 170 may acquire, as the position information, the position of the sound source derived based on the installation position of the sound collection unit 200. Alternatively, the position information acquisition unit 170 may acquire, as the position information, the position of the sound source derived based on sound information relating to inaudible sound emitted from an object whose installation position does not change. Alternatively, the position information acquisition unit 170 may acquire, as the position information, the position of the sound source derived from the direction of the sound source identified based on the directivities of the inaudible sounds collected by two or more sound collection units 200.
[0112] Next, the estimation unit 140a inputs the sound information acquired in step S201 and the sound source position information acquired in step S202 into a trained model 130a indicating the relationship between the sound information, the sound source position information, and behavioral information related to human behavior, and estimates the output result as human behavioral information (S203). For example, when the acquisition unit 110 acquires sound information and the sound source position information acquires sound source position information by the position information acquisition unit 170, the behavior estimation device 100a inputs the acquired sound information and sound source position information into the trained model 130a. The sound information and the form of the sound information input into the trained model 130a, i.e., the data format of the sound information, are the same as those described in the first embodiment, and therefore will not be described here. The trained model 130a is constructed by machine learning in the learning unit 120a using, for example, one or more pairs of sound information, sound source position information, and behavioral information related to human behavior corresponding to the sound information and sound source position information as training data.
[0113] Although not shown, the estimation unit 140a may output the estimated behavioral information of the person to the output unit 150. At this time, the estimation unit 140a may store the estimated behavioral information of the person in the storage unit 160a. For example, the estimation unit 140a may associate the estimated behavioral information with the sound information acquired by the acquisition unit 110 and the position information of the sound source acquired by the position information acquisition unit 170, and store the associated information in the storage unit 160a.
[0114] The behavior inference device 100a repeatedly executes the above processing flow every time the acquisition unit 110 acquires sound information.
[0115] [3. Specific examples of behavioral inference] Hereinafter, human behavior information estimated by the behavior estimation device 100a according to the second embodiment will be described with reference to FIGS. 3 and 5 again.
[0116] [Example 1] 3 shows an example of sound information related to inaudible sounds generated when a person puts on or takes off clothes. In the first embodiment, the sound information shown in FIG. 3 is input to the trained model 130, and the output result obtained is, for example, "undressing" or "changing clothes."
[0117] In the second embodiment, in addition to the sound information shown in FIG. 3, the position information of the sound source is input to the trained model 130a. For example, if the position information of the sound source is a changing room, the human behavior information output from the trained model 130a is, for example, "bathing" or "undressing." Also, for example, if the position information of the sound source is a living room or a closet, the human behavior information output from the trained model 130a is "changing clothes." Also, for example, if the position information of the sound source is a bedroom or a bed, the human behavior information output from the trained model 130a is sleep-related behavior such as "going to bed," "waking up," "turning over," "going to sleep," or "taking a nap."
[0118] [Example 2] 5 shows an example of sound information related to the inaudible sound generated when water is trickling from a water faucet. In the first embodiment, the sound information shown in FIG. 5 is input to the trained model 130, and the output result obtained is, for example, "washing hands."
[0119] In the second embodiment, in addition to the sound information shown in Fig. 5, the location information of the sound source is input to the trained model 130a. For example, if the location information of the sound source is a bathroom, the human behavior information output from the trained model 130a is, for example, "washing hands," "brushing teeth," or "washing face."
[0120] [Example 3] For example, when the acquisition unit 110 acquires the sound information shown in Figure 3 and the sound information shown in Figure 5, and the location information acquisition unit 170 acquires the bathroom (sound of running water) and the changing room (taking off and putting on clothes) as location information of the sound source based on sound information regarding inaudible sounds emitted from an object whose installation position does not change (here, the water faucet shown in Figure 5), the human behavior information output from the trained model 130a is, for example, "bathing."
[0121] [4. Effects, etc.] As described above, the behavior estimation device 100a further includes a position information acquisition unit 170 that acquires position information regarding the position of a sound source emitting inaudible sound, and the estimation unit 140a inputs the position information acquired by the position information acquisition unit 170, in addition to the sound information acquired by the acquisition unit 110, into the trained model 130a, and estimates the output result as human behavior information.
[0122] Such a behavior estimation device 100a can estimate more detailed behavior that a person may take depending on the location where the sound is generated, even if the sound information has the same characteristics, and therefore can estimate a person's behavior with higher accuracy.
[0123] (Modification 1 of Embodiment 2) Next, a first modification of the second embodiment will be described. In the second embodiment, sound information and sound source position information are input to the trained model 130a, and the output result obtained is estimated as human behavior information. However, in the first modification of the second embodiment, human behavior information is estimated by determining whether the output result of the trained model 130a is likely based on a database, which is different from the second embodiment. The following mainly describes the differences from the second embodiment. Note that explanations of content that overlaps with the first and second embodiments will be omitted or simplified.
[0124] [1. Configuration] Here, only the configuration different from that of Embodiment 2 will be described. Referring again to Fig. 10, the behavior estimation device 100a according to Variation 1 of Embodiment 2 differs from Embodiment 2 in that it further includes a database 162 in which position information of a sound source, sound information related to inaudible sounds emitted from the sound source, and behavior information of a person are stored in association with each other.
[0125] Fig. 12 is a diagram showing an example of the database 162. As shown in Fig. 12, the database 162 stores sound information having the same characteristics, but linked to different behavioral information depending on the location information of the sound source. The database 162 is used by the estimation unit 140a when determining whether the output result of the trained model 130a is likely or not.
[0126] [2. Operation] Next, the operation of the behavior inference device 100a according to Modification 1 of Embodiment 2 will be described with reference to Fig. 10 and Fig. 13. Fig. 13 is a flowchart showing an example of the operation of the behavior inference device 100a according to Modification 1 of Embodiment 2. In Fig. 13, step S201 and step S202 in Fig. 11 are described as one step S301.
[0127] First, the acquisition unit 110 acquires sound information relating to the inaudible sound collected by the sound collection unit 200. Then, the position information acquisition unit 170 acquires position information relating to the position of the sound source that emits the inaudible sound (S301).
[0128] Next, the estimation unit 140a inputs the sound information and position information acquired in step S301 into the trained model 130a and acquires an output result (S302).
[0129] Next, the estimation unit 140a determines whether the output result of the trained model 130a is likely based on the database 162 (S303). In step S303, the determination of whether the output result is likely is made based on whether a pair of the sound information and position information input to the trained model 130a and the behavioral information that is the output result exists in the database 162. If the estimation unit 140a determines that the output result of the trained model 130a is likely (Yes in S303), it estimates the output result as human behavioral information (S304). On the other hand, if the estimation unit 140a determines that the output result of the trained model 130a is not likely (No in S303), it stores the determination result in the storage unit 160a (S305). At this time, the estimation unit 140a may associate the sound information and position information input to the trained model 130a with the output result and the determination result and store them in the storage unit 160a. The learning unit 120a may, for example, re-learn the trained model 130a using the stored information.
[0130] [3. Specific examples of behavioral inference] Hereinafter, human behavior information estimated by the behavior estimation device 100a according to the first modification of the second embodiment will be described with reference to FIGS. 3, 5, and 12 again.
[0131] [Example 1] FIG. 3 shows an example of sound information relating to inaudible sounds generated when a person puts on or takes off clothes, that is, sound information of the rustling sound of fabrics shown in FIG.
[0132] In the second embodiment, the output result obtained by inputting the location information of the sound source (e.g., a changing room) into the trained model 130a in addition to the sound information shown in Fig. 3 is, for example, "bathing" or "undressing." In the first variation of the second embodiment, by determining whether the output result is likely or not using the database 162, it is determined that the output result is likely to be "undressing," and "undressing" is estimated as the behavior information.
[0133] [Example 2] FIG. 5 shows an example of sound information relating to the inaudible sound generated when water trickles from a water faucet, that is, the sound information of the sound of running water shown in FIG.
[0134] In the second embodiment, the output result obtained by inputting the location information of the sound source (e.g., the bathroom) in addition to the sound information shown in Fig. 5 into the trained model 130a is, for example, "hand washing," "tooth brushing," or "face washing." In the first modification of the second embodiment, by determining whether the output result is likely or not using the database 162, it is determined that the output result is likely to be "hand washing," and "hand washing" is estimated as the behavior information.
[0135] [4. Effects, etc.] As described above, the behavior estimation device 100a further includes a database 162 in which location information of a sound source, sound information related to inaudible sounds emitted from the sound source, and human behavior information are stored in association with each other, and the estimation unit 140a further estimates the human behavior information by determining whether the output result of the trained model 130a is likely based on the database 162.
[0136] The behavior estimation device 100a as described above determines whether the output result of the trained model 130a is likely based on the database 162, and therefore can estimate human behavior with higher accuracy.
[0137] (Embodiment 3) Next, a behavior inference device 100b according to Embodiment 3 will be described. The behavior inference device 100b according to Embodiment 3 differs from Embodiments 1, 2, and Modification 1 of Embodiment 2 in that the behavior inference device 100b according to Embodiment 3 adjusts the frequency of sound collection by the sound collection unit 200 between time periods when a person is active and time periods when a person is not active. The following description will focus on the differences from the above embodiments. Note that description of content that overlaps with the above embodiments will be omitted or simplified.
[0138] [1. Configuration] 14 is a block diagram showing an example of the configuration of a behavior inference device 100b according to Embodiment 3. The behavior inference device 100b according to Embodiment 3 includes, for example, an acquisition unit 110, a learning unit 120a, a trained model 130a, an estimation unit 140a, an output unit 150, a storage unit 160b, a position information acquisition unit 170, and an adjustment unit 180.
[0139] The third embodiment differs from the above-described embodiments in that it includes a date and time information recording unit 164 and an adjustment unit 180.
[0140] [Date and time information recording section] The date and time information recording unit 164 records date and time information relating to the date and time when the inaudible sound was collected by the sound collection unit 200. The date and time information recording unit 164 may, for example, record the date and time information in association with sound information relating to the inaudible sound collected by the sound collection unit 200. In the example of Fig. 14, the date and time information recording unit 164 is stored in the storage unit 160b, but it may also be a recording device provided separately from the storage unit 160b.
[0141] [Adjustment section] The adjustment unit 180 adjusts the sound collection frequency of the sound collection unit 200 by weighting the sound collection frequency of the sound collection unit 200 based on the number of times that the human behavior information is estimated by the estimation unit 140a and the date and time information recorded in the date and time information recording unit 164. For example, the adjustment unit 180 may adjust the sound collection frequency using a predetermined arithmetic expression. The sound collection frequency may be adjusted every predetermined period, such as every week, every month, or every three months. Hereinafter, the adjustment of the sound collection frequency using the arithmetic expression will be specifically described with reference to FIG. 15. FIG. 15 is a diagram showing an example of the adjustment of the sound collection frequency by the behavior estimation device 100b according to Embodiment 3. Hereinafter, the adjustment may also be referred to as optimization.
[0142] As shown in Fig. 15, the sound collection unit 200 is installed in, for example, the living room, kitchen, and bathroom of a house. The number of behavioral estimations is, for example, the average number of behavioral estimations in each time slot from midnight to 11pm over a certain period of time in the past (for example, one week). In the example of Fig. 15, (A) the number of behavioral estimations is the average number of behavioral estimations performed over one week when sound pressure of -40 dB or more was detected when the sound collection unit 200 performed measurements for one minute at six-minute intervals in each time slot ((B) in the figure, sound collection frequency before optimization).
[0143] The sound collection frequency after optimization (C1) in the figure is derived using the following equation (1).
[0144] Sound collection frequency after optimization = number of estimated actions / sound collection frequency before optimization × 10 + 3 (1)
[0145] Here, the sound collection frequency is the number of times sound is collected per hour.
[0146] After adjusting the sound collection frequency of each sound collection section 200, the adjustment section 180 outputs information about the adjusted sound collection frequency, in other words, the optimized sound collection frequency, to the output section 150. The information about the sound collection frequency may be, for example, information about the time when the sound collection section 200 collects sound.
[0147] Furthermore, the adjustment unit 180 may adjust the sound collection frequency using, for example, a neural network model (not shown) constructed by machine learning. The neural network model may be, for example, a multi-layer neural network model that indicates the relationship between the number of estimated behaviors before optimization and the sound collection frequency after optimization in each time period. The machine learning may be supervised learning, unsupervised learning, or reinforcement learning. For example, when supervised learning is performed, training data may be created for each space in which the sound collection unit 200 is installed. Furthermore, the algorithm used in reinforcement learning may be, for example, a Deep Q Network.
[0148] Adjustment of the sound collection frequency using a neural network model will be specifically described below with reference to Fig. 16. Fig. 16 is a diagram showing another example of adjustment of the sound collection frequency by the behavior estimation device 100b according to the third embodiment.
[0149] The input of the neural network model is, for example, the time periods and (A) the estimated number of actions in each time period in Fig. 16. The output of the neural network model is the adjusted sound collection frequency, for example, (C2) the optimized sound collection frequency in Fig. 16.
[0150] In the example of Figure 16, (C2) the sound collection frequency after optimization is adjusted so that the sum does not exceed 30 by inserting a softmax function into the output only when the sum of the outputs of the neural network model exceeds 30.
[0151] [Output section] In the third embodiment, the output unit 150 outputs information relating to the sound collection frequency adjusted by the adjustment unit 180 to the sound collection unit 200. As described in the first embodiment, the output unit 150 is connected to the sound collection unit 200 via the wide area communication network 50. The output unit 150 is a communication module for communicating with the sound collection unit 200 and the external terminal 300. The communication may be wireless communication or wired communication. There is no particular limitation on the communication standard used for the communication.
[0152] [2. Operation] Next, the operation of the behavior estimation device 100b according to Embodiment 3 will be described with reference to Fig. 14 and Fig. 17. Fig. 17 is a flowchart showing an example of the operation of the behavior estimation device 100b according to Embodiment 3. Here, the flow of adjusting the sound collection frequency will be described.
[0153] First, the adjustment unit 180 determines whether or not a predetermined period has elapsed (S401). If the adjustment unit 180 determines that the predetermined period has not elapsed (No in S401), the adjustment unit 180 returns to the processing of step S401.
[0154] On the other hand, when it is determined that the predetermined period has elapsed (Yes in S401), the adjustment unit 180 acquires the number of times that the estimation unit 140a estimated the person's behavior information during the predetermined period (so-called the number of estimated behaviors) and date and time information relating to the date and time when the inaudible sound was collected by the sound collection unit 200 (S402). For example, the adjustment unit 180 may read the number of estimations performed by the estimation unit 140a from the storage unit 160b and read the date and time information from the date and time information recording unit 164, or the date and time information and the number of estimated behaviors during the predetermined period may be recorded in the date and time information recording unit 164, and this information may be read from the date and time information recording unit 164.
[0155] Next, the adjustment unit 180 adjusts the sound collection frequency by weighting the sound collection frequency of the sound collection unit 200 based on the acquired estimated number of behaviors and date and time information (S403). As described above, the adjustment unit 180 may adjust the sound collection frequency using an arithmetic expression or a neural network model.
[0156] Subsequently, the adjustment unit 180 outputs information relating to the adjusted sound collection frequency to the output unit 150 (not shown). The output unit 150 outputs the acquired information relating to the sound collection frequency to the sound collection unit 200 (S404).
[0157] [3. Effects, etc.] As described above, the behavior estimation device 100b further includes a date and time information recording unit 164 that records date and time information related to the date and time when inaudible sound was collected by the sound collection unit 200, an adjustment unit 180 that adjusts the sound collection frequency of the sound collection unit 200 by weighting the sound collection frequency of the sound collection unit 200 based on the number of times that human behavior information was estimated by the estimation unit 140a and the date and time information recorded in the date and time information recording unit 164, and an output unit 150 that outputs information related to the sound collection frequency adjusted by the adjustment unit 180 to the sound collection unit 200.
[0158] The behavior estimation device 100b adjusts the sound collection frequency based on the date and time information when the inaudible sound is collected by the sound collection unit 200 and the estimated number of times the human behavior information is estimated by the estimation unit 140a. Therefore, for example, rather than collecting sound at a fixed frequency, sound can be collected according to the time period and activity pattern of the human activity. Therefore, it is possible to efficiently collect sound and estimate the human behavior while suppressing unnecessary power consumption. Furthermore, optimizing the sound collection frequency can suppress the temperature rise of the sound collection unit 200 and the behavior estimation device 100b, thereby extending the life of the device. Furthermore, by appropriately adjusting the sound collection frequency, the load is reduced, thereby realizing faster processing.
[0159] (Fourth embodiment) Next, a behavior inference device 100c according to Embodiment 4 will be described. Embodiment 4 differs from the above-described embodiments and modifications in that the behavior inference device 100c creates display information including acquired information and derived information and outputs the display information to the external terminal 300. The following description will focus on the differences from Embodiment 3.
[0160] [1. Configuration] Fig. 18 is a block diagram showing an example of the configuration of a behavior inference device 100c according to Embodiment 4. As shown in Fig. 18, the behavior inference device 100c according to Embodiment 4 differs from Embodiment 3 in that it includes a display information generation unit 190.
[0161] [Display information generation section] The display information generating unit 190 generates display information by superimposing, on layout information indicating the arrangement of a plurality of rooms in a building in which the sound collection unit 200 is installed and indicating in which of the plurality of rooms the sound collection unit 200 is installed, at least one of operation information related to the operation of the sound collection unit 200 and human behavior information estimated based on sound information related to inaudible sounds collected by the sound collection unit 200. Furthermore, the display information generating unit 190 may change the information and display format included in the display information, for example, based on instruction information input by the user to the external terminal 300. For example, when the adjustment unit 180 adjusts the sound collection frequency of the sound collection unit 200, the display information generating unit 190 may generate display information that displays the sound collection efficiency at the sound collection frequency before the adjustment and a predicted value of the sound collection efficiency at the sound collection frequency after the adjustment.
[0162] [Output section] In the fourth embodiment, the output unit 150 further outputs the display information generated by the display information generating unit 190 to the external terminal 300.
[0163] [2. Operation] Next, the operation of the behavior inference device 100c according to the fourth embodiment will be described with reference to Fig. 18 and Fig. 19. Here, an example of the operation of creating display information and outputting the display information will be described. Fig. 19 is a diagram showing an example of display information.
[0164] When the estimation unit 140a estimates a person's behavior, the display information generation unit 190 acquires the estimated behavior information. The display information generation unit 190 then generates display information by superimposing, on layout information indicating the layout of multiple rooms in a building in which the sound collection unit 200 is installed and indicating in which of the multiple rooms the sound collection unit 200 is installed, at least one of operation information regarding the operation of the sound collection unit 200 and human behavior information estimated based on sound information regarding inaudible sounds collected by the sound collection unit 200. As shown in FIG. 19 , when an x is displayed next to a speaker mark, this indicates that the sound collection unit 200 is not operating, and when a waveform is displayed next to the speaker mark, this indicates that the sound collection unit 200 is operating. "Operating" means that the sound collection unit 200 is collecting sound. The human behavior information may be displayed, for example, next to the speaker mark of the sound collection unit 200 that collected the inaudible sound. For example, the behavior information may be displayed when the user touches the speaker mark. Furthermore, in addition to the behavior information of a person, time information such as the time when the behavior was estimated by the behavior estimation device 100c or the time when the person performed the behavior may be displayed.
[0165] In this way, the behavior inference device 100c according to the fourth embodiment outputs display information including the acquired information and derived information to the external terminal 300, so that the user can check the display information by displaying it on a display unit (not shown) of the external terminal 300.
[0166] [3. Effects, etc.] As described above, the behavior estimation device 100c further includes a display information generation unit 190 that generates display information by superimposing at least one of operation information related to the operation of the sound collection unit 200 and human behavior information estimated based on sound information related to inaudible sounds collected by the sound collection unit 200 on layout information that indicates the arrangement of multiple rooms in a building in which the sound collection unit 200 is installed and that indicates in which of the multiple rooms the sound collection unit 200 is installed, and the output unit 150 further outputs the display information generated by the display information generation unit 190 to the external terminal 300.
[0167] The behavior inference device 100c outputs display information to be displayed on the external terminal 300, and therefore, once behavior information is estimated, the user can confirm the information via the external terminal 300.
[0168] (Other embodiments) While the behavior estimation device and behavior estimation method according to one or more aspects of the present disclosure have been described based on the above-mentioned embodiments, the present disclosure is not limited to these embodiments. As long as they do not deviate from the gist of the present disclosure, various modifications conceivable by those skilled in the art and embodiments formed by combining components of different embodiments may also be included within the scope of one or more aspects of the present disclosure.
[0169] For example, some or all of the components included in the behavior inference device according to the above-described embodiments may be configured as one system LSI (Large Scale Integration). For example, the behavior inference device may be configured as a system LSI having an acquisition unit, a learning unit, a trained model, an estimation unit, and an output unit. Note that the system LSI does not necessarily have to include the learning unit.
[0170] A system LSI is an ultra-multifunctional LSI manufactured by integrating multiple components on a single chip, and is specifically a computer system consisting of a microprocessor, ROM (Read Only Memory), RAM (Random Access Memory), etc. Computer programs are stored in the ROM. The system LSI achieves its functions when the microprocessor operates in accordance with the computer program.
[0171] Although we refer to it as a system LSI here, it may also be called an IC, LSI, super LSI, or ultra LSI depending on the level of integration. Furthermore, the method of integration is not limited to LSI, but may be realized using dedicated circuits or general-purpose processors. It is also possible to use FPGAs (Field Programmable Gate Arrays), which can be programmed after the LSI is manufactured, or reconfigurable processors, which allow the connections and settings of circuit cells within the LSI to be reconfigured.
[0172] Furthermore, if an integrated circuit technology that can replace LSI emerges due to advances in semiconductor technology or other derivative technologies, it is natural that such technology may be used to integrate functional blocks. The application of biotechnology, etc. is also a possibility.
[0173] Furthermore, one aspect of the present disclosure may be not only such a behavior estimation device, but also a behavior estimation method having steps corresponding to characteristic components included in the device. Another aspect of the present disclosure may be a computer program that causes a computer to execute each of the characteristic steps included in the behavior estimation method. Another aspect of the present disclosure may be a computer-readable non-transitory recording medium on which such a computer program is recorded. [Industrial Applicability]
[0174] According to the present disclosure, human behavior can be estimated based on inaudible sounds, and therefore a wider variety of behaviors can be estimated while protecting privacy, making it possible to use the system in a variety of locations, such as homes, workplaces, schools, or commercial facilities. [Explanation of symbols]
[0175] 50 Wide Area Communication Network 100, 100a, 100b, 100c Behavior estimation device 110 Acquisition Department 120, 120a Learning Section 130, 130a Trained model 140, 140a Estimated part 150 Output section 160, 160a, 160b storage section 162 databases 164 Date and time information recording section 170 Location information acquisition unit 180 Adjustment section 190 Display information generation section 200 sound pickup unit 300 External Terminal
Claims
1. an acquisition unit that acquires sound information regarding inaudible sounds that are sounds in the ultrasonic band collected by the sound collection unit; an estimation unit that inputs the sound information acquired by the acquisition unit into a trained model that indicates a relationship between the sound information and behavioral information related to human behavior, and estimates the output result as the behavioral information of the person; a date and time information recording unit that records date and time information regarding the date and time when the inaudible sound is picked up by the sound pickup unit; an adjustment unit that adjusts the sound collection frequency of the sound collection unit based on the number of times the behavior information of the person is estimated by the estimation unit and the date and time information recorded by the date and time information recording unit; an output unit that outputs information about the sound collection frequency adjusted by the adjustment unit to the sound collection unit; Equipped with Behavior estimation device.
2. The sound information input to the trained model includes at least one of a frequency band of the inaudible sound, a duration of the inaudible sound, a sound pressure of the inaudible sound, and a waveform of the inaudible sound. The behavior estimation device according to claim 1 .
3. The form of the sound information input to the trained model is time-series numerical data of the inaudible sound, a spectrogram image, or a frequency characteristic image. The behavior estimation device according to claim 1 or 2.
4. the behavior estimation device further includes a position information acquisition unit that acquires position information regarding a position of a sound source emitting the inaudible sound; The estimation unit estimates the output result obtained by inputting the sound information and the location information acquired by the location information acquisition unit into the trained model as the behavioral information of the person. The behavior estimation device according to any one of claims 1 to 3.
5. the position information acquisition unit acquires, as the position information, a position of the sound source derived based on an installation position of the sound collection unit that collected the inaudible sound; The behavior estimation device according to claim 4 .
6. The position information acquisition unit further acquires, as the position information, a position of the sound source derived based on sound information regarding inaudible sound emitted from an object whose installation position does not change and which is acquired by the acquisition unit. The behavior inference device according to claim 4 or 5.
7. the position information acquisition unit acquires, as the position information, a position of the sound source derived from a direction of the sound source identified based on directivities of the inaudible sounds collected by two or more of the sound collection units; The behavior estimation device according to any one of claims 4 to 6.
8. the behavior estimation device further includes a database in which the position information of the sound source, the sound information related to the inaudible sound emitted from the sound source, and the behavior information of the person are stored in association with each other; The estimation unit further estimates the behavioral information of the person by determining whether the output result of the trained model is likely based on the database. The behavior estimation device according to any one of claims 4 to 7.
9. the behavior estimation device further includes a display information generation unit that generates display information by superimposing, on layout information indicating an arrangement of rooms in a building in which the sound collection unit is installed, at least one of operation information related to an operation of the sound collection unit and the behavior information of the person estimated based on the sound information related to the inaudible sound collected by the sound collection unit; The output unit further outputs the display information generated by the display information generation unit to an external terminal. The behavior estimation device according to claim 1 .
10. an acquisition step of acquiring sound information regarding inaudible sound, which is sound in the ultrasonic band, collected by the sound collection unit; an estimation step of inputting the sound information acquired by the acquisition step into a trained model showing the relationship between the sound information and behavioral information related to human behavior, and estimating the output result as the behavioral information of the person; a date and time information recording step of recording date and time information regarding the date and time when the inaudible sound was picked up by the sound pickup unit; an adjustment step of adjusting a sound collection frequency of the sound collection unit based on the number of times the behavior information of the person is estimated by the estimation step and the date and time information recorded by the date and time information recording step; an output step of outputting information about the sound collection frequency adjusted by the adjustment step to the sound collection unit; Including, Behavior estimation method.
11. A method for causing a computer to execute the behavior estimation method according to claim 10, program.
Citation Information
Patent Citations
Method for estimating behavior of in-house users, device and program
JP2019095517A