Behavior estimation systems, behavior estimation methods, programs
The behavior estimation system uses combined image and state data with machine learning to accurately estimate moving object behavior, addressing the issue of partial obscuration in image data.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- IWATE PREFECTURAL UNIVERSITY
- Filing Date
- 2022-03-10
- Publication Date
- 2026-06-04
Smart Images

Figure 0007870002000002 
Figure 0007870002000003 
Figure 0007870002000004
Abstract
Description
[Technical Field]
[0001] This invention relates to a behavior estimation system, a behavior estimation method, and a program. [Background technology]
[0002] In recent years, research has been progressing in fields such as healthcare, nursing care, security, business and behavioral management, and information provision services, on technologies that estimate the behavior of moving objects (e.g., people, animals, machines, etc.) based on images obtained by capturing those objects.
[0003] For example, the technology described in Patent Document 1 detects a person's region from a series of input images obtained by continuously photographing a person (moving object), and if the person remains still for a predetermined time or longer, it is presumed that the person has fallen over. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Re-tabled publication No. 2018-030024 [Overview of the Initiative] [Problems that the invention aims to solve]
[0005] However, with such technology, if another object (for example, another moving object or a fixture) is present in front of the moving object to be detected, and at least a part of the moving object is obscured by the other object during imaging (resulting in image data where at least a part of the moving object is missing), it may be impossible to accurately detect the region of the moving object to be detected. As a result, it may become difficult to accurately estimate the movement of the moving object.
[0006] The present invention has been made in view of the above problems, and aims to provide an action estimation system, action estimation method, and program that can accurately estimate the action of a moving object even when image data is obtained in which at least a part of the moving object is missing. [Means for solving the problem]
[0007] To solve the above problems, firstly, the present invention provides an action estimation system comprising: a first acquisition means for acquiring an image of a moving object captured by an imaging device; a second acquisition means for acquiring information relating to the state of the moving object measured by a measuring device; and an estimation means for estimating the action of the moving object within a predetermined period based on the image of the moving object captured within a predetermined period and the information relating to the state of the moving object within the predetermined period (Invention 1).
[0008] According to this invention (Invention 1), the behavior of a moving object during a predetermined period is estimated based on an image of the moving object captured during that period and information regarding the state of the moving object during that period. For example, even if at least a part of the moving object is missing from the captured image, it becomes possible to estimate the behavior of the moving object by using information regarding the state of the moving object in addition to the image of the moving object. As a result, even if image data in which at least a part of the moving object is missing is obtained, the behavior of the moving object can be accurately estimated.
[0009] In the above invention (Invention 1), the estimation means may estimate the behavior of the moving object during the predetermined period based on an image of the moving object captured during the predetermined period, information regarding the state of the moving object during the predetermined period, and a trained model based on machine learning that uses the image of the moving object and / or the information regarding the state of the moving object as training data (Invention 2).
[0010] According to this invention (Invention 2), the behavior of a moving object can be easily estimated based on the behavior of the moving object estimated by a trained model using images of the moving object and / or information about the state of the moving object.
[0011] In the above invention (Invention 2), the estimation means may estimate the behavior of the moving object within the predetermined period by newly calculating the estimated probability of a predetermined action using the following: the estimated probability of a predetermined action of the moving object obtained by inputting the image of the moving object captured within the predetermined period into a first trained model, which is a trained model based on machine learning that uses the image of the moving object as training data; and the estimated probability of a predetermined action of the moving object obtained by inputting information about the state of the moving object within the predetermined period into a second trained model, which is a trained model based on machine learning that uses the information about the state of the moving object as training data (Invention 3).
[0012] According to this invention (Invention 3), the behavior of a moving object within a predetermined period can be estimated using the estimated probability of a predetermined action of the moving object obtained by inputting images of the moving object captured within a predetermined period into a first trained model, and the estimated probability of a predetermined action of the moving object obtained by inputting information about the state of the moving object within a predetermined period into a second trained model.
[0013] In the above invention (Invention 2), the estimation means may estimate the behavior of the moving object during the predetermined period by inputting the image of the moving object captured during the predetermined period and information regarding the state of the moving object during the predetermined period into a third trained model, which is a trained model based on machine learning that uses the image of the moving object and the information regarding the state of the moving object as training data (Invention 4).
[0014] According to this invention (Invention 4), by inputting images of a moving object captured within a predetermined period and information regarding the state of the moving object within the predetermined period into a third trained model, the behavior of the moving object within a predetermined period can be estimated.
[0015] In the above inventions (inventions 1 to 4), a detection means is provided to detect at least one part of the moving body based on the acquired image, and the estimation means may estimate the behavior of the moving body during the predetermined period based on information relating to the movement pattern of at least one part of the moving body during the predetermined period and information relating to the state of the moving body during the predetermined period (invention 5).
[0016] According to this invention (Invention 5), by using information regarding the movement patterns of at least one part of the moving body within a predetermined period and information regarding the state of the moving body within a predetermined period, the behavior of the moving body within a predetermined period can be estimated more accurately.
[0017] In the above invention (Invention 5), the detection means may detect at least one joint of the moving body as at least one part of the moving body (Invention 6).
[0018] According to this invention (Invention 6), the behavior of a moving body can be estimated using information regarding the movement of at least one joint of the moving body.
[0019] In the above inventions (Inventions 1 to 6), the information relating to the state of the moving body may include at least one of the following: the position of the moving body, the acceleration of the moving body, the angular velocity of the moving body, and the temperature of the moving body (Invention 7).
[0020] According to this invention (Invention 7), the behavior of a moving body can be estimated using at least one of the following: the position of the moving body, the acceleration of the moving body, the angular velocity of the moving body, and the temperature of the moving body.
[0021] Secondly, the present invention provides a method for estimating the behavior of a moving object, wherein a computer performs the following steps: acquiring an image of a moving object captured by an imaging device; acquiring information regarding the state of the moving object measured by a measuring device; and estimating the behavior of the moving object within a predetermined period based on the image of the moving object captured within a predetermined period and the information regarding the state of the moving object within the predetermined period (Invention 8).
[0022] Thirdly, the present invention provides a program for causing a computer to realize a function of acquiring an image of a moving object captured by an imaging device, a function of acquiring information regarding the state of the moving object measured by a measuring device, and a function of estimating the behavior of the moving object within a predetermined period based on the image of the moving object captured within the predetermined period and the information regarding the state of the moving object within the predetermined period (Invention 9).
Advantages of the Invention
[0023] According to the behavior estimation system, behavior estimation method, and program of the present invention, even when image data in which at least a part of a moving object is missing is obtained, the behavior of the moving object can be accurately estimated.
Brief Description of the Drawings
[0024] [Figure 1] It is a diagram schematically showing the basic configuration of a behavior estimation system according to an embodiment of the present invention. [Figure 2] It is a block diagram showing the configuration of an estimation device. [Figure 3] It is a functional block diagram for explaining functions that play a major role in a behavior estimation system. [Figure 4] It is a diagram showing a configuration example of first acquired data. [Figure 5] It is a diagram showing a configuration example of second acquired data. [Figure 6] It is a diagram showing an example of the relationship between the distance from a measuring device and the received signal strength. [Figure 7] It is a diagram showing an example of a detection result of at least one part of a moving object. [Figure 8] It is a flowchart showing an example of the processing of an estimation means. [Figure 9] It is a diagram showing a configuration example of first learning data. [Figure 10] It is a diagram showing a configuration example of second learning data. [Figure 11] It is a diagram showing an example of the relationship between an estimation coefficient and an estimation accuracy. [Figure 12] This figure shows an example of how estimated data can be structured. [Figure 13] This flowchart shows an example of the main processing steps of the behavior estimation system according to one embodiment of the present invention. [Figure 14] This figure shows an example of how behavioral data can be structured. [Figure 15] (a) and (b) are diagrams illustrating examples of the division of labor between the estimation device and the estimation server for each function of the behavior estimation system. [Modes for carrying out the invention]
[0025] One embodiment of the present invention will be described in detail below with reference to the accompanying drawings. However, this embodiment is illustrative and the present invention is not limited thereto.
[0026] (1) Basic configuration of the behavior estimation system Figure 1 is a schematic diagram showing the basic configuration of an action estimation system according to one embodiment of the present invention. As shown in Figure 1, in the action estimation system according to this embodiment, when a moving object (in this embodiment, a subject T) is present in a predetermined space SP such as indoors, the estimation device 30 acquires an image of the subject T captured by an imaging device 10 installed at a predetermined position in the space SP (in the example of Figure 1, the upper center of the space SP), and information regarding the state of the subject T measured by a measuring device 20.
[0027] Furthermore, the estimation device 30 estimates the actions of subject T during a predetermined period based on images captured during that period and information regarding the state of subject T during that period. Here, the imaging device 10 and the measurement device 20, and the estimation device 30 are connected to a communication network NW (network), such as the Internet or a LAN (Local Area Network).
[0028] The imaging device 10 may be, for example, an imaging device that captures moving images and / or still images (e.g., a digital camera or a digital video camera), and is configured to capture images of the spatial SP at a predetermined position within the spatial SP. The imaging device 10 is configured to perform imaging processing at a predetermined frame rate (e.g., 30 fps (frames per second)) and transmit the captured images to the estimation device 30 via a communication network NW. The imaging device 10 may also perform imaging processing when it receives a predetermined imaging instruction signal from the estimation device 30.
[0029] Here, the imaging device 10 may be an imaging device that captures omnidirectional images (for example, images of the surrounding 360° (in the example of Figure 1, the surrounding 360° in the horizontal direction)) (for example, an omnidirectional camera, etc.). In this case, by widening the imaging range, the actions of subject T can be grasped over a wide area.
[0030] Furthermore, the imaging device 10 may be an imaging device that captures infrared images (for example, an infrared camera). This makes it possible to detect the subject T in the captured image even when the subject T is in an environment with poor visibility (for example, at night, in a dark place, in bad weather, etc.).
[0031] Furthermore, the imaging device 10 may be an imaging device that captures stereo images (for example, a stereo camera). Here, a stereo image may be, for example, a set of two images having a predetermined parallax.
[0032] In this embodiment, the case of imaging the spatial SP using one imaging device 10 is described as an example, but the spatial SP may also be imaged using multiple imaging devices (for example, two imaging devices spaced apart in the horizontal and / or vertical directions).
[0033] The measuring device 20 is a device that measures the state of subject T continuously or intermittently (for example, at predetermined intervals (e.g., 300 milliseconds or 1 second)). Here, the information regarding the state of subject T may be, for example, information regarding the position of subject T, a value representing the physical state of subject T, a value obtained by substituting the value representing the position and physical state of subject T into a predetermined calculation formula, or information representing the degree of the position and physical state of subject T. Furthermore, the information regarding the position of subject T may include, for example, position information measured using position measurement technology such as GPS (Global Positioning System) (for example, at least one of latitude, longitude, and altitude), or it may include information regarding the received signal strength (RSSI) when a signal transmitted by one device is received by the other device. Here, the information regarding the received signal strength may be, for example, a value of the received signal strength, a value obtained by substituting the value of the received signal strength into a predetermined calculation formula, or information representing the degree of the received signal strength.
[0034] The measuring device 20 may be, for example, a device that measures the position of the subject T (e.g., a GPS sensor), or a device that measures the physical condition of the subject T (e.g., acceleration in three axes (which may be acceleration at a predetermined part of the subject T), angular velocity in three axes (which may be angular velocity at a predetermined part of the subject T), heart rate (pulse), blood pressure, body temperature, amount of sweat, number of steps, walking speed, posture, exercise intensity (e.g., heart rate ÷ maximum heart rate), or calories burned, etc.) (e.g., a heart rate monitor, blood pressure monitor, thermometer, sweat meter, three-axis acceleration sensor, three-axis gyroscope, motion sensor, etc.).
[0035] Furthermore, as shown in Figure 1, the measuring device 20 may consist of a first measuring device 21 that can be held by the subject T or worn on the subject T's body, and a plurality of second measuring devices 22 (three in the example of Figure 1) arranged at different locations within the space SP, the plurality of second measuring devices 22 that communicate wirelessly with the first measuring device 21 when the first measuring device 21 is present within the space SP.
[0036] In this case, the first measuring device 21 may be configured to communicate wirelessly with a plurality of second measuring devices 22 using a predetermined wireless communication method (e.g., wireless LAN (e.g., Wi-Fi®)) when it is located within the spatial SP. The first measuring device 21 may also be configured to transmit wireless signals (e.g., probe requests, etc.) containing its own identification information (e.g., MAC (Media Access Control) address, etc.) at predetermined intervals (e.g., intervals of several hundred milliseconds, etc.) in order to communicate wirelessly with a plurality of second measuring devices 22. Furthermore, the first measuring device 21 may be, for example, a device that can be worn by the subject T (e.g., a wearable device), or a portable device that the subject T can carry. Moreover, the first measuring device 21 may be a communication device operated by an individual user, such as a mobile terminal, smartphone, PDA (Personal Digital Assistant), personal computer, or television receiver with two-way communication capabilities (including so-called multi-functional smart TVs).
[0037] Multiple second measuring devices 22 may be positioned within the spatial SP to enable wireless communication with the first measuring device 21 using a predetermined wireless communication method (e.g., wireless LAN (e.g., Wi-Fi®)). Furthermore, multiple second measuring devices 22 may, for example, be devices that relay wireless communication between two or more first measuring devices 21 located within the spatial SP, or devices that relay wireless communication between the first measuring device 21 and other devices (not shown) located within the spatial SP, or devices that relay communication between the first measuring device 21 and other devices (e.g., estimation device 30) connected via a communication network NW. Additionally, multiple second measuring devices 22 may be packet capture devices.
[0038] Furthermore, either the first measuring device 21 or the plurality of second measuring devices 22 may be provided with an RSSI circuit for detecting the received signal strength (RSSI) when it receives a signal transmitted by the other. In addition, this signal may include identification information (e.g., MAC address, etc.) of the measuring device that transmitted the signal (first measuring device 21 or second measuring device 22), and when the received signal strength of this signal is detected by the RSSI circuit, the detected received signal strength and the identification information of the device that transmitted this signal may be stored in a storage device (not shown) provided in the measuring device that received this signal (first measuring device 21 or second measuring device 22) in a state of mutual association.
[0039] In this explanation, we describe a case where wireless communication is performed between the first measuring device 21 and multiple second measuring devices 22 using Wi-Fi® as an example, but the communication method is not limited to this case. For example, wireless communication methods such as Bluetooth®, ZigBee®, UWB, optical wireless communication (e.g., infrared) may be used, or wired communication methods such as USB may be used.
[0040] The measuring device 20 (first measuring device 21 and / or multiple second measuring devices 22) is configured to transmit information about the detected state to the estimation device 30 via the communication network NW each time it detects the state of the subject T.
[0041] In this embodiment, the case in which one first measuring device 21 is provided to the subject T is described as an example, but multiple first measuring devices 21 may be provided to the subject T. Also, in this embodiment, the case in which three second measuring devices 22 are provided in space SP is described as an example, but the number of second measuring devices 22 provided in space SP may be two or fewer, or four or more.
[0042] Furthermore, the imaging device 10 and the measuring device 20 may be configured to communicate directly with the estimation device 30 using a wired or wireless connection, or they may be configured to communicate with the estimation device 30 via a predetermined relay device (not shown) by transmitting and receiving information using a wired or wireless connection.
[0043] The estimation device 30 communicates with the imaging device 10 and the measurement device 20 via a communication network NW, and is configured to acquire images captured by the imaging device 10 and information regarding the state of subject T measured by the measurement device 20 over time via the communication network NW. The estimation device 30 may be a terminal device operated by an individual user, such as a mobile terminal, smartphone, PDA, personal computer, or television receiver with two-way communication capabilities (including so-called multi-functional smart TVs).
[0044] (2) Configuration of the estimation device The configuration of the estimation device 30 will be described with reference to Figure 2. Figure 2 is a block diagram showing the internal configuration of the estimation device 30. As shown in Figure 2, the estimation device 30 comprises a CPU (Central Processing Unit) 31, a ROM (Read Only Memory) 32, a RAM (Random Access Memory) 33, a storage device 34, a display processing unit 35, a display unit 36, an input unit 37, and a communication interface unit 38, and is provided with a bus 30a for transmitting control signals or data signals between each unit.
[0045] When power is supplied to the estimation device 30, the CPU 31 loads various programs stored in the ROM 32 or storage device 34 into the RAM 33 and executes them. In this embodiment, the CPU 31 reads and executes the programs stored in the ROM 32 or storage device 34 to realize the functions of the first acquisition means 41, second acquisition means 42, detection means 43, and estimation means 44 (shown in Figure 3), which will be described later.
[0046] The storage device 34 may be a non-volatile storage device such as flash memory, SSD (Solid State Drive), magnetic storage device (e.g., HDD (Hard Disk Drive), floppy disk (registered trademark), magnetic tape, etc.), or optical disk, or it may be a volatile storage device such as RAM, and it stores programs executed by the CPU 31 and data referenced by the CPU 31. The storage device 34 also stores the first acquired data (shown in Figure 4), the second acquired data (shown in Figure 5), the first learning data (shown in Figure 8), the second learning data (shown in Figure 9), and the estimated data (shown in Figure 12), which will be described later.
[0047] The display processing unit 35 displays the display data provided by the CPU 31 on the display unit 36. The display unit 36 is, for example, an LCD (Liquid Crystal Display) monitor including thin-film transistors arranged in a matrix on a pixel-by-pixel basis, and displays the data to be displayed on the display screen by driving the thin-film transistors based on the display data.
[0048] If the estimation device 30 is a button-input type device, the input unit 37 includes a group of buttons including a plurality of instruction input buttons such as a direction indicator button and a select button for receiving user operation input, and a group of buttons including a plurality of instruction input buttons such as a numeric keypad, and includes an interface circuit for recognizing the pressing (operation) input of each button and outputting it to the CPU 31.
[0049] If the estimation device 30 is a touch panel input device, the input unit 37 primarily accepts input via a touch panel, such as by touching the display screen with a fingertip or pen. The touch panel input method may be a known method such as a capacitive touch method.
[0050] Furthermore, if the estimation device 30 is a device capable of voice input, the input unit 37 may be configured to include a microphone for voice input, or it may include an interface circuit for outputting voice data input via an external microphone to the CPU 31. In addition, if the estimation device 30 is a device capable of inputting moving images and / or still images, the input unit 37 may be configured to include a digital camera or digital video camera for image input, or it may include an interface circuit for receiving image data captured by an external digital camera or digital video camera and outputting it to the CPU 31.
[0051] The communication interface unit 38 includes an interface circuit for communicating with other devices (for example, the imaging device 10 and the measuring device 20, etc.) via a communication network NW.
[0052] (3) Overview of each function in the behavior estimation system The functions realized in the behavior estimation system of this embodiment will be described with reference to Figure 3. Figure 3 is a functional block diagram illustrating the functions that play a major role in the behavior estimation system of this embodiment. In the functional block diagram of Figure 3, the first acquisition means 41, the second acquisition means 42, and the estimation means 44 correspond to the main components of the behavior estimation system of the present invention. Other means (detection means 43) are not necessarily essential components, but they are components that further enhance the present invention.
[0053] The first acquisition means 41 has the function of acquiring an image of the subject T (moving body) captured by the imaging device 10.
[0054] The function of the first acquisition means 41 is realized, for example, as follows. First, the imaging device 10 performs imaging processing at a predetermined frame rate (e.g., 30fps) when, for example, a subject T is present in space SP, and each time imaging processing is performed, it transmits the image data of the captured image to the estimation device 30 via the communication network NW. Here, the image data of the image captured by the imaging device 10 may be transmitted to the estimation device 30 in association with the date and time of capture and the identification information of the imaging device 10 (e.g., the serial number or MAC (Media Access Control) address of the imaging device 10).
[0055] On the other hand, each time the CPU 31 of the estimation device 30 receives (acquires) image data transmitted from the imaging device 10 via the communication interface unit 38, it stores the received image data in the first acquisition data shown in Figure 4, for example, in association with the date and time the image was captured. The first acquisition data is data that describes the image data of each image in association with the date and time the image was captured. In this way, the first acquisition means 41 can acquire an image of the subject T captured by the imaging device 10.
[0056] The CPU 31 may acquire a stereo image of the subject T (moving body) captured by the imaging device 10. In this case, a three-dimensional model of the subject T can be generated based on the stereo image of the subject T, and in the function of the detection means 43 described later, it becomes possible to determine the position (coordinates) of at least one part of the subject T in three-dimensional space based on this three-dimensional model. This makes it possible to capture the position of at least one part of the subject T more accurately compared to, for example, using a two-dimensional model of the subject T, thereby improving the accuracy of estimating the subject T's actions based on the image of the subject T.
[0057] Furthermore, for example, if two imaging devices are provided in space SP, spaced apart in the horizontal and / or vertical directions, the CPU 31 may acquire two images captured by each imaging device at substantially the same time (for example, two images with the same capture date and time, or two images with a time difference of a predetermined range (for example, a few milliseconds to tens of milliseconds, etc.)) as a set of images having a predetermined parallax (i.e., a stereo image).
[0058] The second acquisition means 42 has the function of acquiring information regarding the state of the subject T (moving body) measured by the measuring device 20.
[0059] Here, information regarding the state of the subject T (moving body) may include at least one of the following: the position of the subject T, the acceleration of the subject T, the angular velocity of the subject T, and the temperature of the subject T. Furthermore, information regarding the position of the subject T may include positional information measured using position measurement technology such as GPS (for example, at least one of latitude, longitude, and altitude), or it may include information regarding the received signal strength (RSSI) when a signal transmitted by one of the first measuring device 21 and the plurality of second measuring devices 22 is received by the other measuring device. This allows the behavior of the subject T to be estimated using at least one of the following: the position of the subject T, the acceleration of the subject T, the angular velocity of the subject T, and the temperature of the subject T.
[0060] The function of the second acquisition means 42 is realized, for example, as follows. First, when the subject T is present in space SP, the first measuring device 21 continuously or intermittently (for example, at predetermined intervals (e.g., 300 milliseconds or 1 second)) measures the state of the subject T (for example, at least one of the following: the position of the subject T, the acceleration of the subject T in three or two axes, the angular velocity of the subject T in three or two axes, and the temperature of the subject T), and transmits information about the measured state to the estimation device 30 via the communication network NW each time. Here, the information about the state of the subject T measured by the first measuring device 21 may be transmitted to the estimation device 30 in association with the measurement date and time and the identification information of the first measuring device 21 (for example, the serial number or MAC address of the first measuring device 21).
[0061] Here, we will describe a case where the information regarding the state of subject T includes information regarding the location of subject T, and the information regarding the location of subject T includes the received signal strength (RSSI) of signals transmitted and received in communication between the first measuring device 21 and the plurality of second measuring devices 22. For example, when subject T is in space SP, the first measuring device 21 communicates wirelessly with each of the plurality of second measuring devices 22, and each time it receives a wireless signal (e.g., a beacon signal) transmitted from each of the plurality of second measuring devices 22 at predetermined intervals (e.g., intervals of several hundred milliseconds), it stores the value of the received signal strength (RSSI) of the wireless signal detected (measured) by the RSSI circuit in a storage device (not shown) provided in the first measuring device 21. Here, the RSSI value may be stored in association with the measurement date and time (for example, the date and time when the first measuring device 21 received the radio signal corresponding to the RSSI value from any of the second measuring devices 22) and the identification information of any of the second measuring devices 22 that transmitted the radio signal corresponding to the RSSI value (for example, the serial number or MAC address of the second measuring device 22). The first measuring device 21 may then transmit to the estimation device 30 via the communication network NW the value of information regarding the state of subject T measured by the first measuring device 21 within the predetermined period (i.e., the RSSI of all radio signals received by the first measuring device 21 from multiple second measuring devices 22 within the predetermined period) at predetermined intervals (for example, every 3 seconds). Here, the information regarding the state of subject T measured by the first measuring device 21 within the predetermined period may be transmitted to the estimation device 30 in association with the identification information of the first measuring device 21 (for example, the serial number or MAC address of the first measuring device 21).
[0062] In this explanation, we have described an example in which the first measuring device 21 transmits the RSSI value when it receives a wireless signal transmitted from each of the multiple second measuring devices 22 to the estimation device 30 as information regarding the state of subject T. However, each of the multiple second measuring devices 22 may also transmit the RSSI value when it receives a wireless signal transmitted from the first measuring device 21 to the estimation device 30 as information regarding the state of subject T.
[0063] Meanwhile, the CPU 31 of the estimation device 30 receives (acquires) information regarding the state of subject T transmitted from the measurement device 20 via the communication interface unit 38, and stores the received state information in the second acquisition data shown in Figure 5, for example, in association with the measurement date and time of the information. The second acquisition data is data that describes information regarding the state of subject T (in the example shown in the figure, "measurement information") in association with each measurement date and time. In this way, the second acquisition means 42 can acquire the state of subject T measured by the measurement device 20 (in this case, at least one of the following: the position of subject T, the acceleration of subject T in the three-axis or two-axis direction, the angular velocity of subject T in the three-axis or two-axis direction, and the temperature of subject T).
[0064] Furthermore, the CPU 31 of the estimation device 30 may acquire at least one of the following based on the function of the second acquisition means 42: the position of the subject T, the average and / or variance of the subject T's acceleration in the three-axis or two-axis direction, the average and / or variance of the subject T's angular velocity in the three-axis or two-axis direction, and the subject T's temperature.
[0065] Furthermore, in this embodiment, the received signal strength (RSSI) of the signals transmitted and received during communication between the first measuring device 21 and the plurality of second measuring devices 22 is included in the information regarding the position of the subject T. Using this received signal strength (RSSI), it is possible to determine the distance between the first measuring device 21 and each of the plurality of second measuring devices 22, and consequently, the position of the first measuring device 21 in space SP (i.e., the position of the subject T). Specifically, the distance between the first measuring device 21 and each of the plurality of second measuring devices 22 can be calculated, for example, by using the following equations (1) and (2). P r =P t +G r +G t -L …(1)
number
[0066] Thereby, it becomes possible to obtain the distance between the first measurement device 21 and each of the plurality of second measurement devices 22 by using the received signal strength (RSSI) of the signal between the first measurement device 21 and each of the plurality of second measurement devices 22. Also, it becomes possible to obtain the position of the first measurement device 21 in the space SP (that is, the position of the subject T) by using the distance between the first measurement device 21 and each of the plurality of second measurement devices 22.
[0067] The detection means 43 has a function of detecting at least one part of the subject T (moving body) based on the acquired image.
[0068] Further, the detection means 43 may detect at least one joint of the subject T (moving body) as at least one part of the subject T. Thereby, it is possible to estimate the behavior of the subject T by using the information regarding the movement mode of at least one joint of the subject T.
[0069] The function of the detection means 43 is implemented, for example, as follows: The CPU 31 of the estimation device 30 may, for example, based on the function of the first acquisition means 41, store the image data transmitted from the imaging device 10 in the first acquisition data, set up nodes in the image corresponding to the positions of at least one part of the subject T, and calculate the coordinates of each node in the image. Here, the setting of each node and the calculation of the coordinates may be performed, for example, using deep learning-based feature point (keypoint) detection technology (e.g., OpenPose). Furthermore, the coordinates of each node may be two-dimensional coordinates, or, for example, three-dimensional coordinates if a three-dimensional model of the subject T is generated from the image data (stereo image).
[0070] For example, if the imaging device 10 captures an omnidirectional image, the CPU 31 may unfold the omnidirectional image captured by the imaging device 10 into a panoramic image and then perform node setting and node coordinate calculation on the unfolded panoramic image.
[0071] Furthermore, if stereo images of subject T are acquired, the CPU 31 may perform preprocessing (e.g., noise reduction) on each image in a set of images having a predetermined parallax, and then perform a well-known stereo matching process to generate a three-dimensional model of subject T.
[0072] As shown in Figure 7, for example, the CPU 31 detects the position (in the example, the X-direction coordinate and the Y-direction coordinate) of at least one part of the subject T (in the example, "neck," "right shoulder," "left shoulder," "right buttock," "left buttock," etc.) in the image captured by the imaging device 10, and sets a node (in the example, 14 nodes) corresponding to each of the detected parts. Here, the coordinates of each node may be stored in the first acquired data in a state where each node is associated with the image in which it was set. Furthermore, if, for example, at least one part of the subject T cannot be detected because it is obscured by another object (for example, another moving object or installed object, etc.) during imaging, the CPU 31 may store the coordinates of the node corresponding to that part as NULL data in the first acquired data.
[0073] In this explanation, we have described an example in which at least one part of subject T (e.g., "neck," "right shoulder," "left shoulder," "right buttock," "left buttock") is detected using the image captured by the imaging device 10. However, for example, at least one joint of subject T (e.g., "right shoulder joint," "right elbow joint," "right hip joint," "right knee joint," "left shoulder joint," "left elbow joint," "left hip joint," and "left knee joint") may also be detected using the image captured by the imaging device 10.
[0074] The estimation means 44 has the function of estimating the actions of the subject T (moving body) during a predetermined period based on images of the subject T (moving body) captured during that predetermined period and information regarding the state of the subject T (moving body) during that predetermined period.
[0075] Here, the estimation means 44 may estimate the actions of the subject T (moving body) during a predetermined period based on images of the subject T (moving body) captured during that predetermined period, information about the state of the subject T (moving body) during that predetermined period, and a trained model based on machine learning that uses the images of the subject T (moving body) and / or the information about the state of the subject T (moving body) as training data. This makes it possible to easily estimate the actions of the subject T based on the actions of the subject T estimated by the trained model using the images of the subject T and / or the information about the state of the subject T.
[0076] Furthermore, the estimation means 44 may estimate the actions of subject T (moving body) within a predetermined period by newly calculating the estimated probability of a predetermined action using the following: the estimated probability of a predetermined action of subject T (moving body) obtained by inputting images of subject T (moving body) captured within a predetermined period into a first trained model, which is a trained model based on machine learning that uses images of subject T (moving body) as training data; and the estimated probability of a predetermined action of subject T (moving body) obtained by inputting information about the state of subject T (moving body) within the predetermined period into a second trained model, which is a trained model based on machine learning that uses information about the state of subject T (moving body) as training data.
[0077] Furthermore, the estimation means 44 may estimate the behavior of the subject T (moving body) within a predetermined period by inputting images of the subject T (moving body) captured within a predetermined period and information regarding the state of the subject T (moving body) within the predetermined period into a third trained model, which is a trained model based on machine learning that uses the images of the subject T (moving body) and information regarding the state of the subject T (moving body) as training data. In this way, the behavior of the subject T within a predetermined period can be estimated by inputting images of the subject T captured within a predetermined period and information regarding the state of the subject T within the predetermined period into the third trained model.
[0078] Furthermore, the estimation means 44 may estimate the behavior of the subject T (moving body) during a predetermined period based on information regarding the movement patterns of at least one part of the subject T (moving body) during that predetermined period and information regarding the state of the subject T (moving body) during that predetermined period.
[0079] The function of the estimation means 44 is realized, for example, as follows. Here, we will explain the case in which the estimation means 44 estimates the behavior of subject T during a predetermined period based on images of subject T taken within that predetermined period, information about the state of subject T during that predetermined period, a first trained model based on machine learning that uses the images of subject T as training data, and a second trained model based on machine learning that uses the information about the state of subject T as training data.
[0080] Referring to the flowchart in Figure 8, an example of the processing of the estimation means 44 in this embodiment will be explained. First, the CPU 31 of the estimation device 30 estimates the actions of the subject T (moving body) within a predetermined period based on the images of the subject T captured within that predetermined period (step S100). Specifically, the CPU 31 of the estimation device 30 accesses the first acquired data and extracts all image data captured within a predetermined period (for example, 3 seconds) starting from a predetermined imaging date and time. Here, all extracted image data is associated with the coordinates of each node corresponding to at least one part of the subject T detected based on the function of the detection means 43. Next, the CPU 31 further extracts the coordinates of each node corresponding to at least one part of the subject T (in this embodiment, "neck", "right shoulder", "left shoulder", "right hip", "left hip", etc.) from all the extracted image data. This provides information on the time progression of the position of at least one part of the subject T within the predetermined period. Here, information regarding the time progression of the position of at least one part of subject T within a predetermined period is an example of "information regarding the movement of at least one part of a moving body within a predetermined period" according to the present invention.
[0081] Next, the CPU 31 estimates the actions of the subject T during a predetermined period by inputting the coordinates of each node corresponding to at least one part of the subject T (information regarding the time progression of the position of at least one part of the subject T during a predetermined period) into a first trained model based on machine learning, which uses information regarding the movement patterns of at least one part of the subject T as first training data. Here, by inputting the coordinates of each node corresponding to at least one part of the subject T into the first trained model, the estimated probability P(skeleton) of one or more actions (in this embodiment, "remove the outer casing," "connect the cable," "apply a sticker to the main unit," "read the barcode," "assemble," etc.) is obtained.
[0082] An example of the first training data is shown in Figure 9. The first training data shown in Figure 9 is data that describes the time progression of the position of at least one body part of subject T (in the example, "neck," "right shoulder," "left shoulder," "right hip," "left hip") while subject T is performing a predetermined action (in the example, "remove the outer casing," "connect the cable," "apply a sticker to the main unit," "read the barcode," "assemble," etc.), with the position of the body part corresponding to the action (action label). As a result of machine learning using the first training data, a first trained model is constructed that shows the time progression of the position of at least one body part of subject T performing a predetermined action over a predetermined period, and its relationship to that action.
[0083] The data described in the first training data may include data from cases where the same subject performs different actions, or data from cases where different subjects perform the same action.
[0084] Furthermore, the CPU 31 may also have a function to learn a model (first trained model) used to estimate the behavior of subject T based on captured images by machine learning using information on the movement patterns of at least one part of subject T as first training data.
[0085] In this case, the CPU 31 may, for example, train a model using the first training data shown in Figure 9 when a predetermined model training instruction is input via the input unit 37. The CPU 31 may, for example, train using a time-series-responsive neural network model. Here, as the time-series-responsive neural network, for example, an RNN (Recurrent Neural Network) or an advanced version of RNN such as LSTM (Long Short-Term Memory) can be applied. The CPU 31 may also train using any of several models, such as a graph neural network (GNN) model, a convolutional neural network (CNN) model, a support vector machine (SVM) model, a fully connected neural network (FNN) model, a gradient boosting (HGB) model, or a wavenet (WN) model. In this embodiment, it is possible to improve the accuracy of action estimation by using any of the following GNN derivatives: a graph convolutional neural network (GCN) model, a graph attention network (GAT) model, or a graph convolutional LSTM (GC-LSTM) model.
[0086] In this way, the CPU 31 inputs the time progression of the position of at least one part of the subject T within a predetermined period into the first trained model, and based on the images of the subject T taken within that predetermined period, it can estimate the subject T's actions within that predetermined period (in this case, "removing the outer casing," "connecting the cable," "applying a sticker to the main unit," "reading a barcode," "assembling," etc.).
[0087] Next, the CPU 31 of the estimation device 30 estimates the actions of subject T (moving object) during a predetermined period based on information regarding the subject T's state during that period (step S102). Specifically, the CPU 31 of the estimation device 30 accesses the second acquired data and extracts measurement information (information regarding the subject T's state) stored in the second acquired data during a predetermined period (e.g., 3 seconds) starting from a predetermined measurement date and time. The measurement date and time used as the starting point when extracting measurement information from the second acquired data may be the same as the imaging date and time used as the starting point when extracting image data from the first acquired data. The CPU 31 then inputs the extracted measurement information (information regarding the time progression of information regarding the subject T's state during the predetermined period) into a second trained model based on machine learning that uses information regarding the subject T's state as second training data, thereby estimating the actions of subject T during the predetermined period. In this process, information about the state of subject T is input into the second trained model, allowing for the estimation of the probability P(sensor) of one or more actions (in this embodiment, "remove the outer casing," "connect the cable," "apply a sticker to the main unit," "read the barcode," "assemble," etc.).
[0088] An example of the second training data is shown in Figure 10. The first training data shown in Figure 10 is data that describes the time progression of the state of subject T (at least one of subject T's position in space SP, acceleration, angular velocity, temperature, etc.) while subject T is performing a predetermined action (in the example in the figure, one of "remove the outer casing", "connect the cable", "apply a sticker to the main unit", "read the barcode", "assemble", etc.), with the state corresponding to the action (action label). As a result of machine learning using the second training data, a second trained model is constructed that shows the time progression of subject T's state during a predetermined period and its relationship to the action.
[0089] Furthermore, the data described in the second training data may include data from cases where the same subject performs different actions, or data from cases where different subjects perform the same action.
[0090] Furthermore, the CPU 31 may also be equipped with a function to learn a model (second trained model) used to estimate the actions of subject T based on the state of subject T, by machine learning using information about the time progression of subject T's state as second training data.
[0091] In this case, the CPU 31 may, for example, train a model using the second training data shown in Figure 10 when a predetermined model training instruction is input using the input unit 37. The CPU 31 may, for example, train using any of several models, such as a time-series neural network model, GNN model, CNN model, SVM model, FNN model, HGB model, WN model, or convolutional LSTM (GC-LSTM) model, similar to the first trained model described above.
[0092] In this way, the CPU 31 inputs the time progression of the subject T's state over a predetermined period into the second trained model, and based on the information regarding the subject T's state over that predetermined period, it can estimate the subject T's actions during that period (in this case, "remove the outer casing," "connect the cable," "apply a sticker to the main unit," "read the barcode," "assemble," etc.).
[0093] In this embodiment, the CPU 31 learns a model (first trained model and second trained model) provided in the estimation device 30 and uses this trained model to estimate the behavior of subject T. However, the present invention is not limited to this case. For example, the first trained model and / or the second trained model may be provided in a device other than the estimation device 30. In this case, the CPU 31 may input information regarding the time progression of the position of at least one part of subject T within a predetermined period, and / or information regarding the time progression of the state of subject T within a predetermined period, to a trained model provided in the other device, and receive (acquire) information regarding the behavior estimated by the trained model from the other device.
[0094] Next, the CPU 31 of the estimation device 30 estimates the actions of the subject T (moving object) within a predetermined period based on the action estimated from the image of the subject T (moving object) and the action estimated from the information regarding the state of the subject T (moving object) (step S104). Specifically, the CPU 31 of the estimation device 30 inputs the image of the subject T captured within the predetermined period into the first pre-trained model to obtain the estimated probability P(skeleton) of a predetermined action of the subject T, and inputs the information regarding the state of the subject T within the predetermined period into the second pre-trained model to obtain the estimated probability P(sensor) of a predetermined action of the subject T, and uses these to estimate the actions of the subject T within the predetermined period. For example, for each of a plurality of actions (here, "remove the exterior packaging", "connect the cable", "attach a sticker to the main body", "read the barcode", "assemble", etc.), the CPU 31 newly calculates the estimated probability P using the following formula (3). P = P(skeleton) × w + P(sensor) × (1 - w) …(3) In formula (3), w (0 < w < 1) represents an estimation coefficient representing the weight of the image data. The CPU 31 may, for example, calculate the estimated probability P for each of the plurality of actions, and estimate the action with the highest estimated probability P among the plurality of actions (for example, "remove the exterior packaging") as the action of the subject T within the predetermined period.
[0095] FIG. 11 shows an example of the relationship between the estimation coefficient w and the standardized estimation accuracy for each loss rate of the part of the subject T in the image data. In the example shown in FIG. 11, within the range of 0.55 < w < 0.65, the estimation accuracy at each loss rate is the highest. Also, as the loss rate of the part of the subject T in the image data increases, the value of the estimation coefficient w at which the highest estimation accuracy is achieved becomes larger.
[0096] In this way, the CPU 31 can estimate the actions of the subject T within a predetermined period based on the action estimated from the image of the subject T and the action estimated from the information regarding the state of the subject T.
[0097] The CPU 31 may store, for example, the estimated data shown in Figure 12, information regarding the time progression of at least one body part of the subject T within a predetermined period, information regarding the time progression of measurement information of the subject T within the predetermined period, and the estimated behavior of the subject T within the predetermined period. The estimated data shown in Figure 12 is data in which the time progression of at least one body part of the subject T within a predetermined period, the time progression of measurement information of the subject T within the predetermined period, and the behavior of the subject T within the predetermined period are associated with each other.
[0098] Furthermore, the CPU 31 may present information regarding the subject T's actions within a predetermined period. For example, if the CPU 31 estimates the subject T's actions within a predetermined period, it may display the estimated information regarding the subject T's actions within that period on, for example, the display unit 36. Here, the information regarding the subject T's actions within a predetermined period may consist of text data or image data. Also, if the information regarding the subject T's actions within a predetermined period consists of audio data, the CPU 31 may output the information regarding the subject T's actions within a predetermined period from an audio output device such as a speaker.
[0099] Furthermore, the CPU 31 may transmit information regarding the actions of the subject T within a predetermined period to other computers (e.g., servers) via a communication network NW.
[0100] (4) Main processing flow of the behavior estimation system of this embodiment Next, an example of the main processing flow performed by the behavior estimation system of this embodiment will be explained with reference to the flowchart in Figure 13.
[0101] First, the CPU 31 of the estimation device 30 acquires an image of the subject T (moving body) captured by the imaging device 10 based on the function of the first acquisition means 41 (step S200). Next, the CPU 31 of the estimation device 30 acquires information regarding the state of the subject T (moving body) measured by the measuring device 20 (for example, at least one of the following: the position of the subject T, the acceleration of the subject T in the three-axis or two-axis direction, the angular velocity of the subject T in the three-axis or two-axis direction, and the temperature of the subject T) based on the function of the second acquisition means 42 (step S202). After step S200, the CPU 31 of the estimation device 30 may detect at least one part of the subject T (moving body) using the acquired image based on the function of the detection means 43.
[0102] Next, the CPU 31 of the estimation device 30 estimates the actions of the subject T (moving object) during a predetermined period, based on the function of the estimation means 44, using images of the subject T (moving object) captured during that predetermined period and information regarding the state of the subject T (moving object) during that predetermined period (step S204). Here, the CPU 31 of the estimation device 30 may estimate the actions of the subject T during that predetermined period based on images of the subject T captured during that predetermined period, information regarding the state of the subject T during that predetermined period, a first trained model based on machine learning using images of the subject T as first training data, and a second trained model based on machine learning using information regarding the state of the subject T as second training data. Alternatively, the CPU 31 of the estimation device 30 may train a model (first trained model) by machine learning using images of the subject T as first training data, or it may train a model (second trained model) by machine learning using information regarding the state of the subject T as second training data.
[0103] Furthermore, after step S204, the CPU 31 of the estimation device 30 may display information regarding the subject T's actions during the estimated predetermined period on, for example, the display unit 36. Alternatively, the CPU 31 of the estimation device 30 may transmit information regarding the subject T's actions during the predetermined period to another computer (for example, a server) via a communication network NW.
[0104] As described above, according to the behavior estimation system, behavior estimation method, and program of this embodiment, the behavior of subject T during a predetermined period is estimated based on images of subject T captured during that predetermined period and information regarding the state of subject T during that predetermined period. For example, even if at least a part (at least one part) of subject T is missing from the captured image, it becomes possible to estimate the behavior of subject T by using information regarding the state of subject T in addition to the image of subject T. This makes it possible to accurately estimate the behavior of subject T even when image data is obtained in which at least a part (at least one part) of subject T is missing.
[0105] The following describes some variations of the embodiments described above. (modified version) In the above embodiment, an example was described in which the estimation means 44 estimates the actions of a subject T during a predetermined period based on images of the subject T captured during the predetermined period, information about the state of the subject T during the predetermined period, a first trained model based on machine learning using the images of the subject T as first training data, and a second trained model based on machine learning using the information about the state of the subject T as second training data. However, the present invention is not limited to this case. For example, the estimation means 44 may estimate that the subject T is performing a predetermined action if the movement pattern of the subject T in the captured images during the predetermined period satisfies a first condition corresponding to a predetermined action, and the information about the state of the subject T during the predetermined period satisfies a second condition corresponding to the predetermined action. Here, the first condition may be, for example, that the position of at least one part of the subject T during the predetermined period is within the range of the position of at least one part of the subject T corresponding to a predetermined action. The second condition may be, for example, that the information about the state of the subject T during the predetermined period is within the range of the information about the state of the subject T corresponding to a predetermined action.
[0106] In this case, the CPU 31 of the estimation device 30 may, by referring to the behavior data shown in Figure 14, estimate the behavior of subject T during a predetermined period based on the captured images taken during that period and information regarding the state of subject T during that period. The behavior data is data that, for each of several behaviors (in the example shown in the figure, "remove the outer casing," "connect the cable," "apply a sticker to the main unit," "read a barcode," "assemble," etc.), is described in a manner that associates the time progression of the range of position of at least one part of subject T (in the above embodiment, "neck," "right shoulder," "left shoulder," "right hip," "left hip," etc.) when it is estimated that the corresponding behavior is being performed with the time progression of the range of measurement information of subject T when it is estimated that the corresponding behavior is being performed. The behavior data may be stored in, for example, a storage device 34.
[0107] For example, the CPU 31 may estimate that subject T performed one of the actions (e.g., "remove the outer casing") within a predetermined period if it determines that the time progression of the position of at least one part of subject T within a predetermined period falls within the range of the positions of each part corresponding to any action in the action data (satisfying the first condition), and that the time progression of information regarding subject T's state within the predetermined period falls within the range of measurement information corresponding to any action (satisfying the second condition). The CPU 31 may also store the estimated information regarding the action in, for example, RAM 33 or storage device 34.
[0108] Thus, the behavior estimation system, behavior estimation method, and program according to this modified example can achieve the same effects and advantages as the embodiments described above.
[0109] The program of the present invention may be stored on a computer-readable storage medium. The storage medium on which this program is recorded may be the ROM 32, RAM 33, or storage device 34 of the estimation device 30 shown in Figure 2. Alternatively, the storage medium may be a CD-ROM or the like that can be read by being inserted into a program reading device such as a CD-ROM drive. Furthermore, the storage medium may be magnetic tape, cassette tape, flexible disk, MO / MD / DVD, or semiconductor memory.
[0110] The embodiments and modifications described above are provided to facilitate understanding of the present invention and are not intended to limit it. Accordingly, each element disclosed in the above embodiments and modifications is intended to include all design changes and equivalents that fall within the technical scope of the present invention.
[0111] For example, in the embodiment described above, the CPU 31 of the estimation device 30 estimated the actions of the subject T during a predetermined period based on the functions of the estimation means 44, using images of the subject T captured during the predetermined period, information about the state of the subject T during the predetermined period, a first trained model based on machine learning that uses the images of the subject T as training data, and a second trained model based on machine learning that uses the information about the state of the subject T as training data. However, the present invention is not limited to this case. For example, the CPU 31 of the estimation device 30 may estimate the actions of the subject T during a predetermined period by inputting images of the subject T captured during the predetermined period and information about the state of the subject T during the predetermined period into a third trained model, which is a trained model based on machine learning that uses the images of the subject T and the information about the state of the subject T as training data, based on the functions of the estimation means 44. In this case, the CPU 31 of the estimation device 30 may have the function of learning a model (third trained model) used to estimate the actions of subject T based on the captured image and state information of subject T, by machine learning using the image and state information of subject T as training data (for example, training data combining the first and second training data). Here, the model used for machine learning may be the same as the first trained model and the second trained model.
[0112] Furthermore, although the above-described embodiment explained the case where the moving object is a person (subject T) as an example, it is not limited to this case. The moving object can be anything that can be the subject of behavior estimation, such as an animal other than a person, or an object such as a work machine, vehicle, or flying object.
[0113] Furthermore, in the embodiment described above, the case in which the neck, right shoulder, left shoulder, right buttock, and left buttock of subject T (moving body) are detected as body parts of subject T was explained as an example, but other body parts (for example, head, hands, feet, joints, etc.) may also be detected.
[0114] Furthermore, while the above-described embodiment explained the case in which the behavior of one moving object (subject T) within space SP is estimated as an example, the behavior of each of multiple moving objects within space SP may also be estimated.
[0115] Furthermore, in the embodiments described above, we explained as an example the case in which any of the following actions are estimated to be the actions of the moving object (subject T): "removing the outer casing," "connecting the cable," "applying a sticker to the main unit," "reading a barcode," or "assembling." However, the content of the actions is not limited to these.
[0116] Furthermore, although the above-described embodiment explained the case in which one estimation device 30 is provided as an example, it is not limited to this case. For example, multiple estimation devices 30 may be provided, in which case the operation content and processing results on any of the estimation devices 30 may be displayed in real time on other estimation devices 30, or the processing results on any of the estimation devices 30 may be shared among the multiple estimation devices 30.
[0117] Furthermore, in the embodiments described above, the case in which information relating to the time progression of the position of at least one part of the subject T within a predetermined period corresponds to the "information relating to the movement of at least one part of a moving body within a predetermined period" of the present invention was explained as an example, but the invention is not limited to this case. The "information relating to the movement of at least one part of a moving body within a predetermined period" may be, for example, information relating to the time progression of the angle of at least one part of the subject T within a predetermined period, or information relating to the time progression of the distance between multiple parts of the subject T within a predetermined period, etc.
[0118] Furthermore, in the above-described embodiment, the estimation device 30 is configured to realize the functions of the first acquisition means 41, the second acquisition means 42, the detection means 43, and the estimation means 44, but the configuration is not limited to this. For example, a computer (e.g., a general-purpose personal computer or server) that is connected to the estimation device 30 in a communication network such as the Internet or a LAN may realize the function of at least one of the above means 41 to 44. For example, each function in the functional block diagram shown in Figure 3 may be arbitrarily divided between the estimation device 30 and an estimation server, which is an example of a computer connected to the estimation device 30 in a communication network, as shown in Figures 15(a) and (b).
[0119] Furthermore, the function of at least one of the above means 41 to 44 may be realized by the imaging device 10 and / or the measuring device 20. [Industrial applicability]
[0120] The behavior estimation system, behavior estimation method, and program of the present invention described above can accurately estimate the behavior of a moving object even when image data is obtained in which at least a part of the moving object is missing. For example, they can be suitably used in business management systems that perform behavioral analysis of moving objects (e.g., workers, work machines, etc.), information provision systems that provide appropriate information according to the behavior of moving objects, and monitoring systems for moving objects (e.g., hospitalized patients, residents of facilities, pets, etc.), and therefore have extremely great industrial applicability. [Explanation of Symbols]
[0121] 10…Imaging device 20... Measuring device 30…Estimation device 41…First acquisition means 42…Second acquisition means 43...Detection means 44...Estimation means T…Target person
Claims
1. A first acquisition means for acquiring an image of a moving object captured by an imaging device, A second acquisition means for acquiring information regarding the state of the moving body measured by the measuring device, Estimation means for estimating the behavior of the moving object during a predetermined period based on images of the moving object captured during a predetermined period and information regarding the state of the moving object during the predetermined period, Equipped with, The estimation means is an action estimation system that estimates the actions of the moving object during the predetermined period based on images of the moving object captured during the predetermined period, information regarding the state of the moving object during the predetermined period, and a trained model based on machine learning that uses the images of the moving object and the information regarding the state of the moving object as training data.
2. The behavior estimation system according to claim 1, wherein the estimation means estimates the behavior of the moving object within a predetermined period by newly calculating the estimated probability of a predetermined action of the moving object using the following: the estimated probability of a predetermined action of the moving object obtained by inputting images of the moving object captured within the predetermined period into a first trained model, which is a trained model based on machine learning that uses images of the moving object as training data; and the estimated probability of a predetermined action of the moving object obtained by inputting information about the state of the moving object within the predetermined period into a second trained model, which is a trained model based on machine learning that uses information about the state of the moving object as training data.
3. The behavior estimation system according to claim 1, wherein the estimation means estimates the behavior of the moving object during the predetermined period by inputting images of the moving object captured during the predetermined period and information regarding the state of the moving object during the predetermined period into a third trained model, which is a trained model based on machine learning that uses the images of the moving object and the information regarding the state of the moving object as training data.
4. The system includes detection means for detecting at least one part of the moving body based on the acquired image, The behavior estimation system according to any one of claims 1 to 3, wherein the estimation means estimates the behavior of the moving body during the predetermined period based on information relating to the movement patterns of at least one part of the moving body during the predetermined period and information relating to the state of the moving body during the predetermined period.
5. The behavior estimation system according to claim 4, wherein the detection means detects at least one joint of the moving body as at least one part of the moving body.
6. The behavior estimation system according to any one of claims 1 to 5, wherein the information relating to the state of the moving body includes at least one of the position of the moving body, the acceleration of the moving body, the angular velocity of the moving body, and the temperature of the moving body.
7. Computers The steps include acquiring an image of a moving object captured by an imaging device, The steps include: acquiring information regarding the state of the moving body measured by a measuring device; A step of estimating the behavior of the moving object during a predetermined period based on an image of the moving object captured during the predetermined period and information regarding the state of the moving object during the predetermined period. Perform each step, An action estimation method that estimates the behavior of a moving object during a predetermined period, based on an image of the moving object captured during the predetermined period, information regarding the state of the moving object during the predetermined period, and a trained model based on machine learning that uses the image of the moving object and the information regarding the state of the moving object as training data.
8. On the computer, A function to acquire images of moving objects captured by an imaging device, A function to acquire information regarding the state of the moving object measured by the measuring device, A function for estimating the behavior of the moving object during a predetermined period based on images of the moving object captured during that predetermined period and information regarding the state of the moving object during that predetermined period. This is a program to achieve this. A program that estimates the behavior of a moving object during a predetermined period based on images of the moving object captured during that predetermined period, information regarding the state of the moving object during that predetermined period, and a trained model based on machine learning that uses the images of the moving object and the information regarding the state of the moving object as training data.