Detection system, trained model generation method, detection method, and program

The detection system uses a trained model to enhance the estimation of human behavior within a facility by accurately identifying and analyzing objects, simplifying the detection process.

JP7759605B2Active Publication Date: 2025-10-24PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2020116692
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-07-06
Publication Date
2025-10-24
Estimated Expiration
2040-07-06

AI Technical Summary

Technical Problem

Existing systems struggle to easily estimate human behavior within a facility, lacking accuracy in detecting and identifying objects of interest.

Method used

A detection system utilizing an acquisition unit, estimation unit, and output unit, employing a machine-learned trained model to analyze images captured by an imaging unit, estimating the presence of objects and related behaviors within a detection range.

Benefits of technology

Enhances the accuracy of estimating human behavior within a facility by simplifying the detection process to a single stage, improving the estimation of object presence and behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007759605000001
    Figure 0007759605000001
  • Figure 0007759605000002
    Figure 0007759605000002
  • Figure 0007759605000003
    Figure 0007759605000003
Patent Text Reader

Abstract

To provide a detection system capable of easily estimating behavior of a person within a detection range in a facility.SOLUTION: A detection system 100 includes an acquisition unit 11, an estimation unit 12, and an output unit 13. The acquisition unit 11 acquires an image P1 within a facility F1 taken by an imaging unit 2. The estimation unit 12 estimates whether or not a target object A1 is included within a detection range in the image acquired by the acquisition unit 11 by using a machine-learned trained model 121. The output unit 13 outputs a piece of information concerning a related object B1 existing in the facility F1 and related to at least the behavior of person A11 in the facility F1 based on the estimation result by the estimation unit 12.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure generally relates to a detection system, a method for generating a trained model, a detection method, and a program. More specifically, the present disclosure relates to a detection system that detects an object of interest within a detection range within a facility, and a method for generating a trained model used in the detection system. The present disclosure also relates to a detection method that detects an object of interest within a detection range within a facility, and a program for executing the detection method. [Background technology]

[0002] Patent Document 1 discloses an image sensor capable of identifying the presence or absence of a moving object and identifying a specific object. This image sensor includes an imaging unit, an extraction unit, a motion identification unit, and an object identification unit. The imaging unit captures an image of a target space to acquire image data. The extraction unit extracts features from the image data. The motion identification unit identifies the presence or absence of a moving object in the target space based on the features. The object identification unit identifies whether an object in the target space is a pre-registered object based on the features. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-204165 Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure aims to provide a detection system, a method for generating a trained model, a detection method, and a program that can easily estimate human behavior within a detection range within a facility. [Means for solving the problem]

[0005] A detection system according to one aspect of the present disclosure includes an acquisition unit, an estimation unit, and an output unit. The acquisition unit acquires an image of a facility captured by an imaging unit. The estimation unit uses a trained model that has been machine-learned to estimate whether a person is present within a detection range in the image acquired by the acquisition unit. and Self-moving non-living creatures and Contains 、 The output unit estimates whether or not a target object is included in the facility based on the estimation result of the estimation unit. an object included in the target object, and outputting information about a related object related to at least the person's behavior within the facility. 。

[0006] A method for generating a trained model according to one aspect of the present disclosure includes generating the trained model used in the above-described detection system by: The above-mentioned target The image including the object is generated as input data.

[0007] A detection method according to one aspect of the present disclosure includes an acquisition step, an estimation step, and an output step. The acquisition step is a step of acquiring an image of a facility captured by an imaging unit. The estimation step is a step of estimating whether a person is present within a detection range in the image acquired in the acquisition step using a trained model that has been machine-learned. and Self-moving non-living creatures and Contains 、 The output step is a step of estimating whether or not a target object is present in the facility based on the estimation result of the estimation step, an object included in the target object, and outputting information about at least a related object related to the person's behavior within the facility. 。

[0008] A program according to one embodiment of the present disclosure causes one or more processors to execute the above detection method. [Effects of the Invention]

[0009] The present disclosure has the advantage of making it easier to estimate human behavior within a detection range within a facility. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram illustrating an overview of a detection system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a schematic diagram showing an example of a detection range in the detection system. [Figure 3] FIG. 3 is a flowchart showing an example of the operation of the detection system. [Figure 4] FIG. 4 is a schematic diagram showing an example of an output from an output unit in the detection system. [Figure 5] FIG. 5 is a schematic diagram showing another example of output from the output unit in the detection system. [Figure 6] FIG. 6 is a schematic diagram showing still another example of output from the output unit in the detection system. DETAILED DESCRIPTION OF THE INVENTION

[0011] (1) Overview The detection system 100 (see FIG. 1) of this embodiment will be described below with reference to the drawings. However, the following embodiment is merely a part of various embodiments of the present disclosure. The following embodiment can be modified in various ways depending on the design, etc., as long as the object of the present disclosure can be achieved. Furthermore, each figure described in the following embodiment is a schematic diagram, and the ratio of the size and thickness of each component in the figure does not necessarily reflect the actual dimensional ratio.

[0012] As shown in Fig. 1, the detection system 100 of this embodiment is a system for detecting a target object A1 within a detection range P11 in a facility F1. In this disclosure, the "detection range" refers to a part or all of an image P1 (see Fig. 2) captured by an imaging unit 2 (see Fig. 1). In this embodiment, the detection range P11 is a partial area in the image P1.

[0013] In this disclosure, "facility" includes non-residential facilities such as stores, offices, factories, buildings, schools, welfare facilities, or hospitals, as well as residential facilities such as detached houses, apartment buildings, or individual dwelling units in detached houses or apartment buildings. Non-residential facilities also include theaters, movie theaters, public halls, amusement parks, complexes, department stores, hotels, inns, kindergartens, libraries, museums, art galleries, underground shopping malls, stations, and airports. Furthermore, in this disclosure, "facility" includes not only buildings (structures), but also outdoor facilities such as baseball stadiums, gardens, parking lots, grounds, and parks. In this embodiment, an office will be described as an example of facility F1.

[0014] As shown in FIG. 1, the detection system 100 includes an acquisition unit 11, an estimation unit 12, and an output unit 13.

[0015] The acquisition unit 11 acquires an image P1 of the inside of the facility F1 captured by the imaging unit 2. The imaging unit 2 is, for example, a camera installed on the ceiling or wall of the facility F1. The acquisition unit 11 acquires the image P1 by receiving a signal including the image P1 transmitted from the imaging unit 2. In this embodiment, as already mentioned, the facility F1 is an office, and therefore the image P1 is an overhead image of part or the entire office.

[0016] The estimation unit 12 uses a trained model 121 that has been machine-learned to estimate whether or not a target object A1 is included in a detection range P11 in an image P1 acquired by the acquisition unit 11. Specifically, in the estimation unit 12, the image P1 is input to the trained model 121. Then, the trained model 121 estimates, based on the input image P1, whether or not the target object A1 is included in a detection range P11 in the image P1.

[0017] In this embodiment, the object A1 that can be estimated by the trained model 121 is set in advance at the machine learning stage. Of course, when the trained model 121 is retrained, the object A1 that can be estimated by the trained model 121 may be added.

[0018] The output unit 13 outputs information about a related object B1 that exists within the facility F1 and is related to at least the behavior of the person A11 within the facility F1, based on the estimation result of the estimation unit 12. The "related object" in the present disclosure may include, for example, a carried item A121 that the person A11 takes out of the detection range P11. Note that the related object B1 may be the person A11 itself, as long as it is related to the behavior of the person A11.

[0019] Here, the object A1 whose presence or absence is estimated by the estimation unit 12 and the related object B1 to be output by the output unit 13 may be the same or different from each other. That is, the output unit 13 may output information regarding the person A11 as the related object B1 based on the estimation result of the estimation unit 12 regarding the presence or absence of the person A11, or may output information regarding the person A11 as the related object B1 based on the estimation result of the estimation unit 12 regarding the presence or absence of an object A1 other than the person A11.

[0020] As described above, in this embodiment, whether or not the target object A1 is included in the detection range P11 is estimated using the trained model 121. Therefore, in this embodiment, compared to when the trained model 121 is not used, it is expected that the accuracy of estimating the presence or absence of the object A1 in the detection range P11 can be improved, which has the advantage of making it easier to estimate the behavior of the person A11 in the detection range P11 within the facility F1.

[0021] (2)Details The detection system 100 of this embodiment will be described in detail below with reference to the drawings. In this embodiment, the detection system 100 detects each of a plurality of facilities F1 (here, offices). That is, a corresponding imaging unit 2 is installed in each facility F1. For simplicity of explanation, the following description will focus on one facility F1 among the plurality of facilities F1. The following description also applies to the other facilities F1 in the same way.

[0022] As shown in FIG. 1, the detection system 100 includes an acquisition unit 11, an estimation unit 12, an output unit 13, and a setting unit .

[0023] In this embodiment, at least a part of the detection system 100 is realized by a computer system having one or more processors and a memory. The one or more processors execute a program stored in the memory, causing the computer system to function as at least a part of the detection system 100. Here, the program is pre-recorded in the memory, but it may also be provided via a telecommunications line such as the Internet or recorded on a non-transitory recording medium such as a memory card.

[0024] The acquisition unit 11 acquires an image P1 of the facility F1 captured by the imaging unit 2. The acquisition unit 11 is responsible for executing an acquisition step ST1 (see FIG. 3 ), which will be described later. In this embodiment, the acquisition unit 11 is configured to be able to communicate with the imaging unit 2. In the present disclosure, "being able to communicate" means being able to exchange signals directly or indirectly via a network or a repeater, etc., using an appropriate communication method, such as wired communication or wireless communication. The image P1 is transmitted to the acquisition unit 11 via a communication interface provided in the imaging unit 2. The communication interface is connectable to a network, such as the Internet or a LAN (Local Area Network), and has a function of transmitting the image P1 to the acquisition unit 11 via the network. The communication protocol of the communication interface can be selected from various well-known wireless communication standards, such as Wi-Fi (registered trademark).

[0025] The imaging unit 2 is a camera having an imaging element for capturing an image of a subject. Here, the subject is a specific area within the facility F1. As an example, the imaging unit 2 captures a part or all of a room within the facility F1 as its subject. The imaging element is, for example, a two-dimensional image sensor such as a CCD (Charge Coupled Device) image sensor or a CMOS (Complementary Metal-Oxide Semiconductor) image sensor. The imaging unit 2 captures an image P1 by forming an image of light from the subject on the imaging surface (light-receiving surface) of the imaging element using an optical system such as a lens, and converting the light from the subject into an electrical signal using the imaging element. The imaging unit 2 then transmits the captured image P1 to the acquisition unit 11 via a communication interface.

[0026] In this embodiment, the imaging unit 2 periodically captures an image of the subject (for example, every one to several seconds) and transmits the captured image P1 to the acquisition unit 11. Therefore, in this case, the acquisition unit 11 periodically acquires the image P1. Alternatively, the imaging unit 2 may capture an image when it detects movement of the object A1 in the detection range P11 and transmit the captured image P1 to the acquisition unit 11.

[0027] The estimation unit 12 uses the machine-learned trained model 121 to estimate whether or not a target object A1 is included in the detection range P11 in the image P1 acquired by the acquisition unit 11. The estimation unit 12 is an entity that executes an estimation step ST2 (see FIG. 3) described later.

[0028] The trained model 121 may include, for example, a model using a neural network or a model generated by deep learning using a multilayer neural network. The neural network may include, for example, a convolutional neural network (CNN) or a Bayesian neural network (BNN). The trained model 121 is realized by implementing a trained neural network in an integrated circuit such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0029] In this embodiment, when the trained model 121 is used, the estimation unit 12 performs estimation based on a still image as the image P1. In other words, the image P1 acquired by the acquisition unit 11 is a still image of a subject captured at an arbitrary time point, and is not a continuous image (video). Therefore, the estimation unit 12 estimates whether or not the target object A1 is included in the detection range P11 in the image P1, based on the image P1 (still image) acquired by the acquisition unit 11. In other words, in this embodiment, when the trained model 121 is used, the estimation unit 12 performs estimation based on the image P1 (still image) at an arbitrary time point, and does not estimate images P1 before and after that time point. Note that the still image may include a so-called frame-by-frame image, in which one frame is played every second, for example.

[0030] The trained model 121 attempts to extract the feature amounts of the target object A1 in the detection range P11 in the image P1 input as input information. If the trained model 121 extracts the feature amounts of the object A1 in the detection range P11, it estimates that the object A1 is present in the detection range P11. On the other hand, if the trained model 121 does not extract the feature amounts of the object A1 in the detection range P11, it estimates that the object A1 is not present in the detection range P11.

[0031] In this embodiment, the trained model 121 only needs to be able to estimate the type of target object A1. For example, if the target object A1 is a desk, the trained model 121 only needs to be able to estimate that something resembling a desk is present in the detection range P11, and does not need to be able to estimate detailed information about the desk (e.g., design, manufacturer, etc.). Also, for example, if the target object A1 is a person A11, the trained model 121 only needs to be able to estimate that something resembling the person A11 is present in the detection range P11, and does not need to be able to estimate detailed information about the person A11 (e.g., appearance, gender, personal name, etc.).

[0032] Furthermore, in this embodiment, the trained model 121 is capable of extracting feature quantities of one or more types of object A1 from the image P1. That is, the estimation unit 12 is capable of not only estimating one type of object A1 in the detection range P11, but also estimating multiple types of objects A1 in the detection range P11. For example, assume that the objects A1 targeted by the estimation unit 12 are a desk and a chair, and that the desk and chair exist in the detection range P11. In this case, the estimation unit 12 is capable of distinguishing between the desk and the chair, rather than estimating only one type of object, either a desk or a chair, and estimating both.

[0033] The types of object A1 targeted by the estimation unit 12 are listed below. In this embodiment, the estimation unit 12 is capable of estimating all of the types of object A1 listed below. Of course, the estimation unit 12 does not have to be able to estimate all of the types of object A1 listed below, as long as it is capable of estimating at least one type of object A1.

[0034] The object A1 targeted by the estimation unit 12 may include a non-living object. In other words, the estimation unit 12 estimates whether or not the detection range P11 includes a non-living object A12 as the object A1. In the present disclosure, a "non-living object" may include a moving object that moves automatically, such as a robot vacuum cleaner, or a non-moving object. In the present disclosure, a "non-moving object" may include an installed object, such as a pillar, that basically does not move even when force is applied by a person A11, or a passive installed object, such as a fixture, that moves when force is applied by a person A11. In other words, the estimation unit 12 estimates whether or not the detection range P11 includes an installed object as a non-living object A12.

[0035] Furthermore, the passive installed object may include a carried object A121 carried by the person A11, such as a laptop personal computer or a cup. That is, the estimation unit 12 estimates whether or not the detection range P11 includes a carried object A121 that can be carried by the person A11 as an installed object. In addition, the passive installed object may include a non-carried object that is not carried by the person A11, such as a desk, chair, or shelf.

[0036] The object A1 targeted by the estimation unit 12 may include a living thing such as a person A11. In other words, the estimation unit 12 estimates whether or not the person A11 is included as the object A1 in the detection range P11.

[0037] Furthermore, in this embodiment, the estimation unit 12 has a function of estimating the position of the object A1. Specifically, when the detection range P11 includes the feature amount of the target object A1, the estimation unit 12 can recognize the position (coordinates) of the feature amount of the object A1 in the image P1. Therefore, the estimation unit 12 can estimate the position of the object A1 in the image P1. Here, when position (coordinate) data of the detection range P11 in the real space is provided to the estimation unit 12 in advance, the estimation unit 12 can also calculate and estimate the position of the object A1 in the real space based on this position data and the estimated position of the object A1 in the image P1.

[0038] Here, the trained model 121 is generated by machine learning using a large amount of training data. In this embodiment, the trained model 121 is generated by supervised learning. Therefore, the training data is a data set that combines input information input to the trained model 121 and labels assigned to the input information. Specifically, the input information of the training data is an image P1 including a target object A1. Furthermore, the label of the training data is the type of the target object A1. In other words, the method of generating the trained model 121 is a method of generating the trained model 121 to be used in the detection system 100 using an image P1 including the target object A1 as input data.

[0039] The following describes the learning phase in which the trained model 121 is generated by machine learning before the detection system 100 is used. Machine learning in the learning phase is performed, for example, in a learning center. In the learning center, machine learning of the neural network is performed using one or more processors. Before performing machine learning, the weighting coefficients of the neural network are initialized. The term "processor" here may include, for example, a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), as well as a dedicated processor specialized for calculations in the neural network.

[0040] For each of the multiple training data, one or more processors input an image P1 to the input layer of the neural network and perform a calculation. Then, the one or more processors perform backpropagation (error backpropagation) processing using the output values ​​of multiple neurons in the output layer of the neural network and the labels. Here, the multiple neurons in the output layer correspond to multiple types of objects A1, respectively. In the backpropagation processing, the one or more processors update the weighting coefficients of the neural network so as to maximize the output values ​​of neurons in the output layer that correspond to the labels. The one or more processors optimize the weighting coefficients of the neural network by performing the backpropagation processing on all of the training data. The backpropagation processing on all of the training data can be performed not just once, but multiple times until the output values ​​of neurons corresponding to the labels become sufficiently large (i.e., until the error function converges). This completes the learning of the neural network, and a trained model 121 is generated.

[0041] The output unit 13 outputs information about the related object B1 based on the estimation result of the estimation unit 12. The output unit 13 is an entity that executes an output step ST3 (see FIG. 3 ) described later. In this embodiment, the related object B1 is a person A11 (here, an employee working in the office) who acts within the facility F1. That is, in this embodiment, the output unit 13 outputs information about the person A11 as the related object B1 based on the estimation result of the estimation unit 12.

[0042] In this embodiment, the output unit 13 outputs information about the person A11 as the related object B1 depending on whether the estimation result of the estimation unit 12 matches any of a plurality of predetermined patterns. Examples of patterns include the presence or absence of the person A11 in the detection range P11, the presence or absence of an installed object (including a carried object A121), a combination of the person A11 and the installed object, or the positional relationship between the person A11 and the installed object. Here, "matching a pattern" includes a complete match with the pattern, as well as a partial match with the pattern but with a higher degree of match than other patterns. The patterns referenced by the output unit 13 may be added or changed by the user as appropriate.

[0043] Below, examples of output of information relating to related object B1 (here, person A11) by the output unit 13 are listed. In this embodiment, the output unit 13 is capable of outputting all of the output examples shown below. Of course, the output unit 13 does not have to be able to output all of the output examples shown below, as long as it is capable of outputting at least one type of output example. Furthermore, the output unit 13 may be configured to output information other than the output examples listed below.

[0044] 4, it is assumed that a desk and a chair (non-living object A12) are present in the detection range P11, a laptop personal computer (portable object A121) is placed on the desk, and a person A11 (related object B1) is sitting in the chair. In this case, the output unit 13 outputs information that the person A11 is sitting at the desk, based on the estimation result of the estimation unit 12.

[0045] 5, it is assumed that a desk and a chair are present in the detection range P11, and that a laptop personal computer is placed on the desk, but that a person A11 is not sitting in the chair. In this case, the output unit 13 outputs information that the person A11 who is using the desk is away from his / her seat, based on the estimation result of the estimation unit 12. In this case, the output unit 13 also outputs information that the person A11 has left his / her belonging A121, the laptop personal computer, on the desk, and therefore has been away from his / her seat for a relatively short time.

[0046] 6, it is assumed that a desk and a chair are present in the detection range P11, but that person A11 is not sitting in the chair and that no laptop personal computer is placed on the desk. In this case, the output unit 13 outputs information that person A11, who is using the desk, is away from his / her seat, based on the estimation result of the estimation unit 12. In this case, the output unit 13 also outputs information that person A11 has been away from his / her seat for a relatively long time, since the laptop personal computer, which is a carried item A121, is not left on the desk.

[0047] Also, as an example, assume that the person A11 is not sitting in a chair in the detection range P11. In this case, the output unit 13 outputs information that the person A11 is moving, based on the estimation result of the estimation unit 12. In this case, if the estimation unit 12 has also estimated the position and orientation of the person A11, the output unit 13 may also output information regarding the destination of the person A11.

[0048] Also, as an example, the output unit 13 outputs information regarding the layout of the detection range P11. For example, assume that the estimation unit 12 estimates that installed objects, such as desks and chairs, exist in the detection range P11 as non-living objects A12, and estimates the positions of the installed objects. In this case, the output unit 13 generates and outputs information regarding the layout of the detection range P11 by adding information about the estimated installed objects to map information representing a floor plan of the detection range P11 based on the estimation result of the estimation unit 12. The map information is, for example, BIM (Building Information Modeling) data of the facility F1, and is input to the detection system 100 in advance. Also, the map information may be generated based on the image P1 captured by the imaging unit 2.

[0049] Also, as an example, the output unit 13 outputs information regarding the removal status of the carried item A121 from the detection range P11 based on the estimation result of the estimation unit 12. For example, it is assumed that the estimation unit 12 estimates that the carried item A121, such as a laptop personal computer, is present in the detection range P11. In this case, the output unit 13 outputs information that the carried item A121 has not been removed from the detection range P11 based on the estimation result of the estimation unit 12. It is also assumed that the estimation unit 12 estimates that the carried item A121, such as a laptop personal computer, is not present in the detection range P11. In this case, the output unit 13 outputs information that the carried item A121 has been removed from the detection range P11 based on the estimation result of the estimation unit 12.

[0050] The destination of the information output by the output unit 13 may include, for example, an information terminal used by the user or a management system that manages the facility F1. The information terminal may include, for example, a smartphone, a tablet terminal, or a personal computer. The user or the administrator of the management system can understand the status of the related object B1 (here, person A11) by viewing the information output by the output unit 13 on the information terminal. For example, if the user is an office manager, the user can understand the working status of the office employees by viewing the information output by the output unit 13.

[0051] Furthermore, the destination of the information output by the output unit 13 may include, for example, a control system that controls devices installed in the facility F1, such as a BEMS (Building and Energy Management System).

[0052] The equipment may include, for example, facility equipment. Examples of facility equipment include package air conditioners (air conditioning equipment), lighting equipment (including base lights and spotlights), power storage equipment, kitchen equipment (including induction heaters and dishwashers), access control equipment, copy machines, and facsimiles. Furthermore, facility equipment also includes, for example, home appliances such as hot water supply equipment (including EcoCute (registered trademark)), electric shutters, ventilation fans, and 24-hour ventilation systems.

[0053] Furthermore, if the control system is, for example, a Home Energy Management System (HEMS), the devices may include home appliances. Examples of home appliances include television sets, lighting equipment (including ceiling lights), and recorders / players (including DVD recorders with HDDs and external HDDs). Furthermore, home appliances also include washing machines, refrigerators, air conditioners, air purifiers, personal computers, smart speakers, and computer game consoles.

[0054] For example, when the control system receives from the output unit 13 information that person A11 has been leaving the detection range P11 for a long period of time, it turns off the lighting fixtures installed in the detection range P11 (particularly, a certain range including the seat of person A11). Also, when the control system receives from the output unit 13 information that person A11 has only temporarily left the detection range P11, it keeps the lighting fixtures installed in the detection range P11 on instead of turning them off.

[0055] The setting unit 14 sets an algorithm to be used by the estimation unit 12 according to at least the installation environment of the imaging unit 2. In this embodiment, the estimation unit 12 basically estimates whether or not the target object A1 is included in the detection range P11 using the trained model 121. However, depending on the height at which the imaging unit 2 is installed (height relative to the floor surface of the facility F1), for example, the object A1 that may be included in the image P1 captured by the imaging unit 2 may appear small, which may make it difficult for the trained model 121 to recognize the object A1.

[0056] Therefore, in this embodiment, the setting unit 14 sets the algorithm used by the estimation unit 12 according to the installation environment of the imaging unit 2 (here, the height at which the imaging unit 2 is installed). As an example, the height at which the imaging unit 2 is installed can be measured if the imaging unit 2 is a stereo camera or a ToF (Time of Flight) camera and can acquire depth information of the subject in the detection range P11.

[0057] If the height at which the imaging unit 2 is installed is less than a predetermined height, the setting unit 14 sets the algorithm used by the estimation unit 12 to an algorithm that uses the trained model 121. On the other hand, if the height at which the imaging unit 2 is installed is equal to or greater than the predetermined height, the setting unit 14 sets the algorithm used by the estimation unit 12 to another algorithm that does not use the trained model 121. As an example, the other algorithm obtains a difference image between images P1 captured by the imaging unit 2 at two consecutive points in time, and estimates whether or not the target object A1 is included in the detection range P11 based on whether or not the pattern of the object A1 is included in advance in the difference image.

[0058] (3) Operation An example of the operation of the detection system 100 will be described below with reference to Fig. 3. First, the acquisition unit 11 acquires the image P1 (ST1) by receiving the image P1 transmitted from the imaging unit 2. The process ST1 corresponds to an acquisition step ST1.

[0059] Next, the estimation unit 12 uses the image P1 acquired by the acquisition unit 11 as input data and estimates whether or not the target object A1 is included in the detection range P11 using the trained model 121 (ST2). Process ST2 corresponds to estimation step ST2.

[0060] Then, the output unit 13 outputs information about the related object B1 (here, person A11) based on the estimation result of the estimation unit 12 (ST3). Process ST3 corresponds to output step ST3. Thereafter, the detection system 100 executes the above processes ST1 to ST3 every time an image P1 is transmitted from the imaging unit 2.

[0061] (4) Advantages The advantages of the detection system 100 of this embodiment will be described below, along with a comparison with the image sensor described in Patent Document 1. The image sensor described in Patent Document 1 identifies whether or not a specific object is moving through three stages of processing: processing by an extraction unit, processing by a motion identification unit, and processing by an object identification unit. The extraction unit extracts features from image data capturing an image of the target space. The motion identification unit identifies the presence or absence of a moving object in the target space based on the features extracted by the extraction unit. The object identification unit identifies whether or not an object in the target space is a pre-registered object based on the features extracted by the extraction unit.

[0062] As described above, the image sensor described in Patent Document 1 undergoes three stages of processing, whereas in this embodiment, it is possible to estimate whether or not the target object A1 is included in the detection range P11 through one stage of processing using the trained model 121.

[0063] (5) Variations The above-described embodiment is merely one of various embodiments of the present disclosure. The above-described embodiment can be modified in various ways depending on the design, etc., as long as the object of the present disclosure can be achieved. Furthermore, functions similar to those of the detection system 100 may be embodied in a detection method, a (computer) program, a non-transitory recording medium on which a program is recorded, or the like.

[0064] A detection method according to one embodiment of the present disclosure includes an acquisition step ST1, an estimation step ST2, and an output step ST3. The acquisition step ST1 is a step of acquiring an image P1 of a facility F1 captured by the imaging unit 2. The estimation step ST2 is a step of estimating, using a machine-learned trained model 121, whether or not a target object A1 is included in a detection range P11 in the image P1 acquired in the acquisition step ST1. The output step ST3 is a step of outputting information about a related object B1 that exists in the facility F1 and is related to at least the behavior of a person within the facility F1, based on the estimation result of the estimation step ST2. A program according to one embodiment of the present disclosure causes one or more processors to execute the above detection method.

[0065] The following are examples of modifications of the above-described embodiment. The modifications described below can be applied in appropriate combinations.

[0066] The detection system 100 according to the present disclosure includes a computer system, for example, in the estimation unit 12. The computer system is primarily composed of a processor and memory as hardware. The functions of the detection system 100 according to the present disclosure are realized by the processor executing a program stored in the memory of the computer system. The program may be pre-stored in the memory of the computer system, provided via a telecommunications line, or provided in a non-transitory recording medium readable by the computer system, such as a memory card, optical disk, or hard disk drive. The processor of the computer system is composed of one or more electronic circuits, including a semiconductor integrated circuit (IC) or a large-scale integrated circuit (LSI). The integrated circuits, such as ICs and LSIs, are referred to by different names depending on the degree of integration, and include integrated circuits called system LSIs, very large-scale integrations (VLSIs), or ultra-large-scale integrations (ULSIs). Furthermore, field-programmable gate arrays (FPGAs), which are programmable after the LSI is manufactured, or logic devices capable of reconfiguring the connections within the LSI or the circuit partitions within the LSI, can also be used as processors. The electronic circuits may be integrated into one chip or distributed across multiple chips. The chips may be integrated into one device or distributed across multiple devices. The computer system referred to here includes a microcontroller having one or more processors and one or more memories. Therefore, the microcontroller is also composed of one or more electronic circuits including a semiconductor integrated circuit or a large-scale integrated circuit.

[0067] Furthermore, it is not essential for the detection system 100 that multiple functions are integrated into one housing. The components of the detection system 100 may be distributed across multiple housings. Furthermore, at least some of the functions of the detection system 100 may be realized by, for example, a server device and the cloud (cloud computing), etc.

[0068] In the above-described embodiment, there is one trained model 121, but there may be multiple trained models 121. For example, the trained models 121 may be two models: a model for estimating a non-living object A12 and a model for estimating a person A11. In this case, the model for estimating the non-living object A12 is generated by machine learning using, as training data, an image P1 including the non-living object A12 and an image P1 not including the non-living object A12. Furthermore, the model for estimating the person A11 is generated by machine learning using, as training data, an image P1 including the person A11 and an image P1 not including the person A11. In this way, when the estimation unit 12 uses multiple trained models 121, improvement in estimation accuracy of the object A1 can be expected compared to when the estimation unit 12 uses a single trained model 121.

[0069] In the above-described embodiment, the trained model 121 may be machine-learned by unsupervised learning or by reinforcement learning. Furthermore, the trained model 121 may be retrained after the detection system 100 is applied to the facility F1. Whether or not to retrain the trained model 121 may be selectable by the user, for example, via the setting unit 14. In this case, the setting unit 14 may have an interface that accepts user operation input.

[0070] In the above-described embodiment, the estimation unit 12 performs estimation based on a still image as the image P1, but this is not limiting. For example, the estimation unit 12 may perform estimation based on a still image at an arbitrary time point and still images before and after the time point as the image P1. In this embodiment, the trained model 121 may include a model using a neural network such as an RNN (Recurrent Neural Network). Furthermore, for example, the estimation unit 12 may perform estimation based on a difference image as the image P1 (i.e., an image representing the difference between still images at two arbitrary times).

[0071] In the above-described embodiment, the output unit 13 may output information regarding the three-dimensional layout of the detection range P11. In this case, it is preferable that the imaging unit 2 is, for example, a stereo camera or a ToF camera, and is capable of acquiring depth information of the subject in the detection range P11.

[0072] In the above-described embodiment, the detection system 100 may not include the setting unit 14. In this case, the estimation unit 12 estimates whether or not the target object A1 is included in the detection range P11 by a pre-set algorithm (here, processing using the trained model 121).

[0073] In the above-described embodiment, the imaging unit 2 is not included in the components of the detection system 100, but it may be included in the components of the detection system 100. Also, in the above-described embodiment, the imaging unit 2 and the detection system 100 are configured as separate entities, but this is not limited to this. For example, all of the components of the detection system 100 may be included inside the housing that configures the imaging unit 2. In this configuration, imaging and processing of the detection system 100 can be performed in the housing that configures the imaging unit 2, which has the advantage that the detection system 100 can be installed more easily than when the imaging unit 2 and the detection system 100 are separate entities.

[0074] (summary) As described above, the detection system (100) according to the first aspect includes an acquisition unit (11), an estimation unit (12), and an output unit (13). The acquisition unit (11) acquires an image (P1) of the inside of a facility (F1) captured by the imaging unit (2). The estimation unit (12) estimates, using a machine-learned trained model (121), whether or not a target object (A1) is included in a detection range (P11) in the image (P1) acquired by the acquisition unit (11). The output unit (13) outputs, based on the estimation result of the estimation unit (12), information about a related object (B1) that exists within the facility (F1) and is related to at least the behavior of a person (A11) within the facility (F1).

[0075] This embodiment has the advantage that it is easy to estimate the behavior of the person (A11) in the detection range (P11) within the facility (F1).

[0076] In the detection system (100) according to the second aspect, in the first aspect, the estimation unit (12) performs estimation based on a still image as the image (P1).

[0077] This aspect has the advantage that the processing load required for estimation can be easily reduced compared to estimation based on difference images such as moving images.

[0078] In the detection system (100) according to the third aspect, in the first or second aspect, the estimation unit (12) estimates the position of the object (A1).

[0079] This embodiment has the advantage that it is easier to manage the state of the object (A1) compared to when the position of the object (A1) is not estimated.

[0080] In the detection system (100) according to the fourth aspect, in any of the first to third aspects, the estimation unit (12) estimates whether or not a non-living object (A12) is included as an object (A1) in the detection range (P11).

[0081] This embodiment has the advantage that it is expected to make it easier to estimate the behavior of the person (A11) compared to when the non-living thing (A12) is not estimated.

[0082] In the detection system (100) according to the fifth aspect, in the fourth aspect, the estimation unit (12) estimates whether or not an installed object as a non-living object (A12) is included in the detection range (P11). The output unit (13) outputs information regarding the layout of the detection range (P11).

[0083] This embodiment has the advantage that it becomes easier to visually grasp the state of the detection range (P11).

[0084] In the detection system (100) according to the sixth aspect, in the fifth aspect, the estimation unit (12) estimates whether or not a portable object (A121) that can be carried by a person (A11) is included as an installed object within the detection range (P11).

[0085] This embodiment has the advantage that it is expected to make it easier to estimate the behavior of the person (A11) compared to when the carried object (A121) is not estimated.

[0086] In the detection system (100) according to the seventh aspect, in the sixth aspect, the output unit (13) outputs information regarding the removal status of the carried item (A121) from the detection range (P11) based on the estimation result of the estimation unit (12).

[0087] This embodiment has the advantage that the belongings (A121) can be easily managed.

[0088] In the detection system (100) according to the eighth aspect, in any one of the first to seventh aspects, the estimation unit (12) estimates whether or not a person (A11) is included as the object (A1) in the detection range (P11).

[0089] According to this aspect, there is an advantage that the presence or absence of the person (A11) itself in the detection range (P11) is estimated, thereby making it easier to estimate the behavior of the person (A11).

[0090] In the detection system (100) according to the ninth aspect, in the eighth aspect, the output unit (13) outputs the behavior of the person (A11) based on the estimation result of the estimation unit (12).

[0091] This embodiment has the advantage that it becomes easier to understand the behavior of the person (A11).

[0092] The detection system (100) according to a tenth aspect is the detection system according to any one of the first to ninth aspects, further including a setting unit (14). The setting unit (14) sets an algorithm to be used in the estimation unit (12) in accordance with at least the installation environment of the imaging unit (2).

[0093] According to this aspect, the estimation unit (12) can perform estimation using an algorithm suited to the installation environment of the imaging unit (2), which has the advantage that the estimation accuracy of the estimation unit (12) can be expected to improve.

[0094] A method for generating a trained model (121) according to the eleventh aspect generates a trained model (121) to be used in a detection system (100) according to any one of the first to tenth aspects, using an image (P1) including a target object (A1) as input data.

[0095] This embodiment has the advantage of easily generating a trained model (121) that can easily estimate the behavior of a person (A11) in a detection range (P11) within a facility (F1).

[0096] A detection method according to a twelfth aspect includes an acquisition step (ST1), an estimation step (ST2), and an output step (ST3). The acquisition step (ST1) is a step of acquiring an image (P1) of a facility (F1) captured by an imaging unit (2). The estimation step (ST2) is a step of estimating, using a machine-learned trained model (121), whether or not a target object (A1) is included in a detection range (P11) in the image (P1) acquired in the acquisition step (ST1). The output step (ST3) is a step of outputting, based on the estimation result of the estimation step (ST2), information about a related object (B1) that exists in the facility (F1) and is related to at least the behavior of a person within the facility (F1).

[0097] This embodiment has the advantage that it is easy to estimate the behavior of the person (A11) in the detection range (P11) within the facility (F1).

[0098] A program according to a thirteenth aspect causes one or more processors to execute the detection method according to the twelfth aspect.

[0099] This embodiment has the advantage that it is easy to estimate the behavior of the person (A11) in the detection range (P11) within the facility (F1).

[0100] The configurations according to the second to tenth aspects are not essential for the detection system (100) and can be omitted as appropriate. [Explanation of symbols]

[0101] 100 Detection System 11 Acquisition Department 12 Estimation part 121 trained models 13 Output section 14 Setting section 2. Imaging unit F1 Facilities P1 Image P11 Detection range A1 Object A11 people A12 Abiotic A121 Carry-on Items ST1 Acquisition step ST2 Estimation step ST3 Output Step

Claims

1. an acquisition unit that acquires images of the inside of the facility captured by the imaging unit; An estimation unit that uses a machine-learned model to estimate whether or not a target object, including a person and an automatically moving non-living object, is included in the detection range in the image acquired by the acquisition unit; and an output unit that outputs, based on the estimation result of the estimation unit, information about a related object that is present within the facility, is included in the target object, and is related to at least the behavior of the person within the facility; Detection system.

2. the estimation unit performs estimation based on a still image as the image, The detection system of claim 1 .

3. The estimation unit estimates a position of the target object.

3. The detection system according to claim 1 or 2.

4. the estimation unit estimates whether the non-living object is included in the detection range as the target object. The detection system according to any one of claims 1 to 3.

5. The target object further includes an installation, the estimation unit estimates whether the installed object is included in the detection range; the output unit outputs information regarding the layout of the detection range. The detection system of claim 4.

6. The estimation unit estimates whether or not a portable object that can be carried by the person is included as the installed object within the detection range. The detection system of claim 5.

7. the output unit outputs information regarding a state in which the carried item has been taken out of the detection range based on the estimation result of the estimation unit. The detection system of claim 6.

8. the output unit outputs information about the person as the related object based on the estimation result of the estimation unit. The detection system according to any one of claims 1 to 7.

9. further comprising a setting unit that sets an algorithm to be used by the estimation unit in accordance with at least an installation environment of the imaging unit; The detection system according to any one of claims 1 to 8.

10. The trained model used in the detection system according to any one of claims 1 to 9 is generated using the image including the target object as input data. How to generate a trained model.

11. an acquisition step of acquiring an image of the inside of the facility captured by the imaging unit; an estimation step of estimating whether or not a target object, including a person or an automatically moving non-living object, is included in the detection range in the image acquired in the acquisition step using a trained model that has been machine-learned; and an output step of outputting information about a related object that is present in the facility, is included in the target object, and is related to at least the behavior of the person in the facility, based on the estimation result of the estimation step. Detection method.

12. one or more processors, Executing the detection method according to claim 11, program.

Citation Information

Patent Citations

  • Watching support system and method for controlling the same

    JP2019008515A

  • Image sensor, identification method, control system and program

    JP2019204165A

  • Information processing device, information processing system, information processing method, and program

    WO2020116023A1

  • Semiconductor device

    WO2020208995A1