Analysis system and analysis program

The analysis system improves the accuracy of learning and inference in complex environments by generating and using shift features from inclusion features over time, effectively addressing the challenges of dense crowds and obstructions.

WO2025115364A1PCT designated stage expired Publication Date: 2025-06-05HITACHI LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/034026
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-01
Filing Date
2024-09-25
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing analysis systems struggle to accurately analyze human movement and shape information in environments with dense crowds or many obstructions, leading to difficulties in learning or inferring unsteady states.

Method used

The proposed analysis system employs a processor that executes an acquisition process to repeatedly acquire inclusion features, a generation process to generate shift features by combining inclusion features over time, and an output process to output these shift features to an inference model, thereby improving learning and inference accuracy.

Benefits of technology

This approach enhances the learning accuracy and inference accuracy of the inference model, enabling it to effectively infer the state of a space even in complex environments with dense crowds or obstructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024034026_05062025_PF_FP_ABST
    Figure JP2024034026_05062025_PF_FP_ABST
Patent Text Reader

Abstract

This analysis system has a processor that executes a program and a storage device that stores the program, wherein the processor executes: an acquisition process for repeatedly acquiring a containment feature value which contains information pertaining to the motion of an inference target and information pertaining to the direction of the inference target; a generation process for generating a shift feature value group, from a shift feature value at a first time to a shift feature value at a second time, by combining, with each containment feature value from the containment feature value at the first time to the containment feature value at the second time among the plurality of containment feature values repeatedly acquired by the acquisition process, the containment feature value at a third time prior to a prescribed time; and an output process for outputting the shift feature value group generated by the generation process to an inference model which infers the state of the inference target in the period from the first time to the second time.
Need to check novelty before this filing date? Find Prior Art

Description

Analysis systems and analysis programs Incorporation by Reference

[0001] This application claims priority from Japanese Patent Application No. 2023-203745, filed on December 1, 2023, the contents of which are incorporated herein by reference.

[0002] The present invention relates to an analysis system and an analysis program for analyzing data.

[0003] As background art in this technical field, Patent Document 1 discloses a human detection unit and an object detection unit that detect people and objects, a feature inference unit that infers feature amounts based on human movement and shape information detected by the human detection unit and object movement information detected by the object detection unit, a steady state model storage unit that holds a steady state model in advance, a steady state model inference unit that learns steady state feature amounts using the feature inference unit and the steady state model storage unit, a non-steady state inference unit that determines deviation from the steady state model using the feature amount calculated by the feature inference unit and data in the steady state model storage unit, and infers a non-steady state by the non-steady state inference unit.

[0004] Japanese Patent Application Laid-Open No. 2020-149389

[0005] However, the technology disclosed in Patent Document 1 above is unable to accurately analyze human movement and shape information in environments with dense crowds or many obstructions, making it difficult to learn or infer unsteady states.

[0006] The present invention has been made in consideration of the above problems, and aims to improve the learning accuracy and inference accuracy of an inference model that infers the state of a space.

[0007] An analysis system that is a first disclosed technology of the present application is an analysis system having a processor that executes a program and a storage device that stores the program, wherein the processor executes an acquisition process that repeatedly acquires inclusion features that include information about the movement of an inference object and information about the direction of the inference object; a generation process that generates a group of shift features from the shift feature at the first time to the shift feature at the second time by combining an inclusion feature at a third time that is a predetermined time before for each of the multiple inclusion features repeatedly acquired by the acquisition process, from the inclusion feature at a first time to the inclusion feature at a second time; and an output process that outputs the group of shift features generated by the generation process to an inference model that infers the state of the inference object during the period from the first time to the second time.

[0008] According to a representative embodiment of the present invention, the problems, configurations, and effects other than those described above, which aim to improve the learning accuracy and inference accuracy of an inference model that infers the state of space, will be made clear through the description of the following examples.

[0009] FIG. 1 is an explanatory diagram illustrating an example of a system configuration of an analysis system according to a first embodiment. FIG. 2 is a block diagram illustrating an example of a hardware configuration of a computer. FIG. 3 is a block diagram illustrating an example of a functional configuration of the analysis system according to the first embodiment. FIG. 4 is an explanatory diagram illustrating an example of generation of a trajectory ID and a movement vector. FIG. 5 is an explanatory diagram illustrating a speed distribution and a direction distribution. FIG. 6 is an explanatory diagram illustrating an example of calculation of a histogram feature amount at time t. FIG. 7 is an explanatory diagram illustrating an example of generation of a shift feature amount. FIG. 8 is an explanatory diagram illustrating an example of a method for saving learning data. FIG. 9 is an explanatory diagram illustrating an example of space analysis by the analysis system. FIG. 10 is a flowchart illustrating an example of detailed processing procedures of a learning process by a server (learning device) according to the first embodiment. FIG. 11 is a flowchart illustrating an example of a state inference processing procedure by a client (inference device) according to the first embodiment. FIG. 12 is a flowchart illustrating an example of a state notification processing procedure by a terminal (state notification device) according to the first embodiment. FIG. 13 is a block diagram illustrating an example of a functional configuration of an analysis system according to a second embodiment. FIG. 14 is a block diagram illustrating an example of a functional configuration of an analysis system according to a third embodiment. FIG. 15 is a flowchart illustrating a detailed example of a procedure for a learning process by a server (learning device) according to a third embodiment. FIG. 16 is a flowchart illustrating a procedure for a state inference process by a client (inference device) according to a third embodiment. FIG. 17 is a block diagram illustrating a functional configuration example of an analysis system according to a fourth embodiment. FIG. 18 is an explanatory diagram illustrating an example of time-series learning data output from a teacher signal DB. FIG. 19 is a block diagram illustrating a functional configuration example of an analysis system according to a fifth embodiment. FIG. 20 is an explanatory diagram illustrating an example of a shift feature group according to a seventh embodiment. FIG. 21 is a block diagram illustrating a functional configuration example of an analysis system according to an eighth embodiment. FIG. 22 is a flowchart illustrating a detailed example of a procedure for a learning process by a client (learning device) according to the eighth embodiment. FIG. 23 is a flowchart illustrating a procedure for a state inference process by a client (inference device) according to the eighth embodiment.

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all drawings used to describe the embodiments, identical components are generally designated by the same reference numerals, and repeated description thereof will be omitted. In the following embodiments, the components (including element steps, etc.) are not necessarily essential unless otherwise specified or considered to be clearly essential in principle. Furthermore, when the terms "consisting of A," "composed of A," "having A," or "including A" are used, it goes without saying that they do not exclude other elements, unless otherwise specified to include only the relevant element. Similarly, in the following embodiments, when referring to the shape, positional relationship, etc. of components, etc., this includes those that are substantially similar or similar to the shape, etc., unless otherwise specified or considered to be clearly essential in principle.

[0011] The terms "first," "second," "third," and the like used in this specification are used to identify components and do not necessarily limit the number, order, or content of the components. Furthermore, numbers used to identify components are used in different contexts, and a number used in one context does not necessarily indicate the same configuration in another context. Furthermore, this does not prevent a component identified by a certain number from also serving the function of a component identified by another number.

[0012] In order to facilitate understanding of the invention, the position, size, shape, range, etc. of each component shown in the drawings etc. may not represent the actual position, size, shape, range, etc. Therefore, the present invention is not necessarily limited to the position, size, shape, range, etc. disclosed in the drawings etc.

[0013] <Figure 1 Analysis System> Figure 1 is an explanatory diagram showing an example of the system configuration of an analysis system according to Example 1. The analysis system 100 includes a server 101, one or more clients 102, and a terminal 106. The server, clients, and terminals are communicatively connected via a network 105 such as the Internet, a LAN (Local Area Network), or a WAN (Wide Area Network). The server 101 is a computer that manages the clients 102. The clients 102 are connected to sensors 103 and acquire data from the sensors 103. The terminal 106 receives and transmits signals from the clients 102, and inputs and outputs operations and data to and from the server and clients.

[0014] The sensor 103 detects analysis target data from the analysis environment. The sensor 103 is, for example, a camera that captures still images or videos. The camera may be capable of acquiring depth information. The sensor 103 captures images of the analysis environment at a fixed point. Therefore, for example, some objects change in shape or size between time-series frames output as analysis target data from the sensor 103, while other objects are congruent in shape and size between frames. Furthermore, even congruent objects between frames may have different colors.

[0015] The teacher signal DB 104 is a database that stores a combination of image data (also called a frame), which is the learning data to be inspected, and the state of the learning data (for example, correct answer data representing two or more states such as "normal" and "abnormal") as a teacher signal.

[0016] The inspection target is an area whose state is inferred by fixed-point photography, such as the space inside a train, where state inference is required based on the actions of people in the space. The teacher signal DB 104 may be stored in the server 101, or may be connected to a computer that can communicate with the server 101 or the client 102 via the network 105.

[0017] Analysis system 100 has a learning function that uses teacher signal DB 104 and an inference function that uses an inference model obtained by the learning function. The inference model is a learning model that infers the state of an object shown in image data. The learning function and the inference function may be implemented in either server 101 or client 102 as long as they are implemented in analysis system 100. For example, server 101 may implement the learning function, and client 102 may implement the inference function. Alternatively, server 101 may implement the learning function and the inference function, and client 102 may transmit data from sensor 103 to server 101 and receive inference results from server 101 using the inference function.

[0018] Alternatively, the client 102 may implement a learning function and an inference function, and the server 101 may manage the inference model and inference results from the client 102. A computer that implements the learning function is referred to as a learning device, and a computer that implements at least the inference function of the learning function and the inference function is referred to as an inference device. Also, while a client-server type analysis system 100 is shown in FIG. 1 as an example, a standalone inference device may also be used. For convenience of explanation, in Example 1, an analysis system 100 in which the server 101 implements the learning function and the client 102 implements the inference function will be described as an example.

[0019] <FIG. 2: Example of Computer Hardware Configuration> FIG. 2 is a block diagram showing an example of the hardware configuration of a computer (server 101, client 102). The computer 200 has a processor 201, a storage device 202, an input device 203, an output device 204, and a communication interface (communication IF) 205. The processor 201, the storage device 202, the input device 203, the output device 204, and the communication IF 205 are connected via a bus 206. The processor 201 controls the computer 200. The storage device 202 serves as a working area for the processor 201. The storage device 202 is a non-transitory or temporary recording medium that stores various programs and data. Examples of the storage device 202 include a read-only memory (ROM), a random access memory (RAM), a hard disk drive (HDD), and a flash memory. The input device 203 inputs data. Examples of the input device 203 include a keyboard, a mouse, a touch panel, a numeric keypad, and a scanner. The output device 204 outputs data. Examples of the output device 204 include a display, a printer, and a speaker. The communication IF 205 connects to the network 105 and transmits and receives data.

[0020] <Fig. 3: Example of Functional Configuration of Analysis System 100> Fig. 3 is a block diagram showing an example of the functional configuration of the analysis system 100 according to Example 1. The server 101 includes a person detection unit 300, a tracking processing unit 301, a vector information generation unit 302, a speed distribution calculation unit 303, a direction distribution calculation unit 304, a histogram feature generation unit 305, a shift feature generation unit 306, and a learning unit 307. The client 102 includes a person detection unit 310, a tracking processing unit 311, a vector information generation unit 312, a speed distribution calculation unit 313, a direction distribution calculation unit 314, a histogram feature generation unit 315, a shift feature generation unit 316, an inference unit 317, and an inference result output unit 318. The terminal 106 includes an inference result receiving unit 320.

[0021] Specifically, these are realized by, for example, having the processor 201 execute a program stored in the storage device 202 shown in Fig. 2. In Fig. 3, the teacher signal DB 104 exists outside the server 101, but it may also be realized in the storage device 202 in the server 101 or the client 102.

[0022] <Example of Functional Configuration on the Server 101 Side> First, an example of the functional configuration on the server 101 side will be described. The person detection unit 300 generates rectangle information R = (x1, y1, x2, y2) indicating the presence of a person for time-series learning data acquired from the teacher signal DB 104, and outputs the center point of the rectangle information R as the person's position coordinates O = (x, y) to the tracking processing unit 301. If multiple people are present in the learning data, multiple position coordinates O are generated. Note that x1, y1 of the rectangle information R represent the image coordinates of the upper left corner of the rectangle information R, and x2, y2 represent the image coordinates of the lower right corner of the rectangle. The teacher signal is composed of learning data and a state. Details of the teacher signal will be described with reference to FIG. 8.

[0023] The tracking processing unit 301 generates a trajectory ID unique to a person for the position coordinates O of the person acquired from the person detection unit 300. The trajectory IDs of multiple position coordinates O originating from the same person among the acquired multiple learning data have the same value. The tracking processing unit 301 outputs the trajectory ID and the position coordinates O of the person associated with the trajectory ID to the vector information generation unit 302.

[0024] The vector information generating unit 302 generates a movement vector based on the acquired trajectory ID and position coordinates O, and outputs the movement vector to the speed distribution calculating unit 303 and the direction distribution calculating unit 304 .

[0025] [Figure 4: Example of Generating Trajectory IDs and Movement Vectors] Figure 4 is an explanatory diagram showing an example of generating trajectory IDs and movement vectors. Reference numeral 400-t denotes the t-th image data (frame) acquired as learning data. In Figure 4, reference numerals 401(t), 402(t), 401(t-u), and 402(t-u) denote rectangular information R containing image data of a person, but for convenience of explanation, these will be written as people 401(t), 402(t), 401(t-u), and 402(t-u).

[0026] The person detection unit 300 detects person 401(t) and person 402(t), and outputs their respective position coordinates O to the tracking processing unit 301. Frame 400-(t-u) is image data acquired in the t-uth frame. The person detection unit 300 detects person 401(t-u) and person 402(t-u), and outputs their respective position coordinates O to the tracking processing unit 301.

[0027] The tracking processing unit 301 assigns the same trajectory ID to the person detected in frames 400-t and 400-(tu) acquired at different times t and tu. For example, the tracking processing unit 301 considers person 401(t) and person 401(tu) to be the same person, and assigns i to person 401(t) so that it is possible to distinguish at which time the person was detected. t The person 401(tu) is assigned a trajectory ID at time t (the t-th acquired frame 400-t). t-u A trajectory ID at time t−u (frame 400-(t−u) acquired at the t−uth time) is assigned.

[0028] Similarly, the tracking processing unit 301 regards the person 402(t) and the person 402(tu) as the same person, and assigns j to the person 402(t) so that it is possible to distinguish at what time the person was detected. t The person 402(tu) is assigned a trajectory ID at time t (the t-th acquired frame 400-t). t-u A trajectory ID at time t−u (frame 400-(t−u) acquired at the t−uth time) is assigned.

[0029] In addition, in frame 400-t, the position coordinates O of the person 401(t) are O i t = (x i t , y i t ), and the position coordinates O of the person 401(tu) in the frame 400-(tu) are expressed as O i t-u = (x i t-u , y i t-u ) is expressed as

[0030] Similarly, in frame 400-t, the position coordinates O of person 402(t) are O j t = (x j t , y j t ), and the position coordinates O of the person 402(tu) in the frame 400-(tu) are expressed as O j t-u = (x j t-u , y j t-u ) is expressed as

[0031] The tracking processing unit 301 may use authentication processing to identify a person between different frames, or may use tracking processing based on the amount of movement between frames using a Kalman filter or the like, and the person identification method used by the tracking processing unit 301 is not limited. Furthermore, the specific values ​​of t and u, which indicate the order in which frames are acquired for tracking processing, are also not limited. However, since the accuracy of the tracking processing between frames is likely to deteriorate as the value of u increases, it is preferable to set the value of u to a value smaller than a predetermined value.

[0032] The composite frame 410 is image data obtained by superimposing the frame 400-t and the frame 400-(tu). In the composite frame 410, the vector information generating unit 302 calculates the position coordinates O i t and the position coordinates O of the person 401(tu) in the frame 400-(tu). i t-u is used to determine the movement vector 420 of the person 401(t), and the movement vector V i t Generate.

[0033] Similarly, the vector information generating unit 302 calculates the position coordinates O of the person 402(t) in the frame 400-t in the composite frame 410. j t and the position coordinates O of the person 402(tu) in the frame 400-(tu). j t-u is used to determine the movement vector 420 of the person 402(t), j tGenerate.

[0034] Here, the movement vector V i t Using the coordinates of the rectangle information R, the velocity n i t and direction d i t is defined by the following equations (1) and (2): i t is defined by the following formula (3).

[0035]

[0036] Returning to the explanation of FIG. 3, the velocity distribution calculation unit 303 calculates the velocity distribution of the acquired movement vector V i t From the maximum speed n in the learning data at each time i t fastest and maximum speed n i t fastest The movement vector V i t fastest and the velocity distribution 501 (described later with reference to FIG. 5) of all people to whom trajectory IDs have been assigned, and output these to the histogram feature value generation unit 305. In addition, the velocity distribution calculation unit 303 calculates the movement vector V i t fastest to the directional distribution calculation unit 304.

[0037] The direction distribution calculation unit 304 calculates the movement vector V obtained from the vector information generation unit 302. i t and the maximum speed n obtained from the speed distribution calculation unit 303 i t fastest The movement vector V i t fastest Based on this, the direction distribution 502 (described later in FIG. 5) of all people to whom the trajectory ID is assigned and the maximum speed n i t fastest The movement vector V i t fastest Direction d i tfastest and outputs it to the histogram feature value generating unit 305.

[0038] [Fig. 5 Speed ​​distribution 501 and direction distribution 502] Fig. 5 is an explanatory diagram showing a speed distribution 501 and a direction distribution 502. In the speed distribution 501, the horizontal axis represents the trajectory ID, and the vertical axis represents the speed n i t The velocity distribution 501 exists for each time t. The direction distribution 502 has a horizontal axis representing the trajectory ID and a vertical axis representing the direction d. i t The directional distribution 502 also exists for each time t.

[0039] In the velocity distribution 501, the maximum velocity n i t fastest In the direction distribution 502, the direction d of the trajectory ID=9 is i t But the maximum speed n i t fastest Direction d i t fastest is.

[0040] Returning to the explanation of Fig. 3, the histogram feature value generation unit 305 calculates the velocity n i t and direction d i t Histogram feature H t Calculate.

[0041] [Figure 6 Histogram feature H t 6 shows the histogram feature H t 6 is an explanatory diagram showing an example of calculation of the first one-dimensional histogram 601, in which the horizontal axis represents the velocity n i t The second one-dimensional histogram 602 is a one-dimensional histogram with the vertical axis representing the number of people. i t The two-dimensional histogram 603 is a one-dimensional histogram with the vertical axis representing the number of people. it , the y-axis is in the direction d i t The first one-dimensional histogram 601, the second one-dimensional histogram 602, and the two-dimensional histogram 603 are generated by the histogram feature generator 305.

[0042] speed n i t Velocity n in the first one-dimensional histogram 601 i t Number of people in each direction d i t The direction d in the second one-dimensional histogram 602 i t The combination of the number of people at each time point and the number of people at each time point is expressed as a one-dimensional histogram feature H1 t The speed n i t and direction d i t Velocity n in the two-dimensional histogram 603 i t and direction d i t The number of people for each combination is expressed as the two-dimensional histogram feature H2 at time t. t The one-dimensional histogram feature H1 t and two-dimensional histogram feature H2 t If no distinction is made between t It is called.

[0043] The histogram feature value generating unit 305 generates a histogram feature value Ht and a maximum speed n i t fastest The movement vector V i t fastest Maximum speed n i t fastest and direction d i t fastest and are output to the shift feature generating unit 306 as inclusion features.

[0044] Returning to FIG. 3 , the shift feature generation unit 306 combines the inclusion features of each acquired frame 400-t with the inclusion features of a frame a predetermined number of frames prior, and outputs the combined shift features to the learning unit 307. Applying the shift features to learning makes it possible to prevent several types of states from being mixed together when determining the continuous state of a space on a frame-by-frame basis. Specifically, for example, in a 10 fps video, the situation inside a train is unlikely to change on a frame-by-frame basis (every 0.1 seconds), and is expected to change collectively over a certain period of time. Therefore, the inclusion features of the frame 400-t alone may result in a one-off misjudgment. Therefore, by generating shift features by combining the inclusion features of frames a predetermined number of frames prior, it is possible to suppress such one-off misjudgments.

[0045] [Figure 7 Shift Features] Figure 7 is an explanatory diagram showing an example of generating shift features. A time-series training data group 700 is composed of multiple time-series frames, specifically, for example, a first frame 400-1, ..., a δ-th frame 400-δ, a δ+1-th frame 400-(δ+1), a δ+2-th frame 400-(δ+2), ..., a t-1-th frame 400-(t-1), and a t-th frame 400-t, where δ<t is a non-negative integer.

[0046] The inclusive feature set 701 is a set of inclusive features 701-1 to 701-t for each frame of the time-series training data set 700.

[0047] The shift feature group 702 is a set of shift features 702-(δ+1) to 702-t accumulated in time series, generated by the shift feature generation unit 306. The shift feature generation unit 306 generates shift features 702-(δ+1) to 702-t for frames 400-(δ+1) to 400-t using inclusive features 701-(δ+1) to 701-t for frames 400-(δ+1) to 400-t and inclusive features 701-1 to 701-(t-δ) for frames 400-1 to 400-(t-δ) that are δ frames before. The specific value of δ is not limited.

[0048] For example, the inclusive feature 701-t generated from the t-th frame 400-t has a maximum speed of n i t fastest and maximum speed n i t fastest The movement vector V i t fastest Direction d i t fastest and the histogram feature H t and

[0049] Similarly, the inclusive feature 701-(t-δ) of the t-δ-th frame 400-(t-δ) is i t-δ fastest and maximum speed n i t fastest The movement vector V i t fastest Direction d i t fastest and the histogram feature H t-δ and

[0050] Therefore, the shift feature 702-t generated from the t-th frame 400-t is the maximum speed n i t fastest and maximum speed n i t fastest The movement vector V i t fastest Direction d i t fastest and the histogram feature H t and the maximum velocity n in the t-δ-th frame i t-δ fastest and maximum speed n i t-δ fastest The movement vector V i t-δ fastest Direction d it-δ fastest and the histogram feature H t-δ However, since there is no previous frame δ frames before the frames 400-1 to 400-δ from t=1 to t=δ, the process of generating shift features is executed for the δ+1th frame 400-(δ+1) and onwards.

[0051] Returning to the explanation of Fig. 3, the learning unit 307 performs learning using the acquired shift feature amount group 702 as an explanatory variable and the state of the learning data of frames 400-(δ+1) to 400-t as a target variable, and generates an inference model by machine learning that infers the state of the target (person) in frames 400-(δ+1) to 400-t using the shift feature amount group 702. The learning unit 307 outputs the inference model generated by learning to the inference unit 317.

[0052] <Example of Functional Configuration on the Client 102 Side> Next, a description will be given of an example of the functional configuration on the client 102 side. The person detection unit 310 acquires analysis target data from the sensor 103, detects a person from the analysis target data by using the results of learning information about people in advance through machine learning, generates rectangle information R = (x1, y1, x2, y2) indicating the presence of a person, and outputs the center point of the rectangle information R to the tracking processing unit 311 as the position coordinates O = (x, y) of the person.

[0053] Although frame 400-t as learning data in teacher signal DB 104 and the frame acquired from sensor 103 as data to be analyzed differ in both time and image data, for ease of explanation, client 102 also uses t as the time symbol and 400 as the frame symbol.

[0054] The tracking processing unit 311 generates a trajectory ID unique to a person for the position coordinates O of the person acquired from the person detection unit 310. The trajectory IDs of multiple position coordinates O originating from the same person among the acquired multiple learning data will have the same value. The tracking processing unit 301 outputs the trajectory ID and the position coordinates O of the person associated with the trajectory ID to the vector information generation unit 312.

[0055] The vector information generating unit 312 generates a movement vector 420 based on the acquired trajectory ID and position coordinates O, and outputs the movement vector 420 to the speed distribution calculating unit 313 and the direction distribution calculating unit 314 .

[0056] The velocity distribution calculation unit 313 calculates the velocity distribution of the acquired movement vector V i t From the maximum speed n in frame 400-t at each time t i t fastest and maximum speed n i t fastest The movement vector V i t fastest The velocity distribution calculation unit 313 calculates the velocity distribution 501 of all people to whom the trajectory IDs are assigned, and outputs the velocity distribution 501 to the histogram feature value generation unit 315. In addition, the velocity distribution calculation unit 313 calculates the velocity distribution 501 of all people to whom the trajectory IDs are assigned, and outputs the velocity distribution 501 to the histogram feature value generation unit 315. i t fastest to the directional distribution calculation unit 314.

[0057] The direction distribution calculation unit 314 calculates the movement vector V obtained from the vector information generation unit 312. i t and the maximum speed n obtained from the speed distribution calculation unit 313 i t fastest The movement vector V i t fastest Based on this, the direction distribution 502 of all people to whom the trajectory ID is assigned and the maximum speed n i t fastest The movement vector V i t fastest Direction d i t fastest and outputs it to the histogram feature value generating unit 315.

[0058] The histogram feature generation unit 315 calculates the velocity n i t and direction d i t Histogram feature H tThe histogram feature value generation unit 315 calculates the histogram feature value H t and maximum speed n i t fastest The movement vector V i t fastest Maximum speed n i t fastest and direction d i t fastest and are output to the shift feature generating unit 316 as inclusion features.

[0059] For each acquired frame 400-t, the shift feature generation unit 316 combines the inclusion feature 701-t of the frame 400-t with the inclusion feature 701-(t-δ) of a past frame 400-(t-δ) that is δ frames before the time t, and outputs the combined result as a shift feature 702-t to the inference unit 317.

[0060] For example, the shift feature amount of the t-th frame 400-t acquired as the analysis target data from the sensor 103 is the maximum speed n i t fastest and maximum speed n i t fastest The movement vector V i t fastest Direction d i t fastest and the histogram feature H t and the maximum speed n at 400-(t-δ) i t-δ fastest and maximum speed n i t-δ fastest The movement vector V i t-δ fastest Direction d i t-δ fastest and the histogram feature H t-δHowever, for frames 400-1 to 400-δ from t=1 to t=δ, there is no past frame δ frames before, so the generation process of the shift feature 702-t is executed for the δ+1th frame 400-(δ+1) and onwards.

[0061] The inference unit 317 inputs the acquired shift feature set 702, which is the shift feature set 702-(δ+1) at time (δ+1) to the shift feature set 702-t at time t, to the inference model acquired from the learning unit 307, and infers the state of the target (person) in frames 400-(δ+1) to 400-t using the shift feature set 702. The inference unit 317 outputs the inference result to the inference result output unit 318.

[0062] If the inference result acquired from the inference unit 317 is a predetermined inference result, the inference result output unit 318 transmits the inference result to the inference result receiving unit 320 in the terminal 106. Here, the predetermined inference result indicates, for example, a state in which additional control is required from a normal or abnormal state as shown in Fig. 9, and the necessary control can be performed by transmitting the predetermined inference result.

[0063] <Example of Functional Configuration on the Terminal 106 Side> Next, a description will be given of an example of a functional configuration on the terminal 106 side. The inference result receiving unit 320 indicates, in accordance with the inference result received from the inference result output unit 318, that additional control is required.

[0064] 8 is an explanatory diagram showing an example of a method for storing learning data. The learning data is stored in a teacher signal storage folder 800 in the teacher signal DB 104. The teacher signal storage folder 800 has a first subfolder 810, a second subfolder 820, and a third subfolder 830.

[0065] The first subfolder 810 stores subfolders 811 and 812. The subfolder 811 stores image data of 000.jpg811a and image data of 001.jpg811b. The subfolder 812 stores image data of 000.jpg812b and image data of 001.jpg812b.

[0066] The second subfolder 820 stores a subfolder 821 and image data of 012.avi 822. The subfolder 821 stores image data of 010.jpg 821a and image data of 011.jpg 821b.

[0067] In the third subfolder 830, image data of 020.avi831, image data of 021.avi832, and image data of 012.avi833 are stored.

[0068] The time-series image data is stored in a first subfolder 810, a second subfolder 820, and a third subfolder 830 for each state. The state is the name of the subfolder. Specifically, for example, the state of the first subfolder 810 is "class A," the state of the second subfolder 820 is "class B," and the state of the third subfolder 830 is "class C." Furthermore, under the state subfolders, sets of image data belonging to different time series are stored in subfolders 811, 812, and 821. The set is the name of the subfolder. Specifically, for example, the set in subfolder 811 is "set01," the set in subfolder 812 is "set02," and the set in subfolder 821 is "set11."

[0069] 8, the number of subfolders divided by state is three, but it may be two or more. Furthermore, the number of image data and the number of subfolders in the first subfolder 810, the second subfolder 820, and the third subfolder 830 are not limited to those shown in FIG.

[0070] 3 acquires time-series image data organized by subfolder, along with the subfolder name indicating the state, and uses it for learning. For example, when training an inference model on a class-by-class basis, by specifying the target subfolder, the server 101 acquires the image data in the subfolder, inputs it into the inference model, compares the output result with the subfolder name indicating the state, calculates the value of the loss function based on the comparison result (match or mismatch), and performs error backpropagation, thereby enabling training of the inference model.

[0071] Furthermore, when adding a new teacher signal to the teacher signal storage folder 800, the server 101 searches the teacher signal storage folder 800 based on the state of the teacher signal to be added, identifies a subfolder to be used as the storage destination, and stores the image data to be added. In this way, the generation and management of teacher signals is simplified.

[0072] <FIG. 9 Example of Space Analysis by Analysis System> FIG. 9 is an explanatory diagram showing an example of space analysis by the analysis system. Images 901 and 902 are images displayed based on image data captured by sensor 103. Image 901 shows the interior of a train car as a space 900 containing people. In space 900 shown in image 901, there may be specific trends in people's movements, such as "people standing still," "walking slowly in a certain direction," "groups of people uniformly distributed throughout the space," or "people congregating in specific locations." Analysis system 100 can be configured to determine a state in which people move in such a specific manner as a normal state. Note that such trends in people's movements may vary depending on the location, the time of day, and other conditions.

[0073] Such human movement patterns can change dynamically due to unexpected events (unexpected incidents). Examples of such events include incidents or accidents that threaten personal safety or make people want to leave the scene of the incident. Examples of such incidents include violent people, the discovery of suspicious objects, and vomit. Examples of such accidents include vehicle breakdowns.

[0074] Image 902 in FIG. 9 shows how the movement tendency of people changes due to the presence of an unexpected event in space 900 shown in image 901. When an unexpected event occurs, as shown in image 902, people move in direction d2, which is different from direction d1 shown in image 901, to move away from source 921 of the unexpected event. Direction d2 is the longitudinal direction of the train car. In other words, people move along direction d2 to move to the next car. This tendency will continue as long as an unexpected event occurs in space 900. Analysis system 100 determines that such a state in which people move in a manner that differs from the normal state as an abnormal state.

[0075] Analysis system 100 can determine whether space 900 is in a normal state or an abnormal state based on the movement trends of people in space 900. Furthermore, if analysis system 100 determines that an abnormal state exists, it can identify the source of an unexpected event caused by the abnormal state using information such as the direction of people's movement and the density of people gathering. Analysis system 100, for example, enables a system administrator to become aware of the occurrence of an unexpected event in space 900, protect the safety of people in space 900, and quickly resolve the unexpected event.

[0076] FIG. 10 is a flowchart illustrating an example of a detailed procedure of the learning process performed by the server 101 (learning device) according to the first embodiment.

[0077] (Step S1000) The server 101 causes the human detection unit 300 to acquire a teacher signal to be used for learning from the teacher signal DB 104.

[0078] (Step S1001) Next, the server 101 causes the person detection unit 300 to detect a person from the acquired teacher signal used for learning, and calculates the position coordinates O of the person in each frame.

[0079] (Step S1002) Next, the server 101 assigns a trajectory ID to the acquired position coordinates O of the person by the tracking processing unit 301. The trajectory IDs of multiple position coordinates O originating from the same person among the acquired multiple frames are set to the same value.

[0080] (Step S1003) Next, the server 101 calculates a movement vector 420 in each frame of the person to whom the trajectory ID is assigned by the vector information generation unit 302. The movement vector 420 includes a velocity n i t and direction d i t Includes:

[0081] (Step S1004) Next, the server 101 calculates the speed n i t The highest movement vector V i t fastest and the highest velocity n i t fastest The direction distribution calculation unit 304 also obtains the velocity n i t The highest movement vector V i t fastest Direction d i t fastest Get.

[0082] (Step S1005) Next, the server 101 causes the speed distribution calculation unit 303 to calculate the speed distribution 501 of the person to whom the trajectory ID has been assigned.

[0083] (Step S1006) Next, the server 101 causes the direction distribution calculation unit 304 to calculate the direction distribution 502 of the person to whom the trajectory ID has been assigned.

[0084] (Step S1007) Next, the server 101 uses the histogram feature value generating unit 305 to calculate the velocity n for each frame based on the acquired velocity distribution 501 and direction distribution 502. i t and direction d i t Histogram feature H t Generate.

[0085] (Step S1008) Next, the server 101 generates a histogram feature H t and maximum speed n i t fastest and maximum speed n i t fastest The movement vector V i t fastest Direction d i t fastest and generate the inclusive feature 701-t for each frame.

[0086] (Step S1009) Next, the server 101 accumulates the inclusion features 701-t of the acquired frames in chronological order using the shift feature generation unit 306.

[0087] (Step S1010) Next, the server 101 determines whether the accumulated inclusion feature 701-t exceeds the preset number of frames δ using the shift feature generation unit 306. If the accumulated inclusion feature 701-t does not exceed the preset number of frames δ (step S1010: No), the process returns to step S1000. On the other hand, if the accumulated inclusion feature 701-t exceeds the preset number of frames δ (step S1010: Yes), the process proceeds to step S1011.

[0088] (Step S1011) Next, the server 101 generates a shift feature 702-t by using the shift feature generation unit 306 to combine, for each of the accumulated inclusion features 701-t, the inclusion feature 701-t with the inclusion feature 701-(t-δ) of a previous frame that is a preset number of frames δ before.

[0089] (Step S1012) Next, the server 101 performs learning using the learning unit 307 with the shift feature 702-t corresponding to frame 400-t as an explanatory variable and the state corresponding to frame 400-t as a target variable, and generates, by machine learning, an inference model that infers the state of the target (person) in frames 400-(δ+1) to 400-t using the shift feature group 702.

[0090] <FIG. 11 Inference Processing> FIG. 11 is a flowchart illustrating an example of a state inference processing procedure by the client 102 (inference device) according to the first embodiment.

[0091] (Step S1100 ) The client 102 acquires the current frame from the sensor 103 using the human detection unit 310 .

[0092] (Step S1101) Next, the client 102 causes the person detection unit 310 to detect people from the acquired current frame and calculate the position coordinates O of each person.

[0093] (Step S1102) Next, the client 102 assigns a trajectory ID to the acquired position coordinates O of the person by the tracking processing unit 311. The trajectory IDs of multiple position coordinates O originating from the same person among the acquired multiple frames are set to the same value.

[0094] (Step S1103) Next, the client 102 calculates the movement vector 420 of the person to whom the trajectory ID is assigned by the vector information generating unit 312. The movement vector 420 includes a speed n i t and direction d i t Includes:

[0095] (Step S1104) Next, the client 102 calculates the speed n i t The highest movement vector V i t fastest Identify the highest speed V i tfastest The direction distribution calculation unit 304 also obtains the velocity n i t The highest movement vector V i t fastest Direction d i t fastest Get.

[0096] (Step S1105) Next, the speed distribution calculation unit 313 of the client 102 calculates the speed distribution 501 of the person to whom the trajectory ID has been assigned.

[0097] (Step S1106) Next, the client 102 calculates the direction distribution 502 of the person to whom the trajectory ID has been assigned, using the direction distribution calculation unit 314.

[0098] (Step S1107) Next, the client 102 calculates the velocity n from the acquired velocity distribution 501 and direction distribution 502 using the histogram feature value generation unit 315. i t and direction d i t Histogram feature H t Generate.

[0099] (Step S1108) Next, the client 102 generates the histogram feature H t and maximum speed n i t fastest and maximum speed n i t fastest The movement vector V i t fastest Direction d i t fastest and generate the inclusive feature 701-t of the current frame.

[0100] (Step S1109) Next, the client 102 accumulates the inclusion features 701-t of the acquired frames 400-t in chronological order using the shift feature generator 316.

[0101] (Step S1110) Next, the client 102 determines whether the accumulated inclusion feature 701-t exceeds the preset number of frames δ using the shift feature generation unit 316. If the accumulated inclusion feature 701-t does not exceed the preset number of frames (step S1110: No), the process returns to step S1100. On the other hand, if the accumulated inclusion feature 701-t exceeds the preset number of frames (step S1110: Yes), the process proceeds to step S1111.

[0102] (Step S1111) Next, the client 102 generates a shift feature 702-t by using the shift feature generation unit 316 to combine, for each of the accumulated inclusion features 701-t, the inclusion feature 701-t of the current frame 400-t with the inclusion feature 701-(t-δ) of the previous frame 400-(t-δ) that is a preset number of frames δ ago.

[0103] (Step S1112) Next, the client 102, through the inference unit 317, provides the shift feature 702-t generated in step S1111 as input to the inference model acquired from the learning unit 307, and infers the state of the target (person) in frames 400-(δ+1) to 400-t based on the shift feature group 702.

[0104] (Step S1113) Next, the client 102, through the inference result output unit 318, transmits the inference result of step S1112 to the terminal 106 if the inference result is a predetermined inference result. A predetermined inference result is, for example, a specific classification class (e.g., class C in FIG. 5) that requires additional control. Note that the client 102 may transmit the inference result to the terminal 106 even if the inference result is not a specific classification class.

[0105] (Step S1114) Next, the client 102 determines whether to end the acquisition of the analysis target data from the sensor 103. If the acquisition is not to be ended (step S1114: No), the process returns to step S1100. On the other hand, if the acquisition is to be ended (step S1114: Yes), the state inference process of the client 102 ends.

[0106] Whether or not to end acquisition is set in advance by an administrator. For example, when the administrator sets the frame acquisition start time and acquisition end time, the client 102 proceeds to step S1100 until the acquisition end time (step S1114: No), and ends the state inference process when the acquisition end time is reached. Then, when the acquisition start time for the next business day arrives, frame acquisition begins (step S1100). For example, when determining whether the inside of a train is normal or abnormal using security cameras inside the train, the frame acquisition start time is set to start from the train's start time and the frame acquisition end time is set to end at the train's end time so that the state determination function works during train operating hours.

[0107] <FIG. 12 Status Notification Processing> FIG. 12 is a flowchart illustrating an example of a status notification processing procedure by the terminal 106 (status notification device) according to the first embodiment.

[0108] (Step S1200) The terminal 106 receives the inference result of step S1112 from the client 102 via the inference result receiving unit 320.

[0109] (Step S1201) The terminal 106 notifies the user of the communication terminal of the status of the image data of the current frame by executing a specific screen display process or audio output process in accordance with the received inference result.

[0110] Thus, according to Example 1, even in an environment where people are moving around densely within the space 900 or in an environment with many obstructions, the state of the space 900 to be inspected can be accurately inferred by detecting the movement trends of people.

[0111] 11, the client 102 transmits the inference result of step S1112 to the terminal 106 (step S1113), but may also transmit the result to a patrol lamp communicatively connected to the network 105. In this case, the patrol lamp lights up in accordance with the received inference result, indicating that additional control is required.

[0112] The second embodiment will be described, focusing on the differences from the first embodiment. Note that the same reference numerals are used to designate the same parts as the first embodiment, and the description thereof will be omitted.

[0113] 13 is a block diagram illustrating an example of a functional configuration of the analysis system 100 according to Example 2. In Example 1, the learning process is executed by the server 101 (learning device), but in Example 2, the learning process is executed by the client 102.

[0114] Therefore, in Example 2, the person detection unit 310 also functions as the person detection unit 300, the tracking processing unit 311 also functions as the tracking processing unit 301, the vector information generation unit 312 also functions as the vector information generation unit 302, the speed distribution calculation unit 313 also functions as the speed distribution calculation unit 303, the direction distribution calculation unit 314 also functions as the direction distribution calculation unit 304, the histogram feature generation unit 315 also functions as the histogram feature generation unit 305, and the shift feature generation unit 316 also functions as the shift feature generation unit 306.

[0115] When outputting the position coordinates O of the detected person to the tracking processing unit 311, the person detection unit 310 adds current processing information indicating whether the processing currently being executed is a learning processing or an inference processing.

[0116] The tracking processing unit 311, the vector information generating unit 312, the velocity distribution calculating unit 313, the direction distribution calculating unit 314, and the histogram feature generating unit 315 execute the same processes as those in the first embodiment.

[0117] The shift feature generation unit 316 generates a shift feature 702-t by combining the inclusion feature 701-t of each acquired frame with the inclusion feature 701-(t-δ) of a previous frame a predetermined number of frames δ prior to the current frame. If the current processing information assigned by the person detection unit 310 is a learning process, the shift feature generation unit 316 outputs the calculated shift feature 702-t to the learning unit 307, and if the current processing information is an inference process, the shift feature generation unit 316 outputs the calculated shift feature 702-t to the inference unit 317.

[0118] Thus, according to Example 2, by executing the learning process on the client 102, it becomes possible to consolidate similar processes that were previously executed on the server 101 and the client 102, thereby simplifying the analysis system 100 and enabling stand-alone state inference to be performed.

[0119] The third embodiment will be described focusing on the differences between the first and second embodiments. t The accuracy of inference based on the histogram feature H may vary depending on the number of movement vectors 420 of people detected from the analysis target data. t The state of the data to be analyzed can be inferred with higher accuracy by switching the generation method of the vectors 420 according to the number of detected movement vectors 420. Note that the same reference numerals are used to designate the same points as in the first and second embodiments, and the description thereof will be omitted.

[0120] 14 is a block diagram illustrating an example of a functional configuration of the analysis system 100 according to Example 3. In Example 3, a people measurement unit 309 is newly added to the server 101, and a people measurement unit 319 is newly added to the client 102.

[0121] The vector information generation unit 302 outputs the acquired trajectory ID and the movement vector 420 calculated from the position coordinate O to the number of people measurement unit 309. The vector information generation unit 312 outputs the acquired trajectory ID and the movement vector 420 calculated from the position coordinate O to the number of people measurement unit 319.

[0122] The number of people measurement unit 309 measures the number of movement vectors 420 output from the vector information generation unit 302 for each piece of training data, and outputs the measured number of movement vectors 420 as the number of people to the histogram feature generation unit 305 and the training unit 307.

[0123] The histogram feature value generation unit 305 uses the velocity distribution 501 acquired from the velocity distribution calculation unit 303 and the direction distribution 502 acquired from the direction distribution calculation unit 304 to generate a histogram feature value H t Specifically, when the number of people acquired from the number of people measurement unit 309 is greater than a predetermined number of people, the histogram feature value generation unit 305 generates the two-dimensional histogram feature value H2 t If the number of people is less than a predetermined number, the one-dimensional histogram feature quantity H1 t Generate.

[0124] If the number of people acquired from the number of people measurement unit 309 is greater than a predetermined number of people, the learning unit 307 converts the shift feature amount 702-t acquired from the shift feature amount generation unit 306 into the two-dimensional histogram feature amount H2 t When the number of people acquired from the number of people measurement unit 309 is equal to or smaller than a predetermined number of people, the learning unit 307 performs learning by converting the shift feature amount 702-t acquired from the shift feature amount generation unit 306 into the one-dimensional histogram feature amount H1 t Learning is performed by using the corresponding explanatory variables of the inference model.

[0125] The number of people measurement unit 319 measures the number of movement vectors 420 output from the vector information generation unit 312 for each analysis target data 500, and outputs the measured number of movement vectors 420 as the number of people to the histogram feature generation unit 315 and the inference unit 317.

[0126] The histogram feature value generation unit 315 uses the velocity distribution 501 acquired from the velocity distribution calculation unit 313 and the direction distribution 502 acquired from the direction distribution calculation unit 314 to generate a histogram feature value H tSpecifically, when the number of people acquired from the number of people measurement unit 319 is greater than a predetermined number of people, the histogram feature value generation unit 315 generates the two-dimensional histogram feature value H2 t If the number of people is less than a predetermined number, the one-dimensional histogram feature quantity H1 t Generate.

[0127] If the number of people acquired from the number of people measurement unit 319 is greater than a predetermined number of people, the inference unit 317 uses the two-dimensional histogram feature amount H2 t When the number of people acquired from the number of people measurement unit 319 is equal to or less than a predetermined number of people, the inference unit 317 infers the state of the analysis target data using an inference model corresponding to the one-dimensional histogram feature quantity H1 t The state of the data to be analyzed is inferred using an inference model corresponding to the above.

[0128] Because the bins in the two-dimensional histogram 603 are the intersection of the bins in the first one-dimensional histogram 601 and the bins in the second one-dimensional histogram 602, the bins in the two-dimensional histogram 603 are divided more finely than the bins in the first one-dimensional histogram 601 and the bins in the second one-dimensional histogram 602. Therefore, the two-dimensional histogram 603 can more accurately reproduce the situation on-site. However, if the number of people acquired is less than a predetermined number, dividing the values ​​into bins according to the feature values ​​results in low heights, such as approximately 0 for some bins, and the histogram feature values ​​become sparse data with scattered data. Sparse data contains many zero values, and the data has little variation, making it difficult to properly train the input data. This is because, for example, when the data is sparse and there is little variation in the data, the model is prone to overtraining on training data with little variation and roughly zero values.

[0129] For example, the bins for speed n in the first one-dimensional histogram 601 indicating speed are 0<n≦20, 20<n≦40, 40<n≦60, ..., n>300. In addition, in the two-dimensional histogram 603, for example, the bin 0<n≦20 is classified as "0<n≦20 and 0<angle≦30", "0<n≦20 and 30<angle≦60", ..., "0<n≦20 and 330<angle≦360" (the same applies to 20<n≦40...).

[0130] In this case, the total number of people in the two-dimensional histogram 603, the total number of people in the first one-dimensional histogram 601, and the total number of people in the second one-dimensional histogram 602 are the same, and because the bins in the two-dimensional histogram 603 are divided into smaller parts, the height of the bins in the two-dimensional histogram 603 is lower than those in the first one-dimensional histogram 601 and the second one-dimensional histogram 602.

[0131] In the above classification example, in the first one-dimensional histogram 601 indicating speed, the movement direction of people falling into each bin (0<n≦20, 20<n≦40, 40<n≦60, ..., n>300) is unknown. In the two-dimensional histogram 603, the movement direction of people falling into each bin ("0<n≦20 and 0<angle≦30", "0<n≦20 and 30<angle≦60", ..., "0<n≦20 and 330<angle≦360") is clear, making it possible to more accurately reproduce the situation on-site.

[0132] Therefore, when the number of people to be acquired is equal to or less than a predetermined number of people, the one-dimensional histogram feature amount H1 based on the first one-dimensional histogram 601 and the second one-dimensional histogram 602 is t By applying this to make an inference, the decrease in inference accuracy is suppressed, and if the number of acquired people is greater than the predetermined number of people, the two-dimensional histogram feature H2 t By applying this to inference, the accuracy of inference can be improved.

[0133] 15 is a flowchart illustrating a detailed example of a processing procedure of a learning process by the server 101 (learning device) according to Example 3. After step S1003, the server 101 causes the number of people measuring unit 309 to calculate the number of movement vectors 420 acquired from the vector information generating unit 302 as the number of people (step S1503). Then, the process proceeds to step S1004.

[0134] After step S1006, the server 101 uses the histogram feature value generation unit 305 to generate a two-dimensional histogram feature value H2 t If the number of people is less than a predetermined number, the one-dimensional histogram feature quantity H1 t (step S1507), and the process proceeds to step S1008.

[0135] After step S1011, the server 101 sets the state constituting the teacher signal as the objective variable by the learning unit 307. If the number of people acquired from the number of people measurement unit 309 is greater than a predetermined number, the server 101 uses the learning unit 307 to convert the shift feature 702-t into the two-dimensional histogram feature H2 t When the number of people is equal to or less than a predetermined number, the shift feature 702-t is set as an explanatory variable of the inference model corresponding to the one-dimensional histogram feature H1 t The server 101 uses the learning unit 307 to perform learning using the set objective variable and explanatory variables, and generates an inference model by machine learning that infers the state of the target (person) in frames 400-(δ+1) to 400-t using the shift feature set 702 (step S1012).

[0136] 16 is a flowchart illustrating an example of a state inference processing procedure by the client 102 (inference device) according to Example 3. After step S1103, the client 102 calculates, using the number of people measuring unit 319, the number of movement vectors 420 acquired from the vector information generating unit 312 as the number of people (step S1603).

[0137] After step S1111, if the number of people acquired from the number of people measurement unit 309 is greater than the predetermined number of people, the client 102 uses the inference unit 317 to calculate the two-dimensional histogram feature quantity H2 t The shift feature 702-t including the following is input to the inference model, and when the number of people is equal to or less than a predetermined number, the one-dimensional histogram feature H1 t The shift feature 702-t including the above is used as an input to the inference model, and the state of the object (person) in frames 400-(δ+1) to 400-t is inferred using the shift feature group 702 (step S1612).

[0138] In this way, according to the third embodiment, the histogram feature quantity H t By switching between these to generate an inference model and switching the inference model according to the number of detected human movement vectors 420, it is possible to achieve highly accurate state inference according to the number of people detected.

[0139] The fourth embodiment will be described, focusing on the differences between the first, second, and third embodiments. The same reference numerals are used to designate the same parts as the first to third embodiments, and the description thereof will be omitted.

[0140] 17 is a block diagram illustrating an example of a functional configuration of the analysis system 100 according to Example 4. In Example 4, the server 101 (learning device) does not use the person detection unit 300, the tracking processing unit 301, or the vector information generation unit 302. The time-series learning data output from the teacher signal DB 104 to the speed distribution calculation unit 303 and the direction distribution calculation unit 304 includes a movement vector 420 of a person at each time.

[0141] 18 is an explanatory diagram showing an example of time-series learning data output from the teacher signal DB 104. A time-series image data group 1800 is made up of t frames 1800-1 to 1800-t.

[0142] The time-series learning data group 1801 is a set of learning data 1801-1 to 1801-t output from the teacher signal DB 104. The learning data 1801-t is a set of velocity vectors (n 1 t , n 2 t , ..., n i t ) and a set of direction vectors at time t (d 1 t , d 2 t , ..., d i t ) and the set of motion vectors (V 1 t , V 2 t , ..., V i t In the time-series learning data group 1801, the learning data 1801-t at each time t is information derived from the frame 1800-t at the same time t in the time-series image data group 1800. For example, the set of motion vectors (V 1 t , V 2 t , ..., V i t ) is the movement vector V of a person (trajectory ID=1 to i) in the t-th frame 1800-t of the time-series image data group 1800. 1 t , V 2 t , ..., V i t is.

[0143] Thus, according to the fourth embodiment, by using learning data 1801-t, which is only numerical information rather than image data (frame 1800-t), it is possible to reduce the size of the learning data and shorten the learning time.

[0144] The fifth embodiment will be described, focusing on the differences from the first to fourth embodiments. The same reference numerals are used to designate the same parts as the first to fourth embodiments, and the description thereof will be omitted.

[0145] 19 is a block diagram showing an example of the functional configuration of the analysis system 100 according to Example 5. In Example 5, an event occurrence location output unit 321 is newly added to the terminal 106. The inference result receiving unit 320 indicates a state in which additional control is required according to the inference results received from one or more clients 102, and outputs the inference results and the location information of the clients 102 from which the inference results originate to the event occurrence location output unit 321.

[0146] Next, the event occurrence location output unit 321 displays the state requiring additional control and the location from which the inference result originates, based on the inference results and location information of the multiple clients 102 acquired from the inference result receiving unit 320.

[0147] Thus, according to Example 5, when processing inference results and location information from multiple clients 102 in a unified manner, it is possible to grasp the location where a state requiring additional control occurs and quickly implement the additional control.

[0148] The sixth embodiment will be described, focusing on the differences from the first to fifth embodiments. The same reference numerals are used to designate the same parts as the first to fifth embodiments, and the description thereof will be omitted.

[0149] In the sixth embodiment, the inclusion features generated by the histogram feature generator 305 and the histogram feature generator 315 include the maximum speed n i t fastest and maximum speed n i t fastest The movement vector V i t fastest Direction d i t fastest and the histogram feature H t In addition to the above, other information may also be included as a feature. Here, the other information may be, for example, the position coordinates of the detected person, the movement vector V i t Speed ​​n i t The acceleration, which is the derivative of it Direction d i t The variance or standard deviation of

[0150] The learning unit 307 learns an inference model using the shift feature 702-t, which includes other information as an explanatory variable, and the corresponding known state, the point of occurrence and the time of occurrence of the state as objective variables. The inference unit 317 infers the state, the point of occurrence and the time of occurrence of the state by inputting the shift feature 702-t, which includes other information, into the inference model.

[0151] In this way, according to the sixth embodiment, the movement vector V i t By inputting information such as the acceleration of the vehicle and the coordinates of the detected person's location into the inference model, it is possible to infer the state and accurately infer the location and time at which the state occurred.

[0152] The seventh embodiment will be described, focusing on the differences from the first to sixth embodiments. The same reference numerals are used to designate the same parts as the first to sixth embodiments, and the description thereof will be omitted.

[0153] 20 is an explanatory diagram illustrating an example of a shift feature group according to Example 7. In Example 7, the shift feature generation unit 306 and the shift feature generation unit 316 calculate the inclusive feature 701-t of the current frame 400-t and the inclusive feature 701-t of a plurality of past frames (δ 1 previous frame, δ 2 previous frame, ..., δ m Inclusive feature 701-(t-δ) of the previous frame 1 ) ~ 701-(t-δ m ) are combined and output as a shift feature 2003-t.

[0154] δ 1 ~δ m is δ 1 <δ 2 <, ..., <δ m and m is an integer of 1 or greater that satisfies the following formula:

[0155] In addition, from the first to the δ mFrames 400-1 to 400-δ m For δ m Since there is no previous frame, the generation process of the shift feature 2003-t is performed by δ m+1 Frame 400-δ after the m+1 is executed against.

[0156] As described above, according to the seventh embodiment, state inference can be realized with high accuracy using feature amounts of a plurality of past frames.

[0157] In the above-described first to seventh embodiments, the normal state and abnormal state shown in FIG. 9 are assumed, but the combination of normal and abnormal states is not limited to this assumption. For example, it is also possible to assume a case in which people move in a specific direction in the normal state, and in the abnormal state, the movement speed of people is slower than the movement speed in the normal state and the movement direction becomes random. Specifically, for example, in a normal state in which a crowd is moving in a fixed direction toward the venue of an outdoor event, a case is assumed in which the speed of people flow slows due to traffic congestion or a riot, causing confusion in the crowd.

[0158] In this case, the inclusive feature 701-t generated from the t-th frame 400-t has a minimum speed n i t slowest and the minimum speed n i t slowest The movement vector V i t slowest Direction d i t slowest and the histogram feature H t In an abnormal state, the maximum speed n i t fastest and minimum speed n i t slowest Not limited to, a specific speed n different from the normal state i t and the direction d at that specific speed i t The inclusion feature 701-t may be generated assuming the above.

[0159] The eighth embodiment will be described, focusing on the differences from the first to seventh embodiments. The same reference numerals are used to designate the same components as the first to seventh embodiments, and their description will be omitted. The eighth embodiment is an example of learning and inferring whether a person is looking in a specific direction (for example, a signage above a vehicle door). This inference model has three-dimensional position information of the signage that is the subject of the gaze.

[0160] 21 is a block diagram illustrating an example of a functional configuration of the analysis system 100 according to Example 8. In Example 8, the person detection unit 310, the vector information generation unit 302, the speed distribution calculation unit 313, the direction distribution calculation unit 314, and the histogram feature generation unit 315 are not used in the client 102. On the other hand, the face detection unit 2101, the eye detection unit 2102, the face direction calculation unit 2103, and the gaze vector generation unit 2104 are newly added to the client 102.

[0161] The face detection unit 2101 generates facial landmarks for the acquired frame and outputs the frame and the facial landmarks to the tracking processing unit 311. Landmarks are feature points for extracting facial features such as the positions of the eyes and nose. When outputting the frame and the facial landmarks to the tracking processing unit 311, the face detection unit 2101 also adds current processing information indicating whether the currently executing processing is a learning process or an inference process.

[0162] The tracking processing unit 311 generates a trajectory ID for the facial landmark acquired from the face detection unit 2101. The trajectory ID corresponds to the three-dimensional position of the face. If the facial landmarks between the acquired frames originate from the same person, the tracking processing unit 311 sets the trajectory ID to the same value. The tracking processing unit 311 outputs the trajectory ID and the facial landmarks associated with the trajectory ID to the eye detection unit 2102.

[0163] The eye detection unit 2102 identifies the position of the person's eyes from the frame using the facial landmarks acquired from the tracking processing unit 311, and outputs the trajectory ID, the frame, and image data of the person's eyes to the face direction calculation unit 2103. The image data of the eyes is used to determine whether the eyes are open or closed. That is, if the eyes are open, it is determined that the person is looking at something, and if the eyes are closed, it is determined that the person is not looking at anything. The image data of the eyes is also used to identify the position of the iris when the eyes are open. The position of the iris is used to generate a gaze vector.

[0164] The face direction calculation unit 2103 calculates the direction in which the face is facing from the frames acquired from the face detection unit 2101, and outputs the direction in which the face is facing, image data of the person's eyes, and a trajectory ID to the gaze vector generation unit 2104. The face direction is movement information that indicates the direction in which the face is moving, and is used to generate the gaze vector together with the position of the pupil.

[0165] The gaze vector generation unit 2104 generates a gaze vector representing the gaze direction using the face direction and image data of the human eye acquired from the face direction calculation unit 2103, and outputs the gaze vector, face direction, image data of the human eye, and landmarks to the shift feature generation unit 316 as inclusive features of each frame.

[0166] The shift feature generation unit 316 generates a shift feature by combining the inclusive feature of the current frame and the inclusive feature of the past frame with respect to the gaze vector, the direction of the face, the image data of the human eyes, and the landmarks acquired from the gaze vector generation unit 2104. If the current processing information is a learning process, the shift feature generation unit 316 outputs the calculated shift feature to the learning unit 307, and if the current processing information is an inference process, the shift feature generation unit 316 outputs the calculated shift feature to the inference unit 317.

[0167] The learning unit 307 performs learning using the acquired shift feature set 702, which is the shift feature set from time (δ+1) to time t, as explanatory variables and the state of the learning data, which is frames 400-(δ+1) to 400-t, as the objective variable, and generates an inference model by machine learning that infers the state of the target (person) in frames 400-(δ+1) to 400-t (whether or not they are looking in a specific direction (for example, the direction of a signage)) based on the shift feature set 702. The learning unit 307 outputs the inference model generated by learning to the inference unit 317.

[0168] The inference unit 317 inputs the acquired shift feature set 702, which is the shift feature set at time (δ+1) to the shift feature set at time t, into the inference model acquired from the learning unit 307, and infers the state of the target (person) in frames 400-(δ+1) to 400-t (whether or not the person is looking in a specific direction (for example, the direction of a signage)) based on the shift feature set 702. The inference unit 317 outputs the inference result to the inference result output unit 318.

[0169] FIG. 22 is a flowchart illustrating a detailed example of a learning process performed by the client 102 (learning device) according to the eighth embodiment.

[0170] (Step S2201) After step S1000, the client 102 causes the face detection unit 2101 to detect a human face.

[0171] (Step S2202) After step S2201, the client 102 causes the face detection unit 2101 to detect landmarks for the detected human face.

[0172] (Step S2203) After step S2202, the client 102 extracts image data of the human eyes using the acquired landmarks from the eye detection unit 2102.

[0173] (Step S2204) After step S2203, the client 102 causes the face direction calculation unit 2103 to calculate the direction in which the face is facing from the acquired image data of the eyes.

[0174] (Step S2205) After step S2204, the client 102 generates a gaze vector representing the gaze direction using the acquired face direction and image data of the human eyes, via the gaze vector calculation unit 333.

[0175] (Step S2209) After step S2205, the client 102 causes the shift feature generation unit 316 to accumulate the gaze vector, the face direction, the image data of the human eyes, and the facial landmarks in the frame in chronological order as inclusion features of the frame.

[0176] <FIG. 23 Inference Processing> FIG. 23 is a flowchart illustrating an example of a state inference processing procedure by the client 102 (inference device) according to the eighth embodiment.

[0177] (Step S2301) After step S1100, the client 102 causes the face detection unit 2101 to detect a human face.

[0178] (Step S2302) After step S2301, the client 102 causes the face detection unit 2101 to detect landmarks for the detected human face.

[0179] (Step S2303) After step S2302, the client 102 extracts image data of the human eye using the acquired landmarks from the eye detection unit 2102.

[0180] (Step S2304) After step S2303, the client 102 causes the face direction calculation unit 2103 to calculate the direction in which the face is facing from the acquired image data.

[0181] (Step S2305) After step S2304, the client 102 uses the acquired face direction and image data of the human eyes to generate a gaze vector that indicates the gaze direction, using the gaze vector calculation unit 333.

[0182] (Step S2309) After step S2305, the client 102 causes the shift feature generation unit 316 to accumulate the gaze vector, the face direction, the image data of the human eyes, and the facial landmarks in each frame in chronological order as inclusive features for each frame.

[0183] In this way, according to the eighth embodiment, it is possible to infer whether or not a person who is a subject is looking at a signage by using information such as a line-of-sight vector.

[0184] In this way, according to the above-described first to eighth embodiments, it is possible to improve the learning accuracy and inference accuracy of the inference model that infers states in space.

[0185] The analytical systems of the first to eighth embodiments described above can also be configured as follows (1) to (11).

[0186] (1) An analysis system 100 having a processor 201 that executes a program and a storage device 202 that stores the program, wherein the processor 201 performs an acquisition process of repeatedly acquiring inclusion features (701-1 to 701-t) that include information on the movement of an inference target (person) and information on the direction of the inference target, and among the multiple inclusion features (701-1 to 701-t) repeatedly acquired by the acquisition process, an inclusion feature (701-(δ+1)) at a first time (δ+1) to an inclusion feature (701-(δ+1)) at a second time (t) are acquired. a generation process for generating a group of shift features (702) from a shift feature (702-(δ+1)) at a first time (δ+1) to a shift feature (702-t) at a second time (t) by combining an inclusion feature 701-(t-δ) at a third time before a predetermined time (δ) for each of the first time (δ) to the second time (701-t), and an output process for outputting the group of shift features (702) generated by the generation process to an inference model that infers the state of the inference target during the period from the first time (δ+1) to the second time (t).

[0187] As a result, during learning, the inference model can be learned using the shift feature set (702) as explanatory variables and the state set from the state at the first time (δ+1) to the state at the second time (t) as objective variables. Therefore, the learning accuracy of the inference model that infers states can be improved.

[0188] Furthermore, at the time of inference, the shift feature set (702) is input to the inference model, thereby making it possible to infer the state during the period from the first time (δ+1) to the second time (t). Therefore, it is possible to improve the inference accuracy of the inference model that infers the state.

[0189] (2) In the analysis system 100 described above in (1), the information regarding the movement of the inference object includes the movement speed of the inference object, and the information regarding the direction of the inference object includes the movement direction in which the inference object moves.

[0190] This enables learning and inference that takes into account the movement tendencies of the inference subject.

[0191] (3) In the analysis system 100 described in (2) above, the movement speed is a specific movement speed of a specific inference target among multiple movement speeds of multiple inference targets, and the movement direction is a specific movement direction in which the specific inference target moves at the specific movement speed.

[0192] This allows learning and inference of a state that takes into account a specific moving direction at a specific moving speed.

[0193] (4) In the analysis system 100 described in (3) above, the specific movement speed is the highest movement speed among the multiple movement speeds, and the specific movement direction is the movement direction in which the specific inference subject moves at the highest movement speed.

[0194] This enables learning and inference that takes into account the direction of movement at the maximum speed.

[0195] (5) In the analysis system 100 described in (2) above, the information regarding the movement of the inference target is a first one-dimensional histogram showing the number of people by movement speed, and the information regarding the direction of the inference target is a second one-dimensional histogram showing the number of people by movement direction, and in the acquisition process, the processor acquires the specific movement speed, the specific movement direction, the first one-dimensional histogram, and the second one-dimensional histogram as the inclusion features.

[0196] This enables learning and inference that takes into account the distribution of people moving at different speeds and in different directions.

[0197] (6) In the analysis system 100 described in (2) above, the information regarding the movement of the inference target is a first one-dimensional histogram showing the number of people by moving speed, and the information regarding the direction of the inference target is a second one-dimensional histogram showing the number of people by moving direction, and in the acquisition process, the processor acquires the specific moving speed, the specific moving direction, and a two-dimensional histogram that integrates the first one-dimensional histogram and the second one-dimensional histogram as the inclusive features.

[0198] This enables learning and inference that takes into account the distribution of people moving at different speeds and in different directions.

[0199] (7) In the analysis system 100 described in (2) above, the information regarding the movement of the inference target is a first one-dimensional histogram indicating the number of people by movement speed, and the information regarding the direction of the inference target is a second one-dimensional histogram indicating the number of people by movement direction, and the processor executes a determination process to determine whether to generate a two-dimensional histogram that integrates the first one-dimensional histogram and the second one-dimensional histogram based on the number of movement vectors indicating that the inference target is moving at the movement speed toward the movement direction. In the acquisition process, if the determination process determines not to generate the two-dimensional histogram, the processor 201 acquires the specific movement speed, the specific movement direction, the first one-dimensional histogram, and the second one-dimensional histogram as the inclusive features, and if the determination process determines to generate the two-dimensional histogram, the processor 201 acquires the specific movement speed, the specific movement direction, and the two-dimensional histogram as the inclusive features.

[0200] This makes it possible to improve the accuracy of state learning and inference.

[0201] (8) In the analysis system 100 described in (1) above, in the generation process, the processor generates a group of shift features from the shift feature at the first time to the shift feature at the second time by combining, for each of the plurality of inclusion features from the inclusion feature at the first time to the inclusion feature at the second time, a plurality of inclusion features at a third time that is a plurality of predetermined times earlier.

[0202] This makes it possible to realize highly accurate state inference using feature amounts from multiple past frames.

[0203] (9) In the analysis system 100 described in (1) above, the information regarding the movement of the inference subject includes information regarding the facial movement of the inference subject, and the information regarding the direction of the inference subject includes the gaze direction of the inference subject.

[0204] This enables learning and inference in a state that takes into account the visual tendency of the inference target, i.e., whether the target is looking at the visual target.

[0205] (10) In the analysis system 100 of (1) above, the processor 201 executes a learning process to learn the inference model using the shift feature set (702) as explanatory variables and the state set from the state at the first time (δ+1) to the state at the second time (t) as objective variables.

[0206] This makes it possible to improve the learning accuracy of the inference model that infers states.

[0207] (11) In the analysis system 100 of (1) above, the processor 201 executes an inference process to infer the state during the period from the first time (δ+1) to the second time (t) by inputting the shift feature group (702) into the inference model.

[0208] This makes it possible to improve the inference accuracy of the inference model that infers the state.

[0209] The present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the spirit and scope of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to configurations including all of the described configurations. Furthermore, part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Furthermore, the configuration of another embodiment may be added to the configuration of one embodiment. Furthermore, part of the configuration of each embodiment may be added to, deleted from, or replaced with other configurations.

[0210] Furthermore, the aforementioned configurations, functions, processing units, processing means, etc. may be realized in hardware, for example, by designing some or all of them as integrated circuits, or may be realized in software, by having processor 201 interpret and execute a program that realizes each function.

[0211] Information such as programs, tables, and files that realize each function can be stored in a storage device such as a memory, hard disk, or SSD (Solid State Drive), or in a recording medium such as an IC (Integrated Circuit) card, SD card, or DVD (Digital Versatile Disc).

[0212] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines that are necessary for implementation. In reality, it can be assumed that almost all components are interconnected.

Claims

1. An analysis system having a processor that executes a program and a storage device that stores the program, wherein the processor executes the following: an acquisition process that repeatedly acquires inclusion features that include information related to the movement of an inference object and information related to a direction of the inference object; a generation process that generates a group of shift features from the shift feature at the first time to the shift feature at the second time by combining an inclusion feature at a third time that is a predetermined time before for each of a plurality of inclusion features repeatedly acquired by the acquisition process, from the inclusion feature at a first time to the inclusion feature at a second time, and an output process that outputs the group of shift features generated by the generation process to an inference model that infers a state of the inference object during the period from the first time to the second time.

2. An analysis system as described in claim 1, characterized in that the information regarding the movement of the inference object includes a moving speed of the inference object, and the information regarding the direction of the inference object includes a moving direction in which the inference object moves.

3. An analysis system as described in claim 2, characterized in that the movement speed is a specific movement speed of a specific inference target among multiple movement speeds of multiple inference targets, and the movement direction is a specific movement direction in which the specific inference target moves at the specific movement speed.

4. An analysis system as described in claim 3, characterized in that the specific movement speed is the maximum movement speed among the multiple movement speeds, and the specific movement direction is the movement direction in which the specific inference subject moves at the maximum movement speed.

5. An analysis system as described in claim 3, wherein the information regarding the movement of the inference target is a first one-dimensional histogram indicating the number of people by moving speed, and the information regarding the direction of the inference target is a second one-dimensional histogram indicating the number of people by moving direction, and in the acquisition process, the processor acquires the specific moving speed, the specific moving direction, the first one-dimensional histogram, and the second one-dimensional histogram as the inclusion features.

6. An analysis system as described in claim 3, wherein the information regarding the movement of the inference target is a first one-dimensional histogram indicating the number of people by moving speed, and the information regarding the direction of the inference target is a second one-dimensional histogram indicating the number of people by moving direction, and in the acquisition process, the processor acquires the specific moving speed, the specific moving direction, and a two-dimensional histogram integrating the first one-dimensional histogram and the second one-dimensional histogram as the inclusion features.

7. An analysis system as described in claim 3, wherein the information regarding the movement of the inference target is a first one-dimensional histogram indicating the number of people by moving speed, and the information regarding the direction of the inference target is a second one-dimensional histogram indicating the number of people by moving direction, and the processor executes a judgment process to judge whether or not to generate a two-dimensional histogram by integrating the first one-dimensional histogram and the second one-dimensional histogram based on the number of movement vectors indicating that the inference target is moving at the moving speed toward the moving direction, and in the acquisition process, if the judgment process determines not to generate the two-dimensional histogram, the processor acquires the specific moving speed, the specific moving direction, the first one-dimensional histogram, and the second one-dimensional histogram as the inclusion features, and if the judgment process determines to generate the two-dimensional histogram, the processor acquires the specific moving speed, the specific moving direction, and the two-dimensional histogram as the inclusion features.

8. An analysis system as described in claim 1, characterized in that in the generation process, the processor generates a group of shift features from the shift feature at the first time to the shift feature at the second time by combining, for each of the multiple inclusion features from the inclusion feature at the first time to the inclusion feature at the second time, multiple inclusion features at a third time that is multiple predetermined times earlier.

9. An analysis system as described in claim 1, characterized in that the information regarding the movement of the inference subject includes information regarding the facial movement of the inference subject, and the information regarding the direction of the inference subject includes the gaze direction of the inference subject.

10. An analysis system as described in claim 1, characterized in that the processor executes a learning process to learn the inference model using the group of shift features as explanatory variables and a group of states from the state at the first time to the state at the second time as objective variables.

11. An analysis program that causes a processor to execute: an acquisition process that repeatedly acquires inclusion features that include information regarding the movement of an inference object and information regarding the direction of the inference object; a generation process that generates a group of shift features from the shift feature at the first time to the shift feature at the second time by combining an inclusion feature at a third time that is a predetermined time earlier with an inclusion feature at a first time to an inclusion feature at a second time among the multiple inclusion features repeatedly acquired by the acquisition process; and an output process that outputs the group of shift features generated by the generation process to an inference model that infers the state of the inference object during the period from the first time to the second time.

Citation Information

Patent Citations

  • Suspicious behavior detection system

    JP2020149389A

  • Analysis system and analysis program

    JP2025088915A

  • Depth neural network interpretable method, visualization method and related device

    CN114419726A

  • State determining device, state determining method and program

    JP2012212407A

  • Moving object recognition system, moving object recognition program, and moving object recognition method

    JP2014029604A