Analysis system and analysis program

The analysis system enhances state inference accuracy by generating shift feature amounts from inclusion feature amounts over time, effectively addressing the challenges of analyzing movement in dense crowds or environments with obstacles.

JP2025088915APending Publication Date: 2025-06-12HITACHI LTD

Patent Information

Application Number
JP2023203745
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing analysis systems face challenges in accurately analyzing movement and shape information of individuals in dense crowds or environments with many obstacles, leading to difficulties in learning and inferring non-steady states.

Method used

The proposed analysis system employs a processor that repeatedly acquires inclusion feature amounts, generates shift feature amounts by combining inclusion feature amounts from multiple time points, and outputs these shift feature amounts to an inference model to infer the state of the inference target over a period.

Benefits of technology

This approach improves the learning accuracy and inference accuracy of the inference model, enabling more precise state inference in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025088915000001_ABST
    Figure 2025088915000001_ABST
Patent Text Reader

Abstract

To improve learning accuracy and inference accuracy of an inference model that infers a state of a space.SOLUTION: An analysis system has a processor that executes a program and a storage device that stores the program. The processor executes: an acquisition process for repeatedly acquiring a containment feature value which contains information pertaining to the motion of an inference target and information pertaining to a direction of the inference target; a generation process for generating a shift feature value group, from a shift feature value at a first time to a shift feature value at a second time, by combining, with each containment feature value from the containment feature value at the first time to the containment feature value at the second time among the plurality of containment feature values repeatedly acquired in the acquisition process, the containment feature value at a third time prior to a prescribed time; and an output process for outputting the shift feature value group generated in the generation process to an inference model which infers the state of the inference target in the period from the first time to the second time.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an analysis system and an analysis program for analyzing data.

Background Art

[0002] As the background art in this technical field, Patent Document 1 discloses a human detection unit and an object detection unit for detecting humans and objects, and a feature quantity inference unit for inferring a feature quantity based on the movement and shape information of the human detected by the human detection unit and the movement information of the object detected by the object detection unit, a steady state model storage unit for storing a steady state model in advance, a steady state model inference unit for learning the feature quantity of the steady state by the feature quantity inference unit and the steady state model storage unit, a non-steady state inference unit for determining the deviation from the steady state model using the feature quantity calculated by the feature quantity inference unit and the data of the steady state model storage unit, and the non-steady state inference unit infers a non-steady state.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the technology disclosed in Patent Document 1 above, in an environment with a dense crowd or many obstacles, it is difficult to accurately analyze the movement and shape information of a person and to learn or infer a non-steady state.

[0005] The present invention has been made in view of the above problems, and an object thereof is to improve the learning accuracy and inference accuracy of an inference model for inferring the state of a space.

Means for Solving the Problems

[0006] The analysis system, which is the first disclosed technology of the present application, is an analysis system having a processor that executes a program and a storage device that stores the program. The processor repeatedly performs an acquisition process of repeatedly acquiring an inclusion feature amount including information regarding the movement of the inference target and information regarding the direction of the inference target, and for each of the inclusion feature amounts from the inclusion feature amount at the first time to the inclusion feature amount at the second time among the plurality of inclusion feature amounts repeatedly acquired by the acquisition process, a generation process of generating a group of shift feature amounts from the shift feature amount at the first time to the shift feature amount at the second time by combining the inclusion feature amount at the third time before a predetermined time, and an output process of outputting the group of shift feature amounts generated by the generation process to an inference model that infers the state of the inference target during the period from the first time to the second time.

Advantages of the Invention

[0007] According to a typical embodiment of the present invention, problems, configurations, and effects other than those described above for improving the learning accuracy and inference accuracy of an inference model for inferring the state of space will be clarified by the description of the following examples.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings for explaining the embodiments, the same members are generally denoted by the same reference numerals, and repeated explanations thereof are omitted. Further, in the following embodiments, it goes without saying that the constituent elements (including element steps, etc.) are not necessarily essential except in cases where it is particularly specified and cases where it is considered clearly essential in principle. Also, when it is said that "comprising A", "consisting of A", "having A", or "including A", it goes without saying that other elements are not excluded except in cases where it is particularly specified that only that element is present. Similarly, in the following embodiments, when referring to the shape, positional relationship, etc. of constituent elements, etc., it is assumed to include those that are substantially approximate or similar to the shape, etc. except in cases where it is particularly specified and cases where it is considered clearly not so in principle.

[0010] Expressions such as "first", "second", "third", etc. in this specification and the like are attached to identify constituent elements, and do not necessarily limit the number, order, or content thereof. Also, the numbers for identifying constituent elements are used for each context, and the numbers used in one context do not necessarily indicate the same configuration in other contexts. Also, it does not prevent a constituent element identified by a certain number from having the functions of a constituent element identified by another number.

[0011] The positions, sizes, shapes, ranges, etc. of each configuration shown in the drawings and the like may not represent the actual positions, sizes, shapes, ranges, etc. in order to facilitate understanding of the invention. For this reason, the present invention is not necessarily limited to the positions, sizes, shapes, ranges, etc. disclosed in the drawings and the like.

Example

[0012] <Figure 1 Analysis System> Figure 1 is an explanatory diagram showing a system configuration example of the analysis system according to Example 1. The analysis system 100 includes a server 101, one or more clients 102, and a terminal 106. The server, clients, and terminal are communicably connected via a network 105 such as the Internet, a LAN (Local Area Network), or a WAN (Wide Area Network). The server 101 is a computer that manages the clients 102. The client 102 is a computer connected to a sensor 103 and acquires data from the sensor 103. The terminal 106 receives and transmits signals from and to the client 102, and inputs and outputs operations and data to and from the server and the client.

[0013] The sensor 103 detects analysis target data from the analysis environment. The sensor 103 is, for example, a camera that captures a still image or a moving image. The camera may be capable of acquiring depth information. The sensor 103 performs fixed-point imaging of the analysis environment. Therefore, for example, there are subjects whose shape and size change between time-series frames output as analysis target data from the sensor 103, and there are also congruent subjects whose shape and size do not change between frames. Also, even among congruent subjects between frames, the colors may be different.

[0014] The teacher signal DB 104 is a database that holds, as teacher signals, combinations of image data (also referred to as frames) that are learning data of the inspection target and the states of the learning data (for example, correct answer data representing two or more states such as "normal" and "abnormal").

[0015] The inspection target is an area whose state is inferred by fixed-point photography, for example, an area such as the space inside a train where state inference is required based on the actions of people in the space. The teacher signal DB104 may be stored in the server 101 or may be connected to a computer that can communicate with the server 101 or the client 102 via the network 105.

[0016] The analysis system 100 has a learning function using the teacher signal DB104 and an inference function using the inference model obtained by the learning function. The inference model is a learning model for inferring the state of the object reflected in the image data. The learning function and the inference function may be implemented in either the server 101 or the client 102 as long as they are implemented in the analysis system 100. For example, the server 101 may implement the learning function and the client 102 may implement the inference function. Also, the server 101 may implement the learning function and the inference function, and the client 102 may send data from the sensor 103 to the server 101 or receive the inference result from the inference function of the server 101.

[0017] Alternatively, the client 102 may implement the learning function and the inference function, and the server 101 may manage the inference model and the inference result from the client 102. Note that a computer implementing the learning function is referred to as a learning device, and a computer implementing at least the inference function among the learning function and the inference function is referred to as an inference device. Also, in FIG. 1, the client-server type analysis system 100 is taken as an example, but a stand-alone type inference device may also be used. In the first embodiment, for convenience of explanation, the analysis system 100 in which the server 101 implements the learning function and the client 102 implements the inference function will be described as an example.

[0018] <Figure 2 Example of the hardware configuration of a computer> FIG. 2 is a block diagram showing an example of the hardware configuration of computers (server 101, client 102). Computer 200 includes a processor 201, a storage device 202, an input device 203, an output device 204, and a communication interface (communication IF) 205. The processor 201, the storage device 202, the input device 203, the output device 204, and the communication IF 205 are connected by a bus 206. The processor 201 controls the computer 200. The storage device 202 serves as a working area for the processor 201. Also, the storage device 202 is a non-temporary or temporary recording medium that stores various programs and data. Examples of the storage device 202 include a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), and a flash memory. The input device 203 inputs data. Examples of the input device 203 include a keyboard, a mouse, a touch panel, a numeric keypad, and a scanner. The output device 204 outputs data. Examples of the output device 204 include a display, a printer, and a speaker. The communication IF 205 connects to the network 105 and transmits and receives data.

[0019] <FIG. 3 Functional Configuration Example of Analysis System 100> FIG. 3 is a block diagram showing a functional configuration example of the analysis system 100 according to the first embodiment. The server 101 includes a person detection unit 300, a tracking processing unit 301, a vector information generation unit 302, a speed distribution calculation unit 303, a direction distribution calculation unit 304, a histogram feature amount generation unit 305, a shift feature amount generation unit 306, and a learning unit 307. The client 102 includes a person detection unit 310, a tracking processing unit 311, a vector information generation unit 312, a speed distribution calculation unit 313, a direction distribution calculation unit 314, a histogram feature amount generation unit 315, a shift feature amount generation unit 316, an inference unit 317, and an inference result output unit 318. The terminal 106 includes an inference result reception unit 320.

[0020] These are specifically realized, for example, by causing the processor 201 to execute the program stored in the storage device 202 shown in FIG. 2. In FIG. 3, although the teacher signal DB104 exists outside the server 101, it may be realized by the storage device 202 in the server 101 or the client 102.

[0021] <Functional configuration example on the server 101 side> First, a functional configuration example on the server 101 side will be described. The human detection unit 300 generates rectangular information R = (x1, y1, x2, y2) indicating the presence of a person for the time-series learning data acquired from the teacher signal DB104, and outputs the center point of the rectangular information R as the position coordinates O = (x, y) of the person to the tracking processing unit 301. When there are multiple persons in the learning data, multiple position coordinates O are generated. Note that x1 and y1 of the rectangular information R represent the upper left image coordinates of the rectangular information R, and x2 and y2 represent the lower right image coordinates of the rectangle. The teacher signal is composed of learning data and a state. The details of the teacher signal will be described with reference to FIG. 8.

[0022] The tracking processing unit 301 generates a unique trajectory ID for the person based on the position coordinates O of the person acquired from the human detection unit 300. The trajectory IDs of the multiple position coordinates O derived from the same person among the acquired multiple learning data have the same value. The tracking processing unit 301 outputs the trajectory ID and the position coordinates O of the person associated with the trajectory ID to the vector information generation unit 302.

[0023] The vector information generation unit 302 generates a movement vector based on the acquired trajectory ID and position coordinates O, and outputs it to the speed distribution calculation unit 303 and the direction distribution calculation unit 304.

[0024] [FIG. 4 Example of generation of trajectory ID and movement vector] FIG. 4 is an explanatory diagram showing an example of generating a trajectory ID and a movement vector. Reference numeral 400-t denotes the image data (frame) acquired as the t-th learning data. In FIG. 4, reference numerals 401(t), 402(t), 401(t-u), and 402(t-u) are rectangular information R including human image data, but for convenience of explanation, they are denoted as humans 401(t), 402(t), 401(t-u), and 402(t-u).

[0025] The person detection unit 300 detects humans 401(t) and 402(t), and outputs their respective position coordinates O to the tracking processing unit 301. Frame 400-(t-u) is the image data acquired at the (t-u)-th time. The person detection unit 300 detects humans 401(t-u) and 402(t-u), and outputs their respective position coordinates O to the tracking processing unit 301.

[0026] The tracking processing unit 301 assigns the same trajectory ID to the same person for the humans detected in frames 400-t and 400-(t-u) acquired at different times t and t-u. For example, the tracking processing unit 301 regards human 401(t) and human 401(t-u) as the same person, and assigns the trajectory ID at time t (frame 400-t acquired at the t-th time) to human 401(t) so that it is possible to distinguish at which time the person was detected. t is assigned to human 401(t), and the trajectory ID at time t-u (frame 400-(t-u) acquired at the (t-u)-th time) t-u is assigned to human 401(t-u).

[0027] Similarly, the tracking processing unit 301 regards human 402(t) and human 402(t-u) as the same person, and assigns the trajectory ID at time t (frame 400-t acquired at the t-th time) to human 402(t) so that it is possible to distinguish at which time the person was detected. t is assigned to human 402(t), and the trajectory ID at time t-u (frame 400-(t-u) acquired at the (t-u)-th time) t-u is assigned to human 402(t-u).

[0028] Also, in frame 400-t, the position coordinates O of person 401(t) are O i t =(x i t ,y i t ), and in frame 400-(t-u), the position coordinates O of person 401(t-u) are O i t-u =(x i t-u ,y i t-u ).

[0029] Similarly, in frame 400-t, the position coordinates O of person 402(t) are O j t =(x j t ,y j t ), and in frame 400-(t-u), the position coordinates O of person 402(t-u) are O j t-u =(x j t-u ,y j t-u ).

[0030] Note that the tracking processing unit 301 may use authentication processing as a method for identifying a person between different frames, or may use tracking processing based on the amount of movement between frames using a Kalman filter or the like. The person identification method of the tracking processing unit 301 is not limited. Also, the specific values of t and u indicating the acquisition order of the frames for performing the tracking processing are not limited. However, since the accuracy of the tracking processing between frames is considered to deteriorate as the value of u increases, it is preferable that the value of u be set to a value smaller than a predetermined value.

[0031] The composite frame 410 is image data obtained by overlapping frame 400-t and frame 400-(t-u). The vector information generation unit 302, in the composite frame 410, the position coordinates O of person 401(t) in frame 400-t i t and the position coordinates O of person 401(t-u) in frame 400-(t-u) it-u Using this, as the movement vector 420 of person 401(t), the movement vector V i t is generated.

[0032] Similarly, the vector information generation unit 302, in the composite frame 410, uses the position coordinates O of person 402(t) in frame 400 - t j t and the position coordinates O of person 402(t - u) in frame 400 - (t - u) j t-u to generate, as the movement vector 420 of person 402(t), the movement vector V j t is generated.

[0033] Here, the general formula of the movement vector V i t is defined. Using the coordinates of the rectangular information R, the speed n i t and the direction d i t are defined by the following formulas (1) and (2). The movement vector V i t is defined by the following formula (3).

[0034]

Equation

[0035] Returning to the description of FIG. 3. The speed distribution calculation unit 303 calculates, from the obtained movement vector V i t the maximum speed n in the learning data at each time i t fastest and the maximum speed n i t fastest of the movement vector V i t fastest and the speed distribution 501 (described later in FIG. 5) of all persons to whom the trajectory ID is assigned, and outputs them to the histogram feature amount generation unit 305. Also, the speed distribution calculation unit 303 calculates the movement vector V i tfastest Output it to the direction distribution calculation unit 304.

[0036] Based on the movement vector V obtained from the vector information generation unit 302 i t and the maximum speed n obtained from the speed distribution calculation unit 303 i t fastest of the movement vector V i t fastest and, for all persons assigned with a trajectory ID, the direction distribution 502 (to be described later in FIG. 5), and the maximum speed n i t fastest of the movement vector V i t fastest calculate the direction d i t fastest and output it to the histogram feature quantity generation unit 305.

[0037] [FIG. 5 Speed distribution 501 and direction distribution 502] FIG. 5 is an explanatory diagram showing the speed distribution 501 and the direction distribution 502. The speed distribution 501 is a histogram with the horizontal axis being the trajectory ID and the vertical axis being the speed n i t and. The speed distribution 501 exists for each time t. The direction distribution 502 is a histogram with the horizontal axis being the trajectory ID and the vertical axis being the direction d i t and. The direction distribution 502 also exists for each time t.

[0038] In the speed distribution 501, the trajectory ID of the maximum speed n i t fastest is "9". In the direction distribution 502, the direction d of the trajectory ID = 9 i t is the direction d of the maximum speed n i t fastest of i t fastest is.

[0039] Return to the description of FIG. 3. For the speed distribution 501 and direction distribution 502 of people in each acquired frame 400-t, the speed n i t and direction d i t histogram feature amount H t is calculated.

[0040] [FIG. 6 Histogram feature amount H t FIG. 6 is an explanatory diagram showing an example of calculating the histogram feature amount H t at time t. In FIG. 6, the first one-dimensional histogram 601 is a one-dimensional histogram with the horizontal axis representing the speed n i t and the vertical axis representing the number of people. The second one-dimensional histogram 602 is a one-dimensional histogram with the horizontal axis representing the direction d i t and the vertical axis representing the number of people. The two-dimensional histogram 603 is a two-dimensional histogram with the x-axis representing the speed n i t , the y-axis representing the direction d i t , and the z-axis representing the number of people. The first one-dimensional histogram 601, the second one-dimensional histogram 602, and the two-dimensional histogram 603 are generated by the histogram feature amount generation unit 305.

[0041] The number of people for each speed n i t in the first one-dimensional histogram 601 of the speed n i t , the number of people for each direction d i t in the second one-dimensional histogram 602 of the direction d i t , and the combination of the number of people for each direction d t are referred to as the one-dimensional histogram feature amount H1 i t at time t. The speed n i t and direction d i t in the two-dimensional histogram 603 of the speed n i t ​The number of people for each combination is the two-dimensional histogram feature amount H2 at time t t and is referred to as such. If the one-dimensional histogram feature amount H1 t and the two-dimensional histogram feature amount H2 t are not distinguished, it is referred to as the histogram feature amount H t and is so called.

[0042] The histogram feature amount generation unit 305 uses, for each frame 400-t, the histogram feature amount Ht and the maximum speed n i t fastest of the movement vector V i t fastest and the maximum speed n i t fastest and the direction d i t fastest in it as inclusion feature amounts and outputs them to the shift feature amount generation unit 306.

[0043] Return to FIG. 3. The shift feature amount generation unit 306 combines the inclusion feature amount of each acquired frame 400-t with the inclusion feature amount of the past frame several predetermined frames before the inclusion feature amount of the frame 400-t, and outputs it to the learning unit 307 as a shift feature amount. By applying the shift feature amount to learning, it becomes possible to prevent several types of states from being mixed up when judging the continuous state of space in frame units. Specifically, for example, in a video at 10 fps, it is unlikely that the situation inside a train changes in frame units (every 0.1 second), and it is expected that the situation changes in a lump over a certain period of time. Using only the inclusion feature amount of the frame 400-t may cause a single-shot misjudgment. Therefore, by combining the inclusion feature amounts of the past frames several predetermined frames before to generate a shift feature amount, such single-shot misjudgments can be suppressed.

[0044] [FIG. 7 Shift Feature Amount] FIG. 7 is an explanatory diagram showing an example of generating shift feature amounts. The time-series learning data group 700 is composed of a plurality of time-series frames, specifically, for example, the first frame 400-1, …, the δ-th frame 400-δ, the (δ + 1)-th frame 400-(δ + 1), the (δ + 2)-th frame 400-(δ + 2), …, the (t - 1)-th frame 400-(t - 1), and the t-th frame 400-t. δ < t is a non-negative integer.

[0045] The inclusion feature amount group 701 is a set of inclusion feature amounts 701-1 to 701-t for each frame of the time-series learning data group 700.

[0046] The shift feature amount group 702 is a set of shift feature amounts 702-(δ + 1) to 702-t accumulated in time series, which are generated by the shift feature amount generation unit 306. The shift feature amount generation unit 306 uses the inclusion feature amounts 701-(δ + 1) to 701-t of the frames 400-(δ + 1) to 400-t and the inclusion feature amounts 701-1 to 701-(t - δ) of the frames 400-1 to 400-(t - δ) δ frames before to generate the shift feature amounts 702-(δ + 1) to 702-t of the frames 400-(δ + 1) to 400-t. The specific value of δ is not limited.

[0047] For example, the inclusion feature amount 701-t generated from the t-th frame 400-t is the maximum speed n i t fastest and the maximum speed n i t fastest of the movement vector V i t fastest in the direction d i t fastest and the histogram feature amount H t and is composed of.

[0048] Similarly, the inclusion feature amount 701-(t - δ) of the (t - δ)-th frame 400-(t - δ) is the maximum speed n i t-δ fastest and the maximum speed n it fastest The movement vector V i t fastest The direction d i t fastest and the histogram feature amount H t-δ are constituted by

[0049] Therefore, the shift feature amount 702-t generated from the t-th frame 400-t is the maximum speed n in the t-th frame 400-t i t fastest and the maximum speed n i t fastest The movement vector V i t fastest The direction d i t fastest and the histogram feature amount H t and the maximum speed n in the (t - δ)-th frame i t-δ fastest and the maximum speed n i t-δ fastest The movement vector V i t-δ fastest The direction d i t-δ fastest and the histogram feature amount H t-δ and are constituted by. However, for the frames 400-1 to 400-δ from the 1st to the δ-th frame, since there are no past frames δ frames before, the generation process of the shift feature amount is executed for frames 400-(δ + 1) and later.

[0050] Return to the explanation of FIG. 3. The learning unit 307 uses the acquired shift feature amount group 702 as an explanatory variable and the state of the learning data which is frames 400-(δ + 1) to 400-t as an objective variable, performs learning, and generates an inference model for inferring the state of the target (person) within frames 400-(δ + 1) to 400-t from the shift feature amount group 702 by machine learning. The learning unit 307 outputs the inference model generated by learning to the inference unit 317.

[0051] <Functional configuration example on the client 102 side> Next, a functional configuration example on the client 102 side will be described. The human detection unit 310 acquires analysis target data from the sensor 103, detects a person from the analysis target data by using the result of previously learning human information through machine learning, generates rectangular information R = (x1, y1, x2, y2) indicating the presence of a person, and outputs the center point of the rectangular information R as the position coordinates O = (x, y) of the person to the tracking processing unit 311.

[0052] Note that the frame 400-t as learning data in the teacher signal DB104 and the frame acquired from the sensor 103 as analysis target data are different in both time and image data. For convenience of explanation, on the client 102 side as well, t is used as the time code and 400 is used as the frame code.

[0053] The tracking processing unit 311 generates a unique trajectory ID for the person based on the position coordinates O of the person acquired from the human detection unit 310. The trajectory IDs of the plurality of position coordinates O derived from the same person among the acquired plurality of learning data have the same value. The tracking processing unit 301 outputs the trajectory ID and the position coordinates O of the person associated with the trajectory ID to the vector information generation unit 312.

[0054] The vector information generation unit 312 generates a movement vector 420 based on the acquired trajectory ID and position coordinates O, and outputs it to the speed distribution calculation unit 313 and the direction distribution calculation unit 314.

[0055] The speed distribution calculation unit 313 calculates the maximum speed n in the frame 400-t at each time t from the acquired movement vector V i t and the movement vector V of the maximum speed n i t fastest and the maximum speed n i t fastest of the movement vector V i t fastestThen, the speed distribution 501 of all the people with trajectory IDs is calculated and output to the histogram feature quantity generation unit 315. Also, the speed distribution calculation unit 313 outputs the movement vector V i t fastest to the direction distribution calculation unit 314.

[0056] The direction distribution calculation unit 314 is based on the movement vector V i t acquired from the vector information generation unit 312, and the maximum speed n i t fastest of the movement vector V i t fastest to calculate the direction distribution 502 of all the people with trajectory IDs and the direction d i t fastest of the movement vector V i t fastest of the maximum speed n i t fastest and output them to the histogram feature quantity generation unit 315.

[0057] The histogram feature quantity generation unit 315 calculates the histogram feature quantity H i t relating to the speed n i t and the direction d t for the speed distribution 501 and direction distribution 502 of people in each acquired frame 400-t. The histogram feature quantity generation unit 315 outputs, for each frame 400-t, the histogram feature quantity H t and the maximum speed n i t fastest of the movement vector V i t fastest and the maximum speed n i t fastest and the direction d i t fastest as inclusion features to the shift feature quantity generation unit 316.

[0058] The shift feature quantity generation unit 316 combines the acquired inclusion feature quantity 701-t of each frame 400-t with the inclusion feature quantity 701-t of the frame 400-t and the inclusion feature quantity 701-(t-δ) of the past frame 400-(t-δ) δ frames before the time t, and outputs it to the inference unit 317 as the shift feature quantity 702-t.

[0059] For example, the shift feature quantity of the t-th frame 400-t acquired from the sensor 103 as the data to be analyzed is the maximum speed n in the t-th frame 400-t i t fastest and the maximum speed n i t fastest of the movement vector V i t fastest of the direction d i t fastest and the histogram feature quantity H t and the maximum speed n in the 400-(t-δ) of the (t-δ)-th i t-δ fastest and the maximum speed n i t-δ fastest of the movement vector V i t-δ fastest of the direction d i t-δ fastest and the histogram feature quantity H t-δ and is more composed. However, for the frames 400-1 to 400-δ from the t = 1st to the δ-th, since there is no past frame δ frames before, the generation process of the shift feature quantity 702-t is executed for the frames after the (δ + 1)-th frame 400-(δ + 1).

[0060] The inference unit 317 inputs the acquired shift feature quantity group 702, which is the shift feature quantities 702-(δ + 1) at time (δ + 1) to 702-t at time t, into the inference model acquired from the learning unit 307, and infers the state of the target (person) within the frames 400-(δ + 1) to 400-t based on the shift feature quantity group 702. The inference unit 317 outputs the inference result to the inference result output unit 318.

[0061] For the inference result acquired by the inference result output unit 318 from the inference unit 317, if it is a predetermined inference result, the inference result output unit 318 transmits the inference result to the inference result receiving unit 320 within the terminal 106. Here, the predetermined inference result indicates, for example, a normal or abnormal state as shown in FIG. 9 or a state where additional control is required. By transmitting the predetermined inference result, necessary control can be performed.

[0062] <Functional configuration example on the terminal 106 side> Next, a functional configuration example on the terminal 106 side will be described. The inference result receiving unit 320 indicates that an additional control is required according to the inference result received from the inference result output unit 318.

[0063] <FIG. 8 Method of storing learning data> FIG. 8 is an explanatory diagram showing an example of a method of storing learning data. The learning data is stored in the teacher signal storage folder 800 of the teacher signal DB104. The teacher signal storage folder 800 has a first subfolder 810, a second subfolder 820, and a third subfolder 830.

[0064] The first subfolder 810 stores a subfolder 811 and a subfolder 812. The subfolder 811 stores the image data of 000.jpg811a and the image data of 001.jpg811b. The subfolder 812 stores the image data of 000.jpg812b and the image data of 001.jpg812b.

[0065] In the second subfolder 820, a subfolder 821 and the image data of 012.avi822 are stored. In the subfolder 821, the image data of 010.jpg821a and the image data of 011.jpg821b are stored.

[0066] In the third subfolder 830, the image data of 020.avi831, the image data of 021.avi832, and the image data of 012.avi833 are stored.

[0067] The time-series image data is stored in the first subfolder 810, the second subfolder 820, and the third subfolder 830 for each state. The state is used as the name of the subfolder. Specifically, for example, the state of the first subfolder 810 is "classA", the state of the second subfolder 820 is "classB", and the state of the third subfolder 830 is "classC". Also, below the subfolder of the state, the sets of image data belonging to different time series are stored in the subfolder 811, the subfolder 812, and the subfolder 821. The set is used as the name of the subfolder. Specifically, for example, the set of the subfolder 811 is "set01", the set of the subfolder 812 is "set02", and the set of the subfolder 821 is "set11".

[0068] Note that in Fig. 8, the number of subfolders divided by state is set to 3, but it may be 2 or more. Also, the number of image data and the number of subfolders in the first subfolder 810, the second subfolder 820, and the third subfolder 830 are not limited to those in Fig. 8.

[0069] The human detection unit 300 shown in FIG. 3 acquires time-series image data grouped by subfolder, together with the subfolder name indicating the state, and uses it for learning. For example, when learning an inference model in units of classes, by specifying the target subfolder, the server 101 acquires the image data in the subfolder, inputs it into the inference model, compares the output result with the subfolder name indicating the state, calculates the value of the loss function based on the comparison result (match or mismatch), and performs error backpropagation, enabling the learning of the inference model.

[0070] Also, when adding a new teacher signal to the teacher signal storage folder 800, the server 101 searches the teacher signal storage folder 800 by the state in the teacher signal to be added to identify the subfolder to be the storage destination, and stores the image data to be added. In this way, the generation and management of teacher signals are simplified.

[0071] <Example of Spatial Analysis by the Analysis System in FIG. 9> FIG. 9 is an explanatory diagram showing an example of spatial analysis by the analysis system. Images 901 and 902 are images displayed based on the image data captured by the sensor 103. Image 901 shows the state inside a train car as a space 900 where people are present. In the space 900 shown in Image 901, there may be specific tendencies in people's movement, such as "people are standing still", "walking slowly in a certain direction", "groups of people are uniformly distributed in the space", or "concentrated only in specific locations". The analysis system 100 can be set to determine such a state where people move with a specific tendency as a normal state. Note that such tendencies in people's movement may vary depending on the location and may also vary according to time zones and other conditions.

[0072] Such tendencies in people's movement can change dynamically due to events that occur suddenly (sudden events). Examples of events include events or accidents that threaten personal safety or make people want to leave the scene of an incident. Examples of events include the appearance of a violent person, the discovery of a suspicious object, or the occurrence of vomiting. Examples of accidents include vehicle breakdowns.

[0073] In the image 902 of FIG. 9, in the space 900 shown in the image 901, it shows how the movement tendency of people changes due to the presence of a sudden event. When a sudden event occurs, as shown in the image 902, people will move in a direction d2 different from the direction d1 shown in the image 901 so as to move away from the source 921 of the sudden event. The direction d2 is the longitudinal direction of the vehicle of the train. That is, people will move along the direction d2 to move to the adjacent vehicle. Such a tendency will continue as long as the sudden event occurs in the space 900. The analysis system 100 determines the state in which people move in a tendency different from such a normal state as an abnormal state.

[0074] The analysis system 100 can determine the normal state and abnormal state of the space 900 according to the movement tendency of people in the space 900. In addition, when the analysis system 100 determines an abnormal state, it is possible to identify the source of the sudden event caused by the abnormal state by using information such as the moving direction of people and the density of people gathering. According to the analysis system 100, for example, the system administrator can know the occurrence of a sudden event in the space 900, protect the safety of people in the space 900, and resolve the sudden event at an early stage.

[0075] <Figure 10 Learning Process> FIG. 10 is a flowchart showing a detailed processing procedure example of the learning process by the server 101 (learning device) according to the first embodiment.

[0076] (Step S1000) The server 101 acquires a teacher signal for learning from the teacher signal DB104 by the person detection unit 300.

[0077] (Step S1001) Next, the server 101 detects people from the acquired teacher signal for learning by the person detection unit 300, and calculates the position coordinates O of each frame of the people.

[0078] (Step S1002) Next, the server 101 assigns a trajectory ID to the acquired position coordinates O of the person by the tracking processing unit 301. The trajectory IDs of the plurality of position coordinates O derived from the same person among the acquired plurality of frames are set to the same value.

[0079] (Step S1003) Next, the server 101 calculates a movement vector 420 in each frame of the person to whom the trajectory ID has been assigned by the vector information generation unit 302. The movement vector 420 includes a speed n i t and a direction d i t is included.

[0080] (Step S1004) Next, the server 101 identifies the movement vector V with the highest speed n in each frame by the speed distribution calculation unit 303 i t and acquires the highest speed n in each frame i t fastest . Also, the direction distribution calculation unit 304 acquires the direction d of the movement vector V with the highest speed n i t fastest . i t The movement vector V with the highest speed n i t fastest of i t fastest is acquired.

[0081] (Step S1005) Next, the server 101 calculates the speed distribution 501 of the person to whom the trajectory ID has been assigned by the speed distribution calculation unit 303.

[0082] (Step S1006) Next, the server 101 calculates the direction distribution 502 of the person to whom the trajectory ID has been assigned by the direction distribution calculation unit 304.

[0083] (Step S1007) ​​Next, the server 101 uses the speed distribution 501 and the direction distribution 502 obtained by the histogram feature amount generation unit 305 to calculate the speed n i t and the direction d i t for each frame, and generates a histogram feature amount H t related to them.

[0084] (Step S1008) Next, the server 101 uses the histogram feature amount H t in each frame, the maximum speed n i t fastest and the maximum speed n i t fastest of the movement vector V i t fastest and the direction d i t fastest to generate the inclusion feature amount 701-t of each frame.

[0085] (Step S1009) Next, the server 101 accumulates the inclusion feature amounts 701-t of the acquired frames in time series by the shift feature amount generation unit 306.

[0086] (Step S1010) Next, the server 101 determines whether or not the accumulated inclusion feature amount 701-t has exceeded a preset number of frames δ by the shift feature amount generation unit 306. If the accumulated inclusion feature amount 701-t has not exceeded the preset number of frames δ (Step S1010: No), the process returns to the process of Step S1000. On the other hand, if the accumulated inclusion feature amount 701-t has exceeded the preset number of frames δ (Step S1010: Yes), the process proceeds to Step S1011.

[0087] (Step S1011) Next, the server 101 generates a shift feature amount 702-t by combining each of the accumulated inclusion feature amounts 701-t with the inclusion feature amount 701-(t-δ) of the past frame δ frames before the preset number of frames, by the shift feature amount generation unit 306.

[0088] (Step S1012) Next, the server 101 performs learning with the shift feature amount 702-t corresponding to the frame 400-t as an explanatory variable and the state corresponding to the frame 400-t as an objective variable by the learning unit 307, and generates an inference model for inferring the state of the target (person) within the frames 400-(δ+1) to 400-t by machine learning using the shift feature amount group 702.

[0089] <Figure 11 Inference Process> FIG. 11 is a flowchart showing an example of a state inference processing procedure by the client 102 (inference device) according to the first embodiment.

[0090] (Step S1100) The client 102 acquires the current frame from the sensor 103 by the person detection unit 310.

[0091] (Step S1101) Next, the client 102 detects a person from the acquired current frame by the person detection unit 310, and calculates the position coordinates O of each person.

[0092] (Step S1102) Next, the client 102 assigns a trajectory ID to the position coordinates O of the acquired person by the tracking processing unit 311. The trajectory IDs of the plurality of position coordinates O derived from the same person among the acquired plurality of frames are set to the same value.

[0093] (Step S1103) Next, the client 102 calculates the movement vector 420 of the person to whom the trajectory ID is assigned by the vector information generation unit 312. The movement vector 420 includes the speed n it and direction d i t is included.

[0094] (Step S1104) Next, the client 102 obtains the speed n from the movement vector 420 acquired from the vector information generation unit 312 by the speed distribution calculation unit 313 i t for the movement vector V with the highest i t fastest speed, and obtains the highest speed V i t fastest Also, the direction distribution calculation unit 304 obtains the direction d i t for the movement vector V with the highest i t fastest speed. i t fastest

[0095] (Step S1105) Next, the client 102 calculates the speed distribution 501 of the person to whom the trajectory ID is assigned by the speed distribution calculation unit 313.

[0096] (Step S1106) Next, the client 102 calculates the direction distribution 502 of the person to whom the trajectory ID is assigned by the direction distribution calculation unit 314.

[0097] (Step S1107) Next, the client 102 uses the speed distribution 501 and direction distribution 502 obtained by the histogram feature amount generation unit 315 to generate the histogram feature amount H i t relating to speed n i t and direction d t .

[0098] (Step S1108) ​Next, the client 102 uses the histogram feature quantity H t and the maximum speed n i t fastest and the maximum speed n i t fastest and the movement vector V i t fastest in the direction d i t fastest to generate the inclusion feature quantity 701-t of the current frame.

[0099] (Step S1109) Next, the client 102 accumulates the inclusion feature quantities 701-t of the acquired respective frames 400-t in time series by the shift feature quantity generation unit 316.

[0100] (Step S1110) Next, the client 102 determines by the shift feature quantity generation unit 316 whether or not the accumulated inclusion feature quantity 701-t has exceeded the preset number of frames δ. If the accumulated inclusion feature quantity 701-t has not exceeded the preset number of frames (Step S1110: No), the process returns to the process of Step S1100. On the other hand, if the accumulated inclusion feature quantity 701-t has exceeded the preset number of frames (Step S1110: Yes), the process proceeds to Step S1111.

[0101] (Step S1111) Next, the client 102 generates the shift feature quantity 702-t by combining, for each of the accumulated inclusion feature quantities 701-t, the inclusion feature quantity 701-t of the current frame 400-t and the inclusion feature quantity 701-(t-δ) of the past frame 400-(t-δ) δ frames before the preset number of frames by the shift feature quantity generation unit 316.

[0102] (Step S1112) Next, the client 102 provides the shift feature amount 702-t generated in step S1111 to the inference unit 317 as the input of the inference model acquired from the learning unit 307, and infers the state of the target (person) within the frames 400-(δ+1) to 400-t based on the shift feature amount group 702.

[0103] (Step S1113) Next, for the inference result of step S1112, if the inference result is a predetermined inference result, the client 102 transmits the inference result to the inside of the terminal 106 by the inference result output unit 318. The predetermined inference result is, for example, a specific classification class (for example, class C in FIG. 5) in a state where additional control is required. Note that the client 102 may also transmit the inference result to the terminal 106 even when the inference result is not a specific classification class.

[0104] (Step S1114) Next, the client 102 determines whether or not the acquisition of the analysis target data from the sensor 103 has ended. If the acquisition has not ended (step S1114: No), the process returns to step S1100. On the other hand, if the acquisition has ended (step S1114: Yes), the state inference process of the client 102 ends.

[0105] Whether or not the acquisition has ended is set in advance by the administrator. For example, when the acquisition start time and acquisition end time of the frame are set by the administrator, the client 102 moves to step S1100 until the acquisition end time (step S1114: No), and ends the state inference process when the acquisition end time is reached. Then, when the acquisition start time of the next business day is reached, the acquisition of the frame starts (step S1100). For example, when determining normality and abnormality inside a train through a security camera inside the train, the acquisition of the frame starts from the operation start time as the acquisition start time of the frame in advance so that the state determination function works during the train operation time zone, and the acquisition of the frame ends at the operation end time as the acquisition end time of the frame.

[0106] <Figure 12 Status Notification Process> Figure 12 is a flowchart showing an example of a status notification processing procedure by the terminal 106 (status notification device) according to the first embodiment.

[0107] (Step S1200) The terminal 106 receives the inference result of step S1112 from the client 102 by the inference result receiver 320.

[0108] (Step S1201) The terminal 106 notifies the user of the communication terminal of the status of the image data of the current frame by executing specific screen display processing or voice output processing according to the received inference result.

[0109] As described above, according to the first embodiment, even in an environment where people in the space 900 move densely or an environment with many obstacles, the state of the space 900 to be inspected can be accurately inferred by detecting the movement tendency of people.

[0110] In the inference process of FIG. 11, the client 102 transmits the inference result of step S1112 to the terminal 106 (step S1113), but it may also be transmitted to a patrol lamp that is communicably connected to the network 105. In this case, the patrol lamp lights up according to the received inference result, indicating that it is in a state where additional control is required.

Embodiment

[0111] The second embodiment will be described mainly focusing on the differences from the first embodiment. Regarding the points common to the first embodiment, the same reference numerals are given and the description thereof is omitted.

[0112] <Figure 13 Functional Configuration Example of the Analysis System 100> Figure 13 is a block diagram showing a functional configuration example of the analysis system 100 according to the second embodiment. In the first embodiment, the learning process executed by the server 101 (learning device) is executed by the client 102 in the second embodiment.

[0113] Therefore, in the second embodiment, the human detection unit 310 also functions as the human detection unit 300, the tracking processing unit 311 also functions as the tracking processing unit 301, the vector information generation unit 312 also functions as the vector information generation unit 302, the speed distribution calculation unit 313 also functions as the speed distribution calculation unit 303, the direction distribution calculation unit 314 also functions as the direction distribution calculation unit 304, the histogram feature amount generation unit 315 also functions as the histogram feature amount generation unit 305, and the shift feature amount generation unit 316 also functions as the shift feature amount generation unit 306.

[0114] When the human detection unit 310 outputs the position coordinates O of the detected person to the tracking processing unit 311, it attaches current processing information indicating whether the currently executed process is a learning process or an inference process.

[0115] The tracking processing unit 311, the vector information generation unit 312, the speed distribution calculation unit 313, the direction distribution calculation unit 314, and the histogram feature amount generation unit 315 execute the same processes as in the first embodiment.

[0116] The shift feature amount generation unit 316 generates a shift feature amount 702-t by combining the inclusion feature amount of each acquired frame with the inclusion feature amount 701-(t-δ) of a past frame δ frames before the inclusion feature amount 701-t of the frame. If the current processing information given by the human detection unit 310 is a learning process, the shift feature amount generation unit 316 outputs the calculated shift feature amount 702-t to the learning unit 307, and if the current processing information is an inference process, it outputs the calculated shift feature amount 702-t to the inference unit 317.

[0117] Thus, according to the second embodiment, by executing the learning process on the client 102, it becomes possible to consolidate the same processes that were respectively executed by the server 101 and the client 102, simplify the analysis system 100, and perform state inference in a stand-alone manner.

Embodiment

[0118] Example 3 will be described focusing on the differences from Example 1 and Example 2. The inference accuracy of the histogram feature quantity H with the highest inference accuracy t The inference accuracy based on may vary depending on the number of human movement vectors 420 detected from the data to be analyzed. Therefore, the analysis system 100 according to Example 3 can infer the state of the data to be analyzed with higher accuracy by switching the generation method of the histogram feature quantity H t according to the number of detected movement vectors 420. Regarding the points common to Examples 1 to 2, the same reference numerals are given and the description thereof is omitted.

[0119] <FIG. 14 Functional configuration example of the analysis system 100> FIG. 14 is a block diagram showing a functional configuration example of the analysis system 100 according to Example 3. In Example 3, a number measurement unit 309 is newly added to the server 101, and a number measurement unit 319 is newly added to the client 102.

[0120] The vector information generation unit 302 outputs the acquired trajectory ID and the movement vector 420 calculated from the position coordinates O to the number measurement unit 309. The vector information generation unit 312 outputs the acquired trajectory ID and the movement vector 420 calculated from the position coordinates O to the number measurement unit 319.

[0121] The number measurement unit 309 measures the number of movement vectors 420 output from the vector information generation unit 302 for each learning data, and outputs the measured number of movement vectors 420 as the number of people to the histogram feature quantity generation unit 305 and the learning unit 307.

[0122] The histogram feature quantity generation unit 305 uses the speed distribution 501 acquired from the speed distribution calculation unit 303 and the direction distribution 502 acquired from the direction distribution calculation unit 304 to generate the histogram feature quantity H t Specifically, when the number of people acquired from the number measurement unit 309 is larger than a predetermined number of people, the histogram feature quantity generation unit 305 generates a two-dimensional histogram feature quantity H2 tGenerate it, and when the number of people is less than a predetermined number, generate the one-dimensional histogram feature amount H1 t Generate it.

[0123] When the number of people obtained by the number measurement unit 309 is greater than the predetermined number of people, the learning unit 307 uses the shift feature amount 702-t obtained from the shift feature amount generation unit 306 as the explanatory variable of the inference model corresponding to the two-dimensional histogram feature amount H2 t Perform learning. When the number of people obtained by the number measurement unit 309 is less than the predetermined number of people, the learning unit 307 uses the shift feature amount 702-t obtained from the shift feature amount generation unit 306 as the explanatory variable of the inference model corresponding to the one-dimensional histogram feature amount H1 t Perform learning.

[0124] The number measurement unit 319 measures the number of movement vectors 420 output from the vector information generation unit 312 for each analysis target data 500, and outputs the measured number of movement vectors 420 as the number of people to the histogram feature amount generation unit 315 and the inference unit 317.

[0125] The histogram feature amount generation unit 315 uses the speed distribution 501 obtained from the speed distribution calculation unit 313 and the direction distribution 502 obtained from the direction distribution calculation unit 314 to generate the histogram feature amount H t Specifically, when the number of people obtained by the number measurement unit 319 is greater than the predetermined number of people, the histogram feature amount generation unit 315 generates the two-dimensional histogram feature amount H2 t Generate it, and when the number of people is less than a predetermined number, generate the one-dimensional histogram feature amount H1 t Generate it.

[0126] When the number of people obtained by the number measurement unit 319 is greater than the predetermined number of people, the inference unit 317 uses the inference model corresponding to the two-dimensional histogram feature amount H2 t To infer the state of the analysis target data. When the number of people obtained by the number measurement unit 319 is less than the predetermined number of people, the inference unit 317 uses the one-dimensional histogram feature amount H1 tInfer the state of the data to be analyzed by using the inference model corresponding thereto.

[0127] Since the bins in the two-dimensional histogram 603 are the product set of the bins in the first one-dimensional histogram 601 and the bins in the second one-dimensional histogram 602, the bins in the two-dimensional histogram 603 are divided more finely than the bins in the first one-dimensional histogram 601 and the bins in the second one-dimensional histogram 602. Therefore, the two-dimensional histogram 603 can reproduce the on-site situation more accurately. However, when the number of acquired people is less than or equal to a predetermined number, if the values are finely divided into each bin according to the feature amount, the height of each bin, such as approximately 0 in that bin, will be low, and the histogram feature amount will become sparse data in which the data is dispersed and sparse. There are many zero values in sparse data, and since there is little change in the data, appropriate learning cannot be performed on the input data. This is because, for example, when the data is sparse and the change in the data is small, the model is likely to overfit to the learning data with little data change and approximately zero values.

[0128] For example, let the bin of speed n in the first one-dimensional histogram 601 indicating speed be 0 < n ≤ 20, 20 < n ≤ 40, 40 < n ≤ 60,..., n > 300, and for the two-dimensional histogram 603, for example, let the bin of 0 < n ≤ 20 be classified into "0 < n ≤ 20 and 0 < angle ≤ 30", "0 < n ≤ 20 and 30 < angle ≤ 60",..., "0 < n ≤ 20 and 330 < angle ≤ 360" (the same applies to 20 < n ≤ 40...).

[0129] In this case, the total number of people in the two-dimensional histogram 603, the total number of people in the first one-dimensional histogram 601, and the total number of people in the second one-dimensional histogram 602 are the same. Since the bins in the two-dimensional histogram 603 are divided finely, the height of the bins in the two-dimensional histogram 603 is lower than that in the first one-dimensional histogram 601 and the second one-dimensional histogram 602.

[0130] In the case of the above classification example, in the first one-dimensional histogram 601 indicating speed, the moving directions of the people entering each bin (0 < n ≤ 20, 20 < n ≤ 40, 40 < n ≤ 60, …, n > 300) are unknown. In the two-dimensional histogram 603, since the moving directions of the people entering each bin (“0 < n ≤ 20 and 0 < angle ≤ 30”, “0 < n ≤ 20 and 30 < angle ≤ 60”, …, “0 < n ≤ 20 and 330 < angle ≤ 360”) are clear, it becomes possible to more accurately reproduce the on-site situation.

[0131] Therefore, when the number of acquired people is less than or equal to a predetermined number, the one-dimensional histogram feature amount H1 based on the first one-dimensional histogram 601 and the second one-dimensional histogram 602 t is applied for inference to suppress a decrease in inference accuracy, and when the number of acquired people is more than the predetermined number, the two-dimensional histogram feature amount H2 t is applied for inference to improve the inference accuracy.

[0132] <Figure 15 Learning Device> FIG. 15 is a flowchart showing a detailed processing procedure example of the learning process by the server 101 (learning device) according to the third embodiment. After step S1003, the server 101 calculates, as the number of people, the number of movement vectors 420 acquired from the vector information generation unit 302 by the number measurement unit 309 (step S1503). Then, the process proceeds to step S1004.

[0133] After step S1006, the server 101, by the histogram feature amount generation unit 305, when the number of people acquired from the number measurement unit 309 is more than a predetermined number, generates the two-dimensional histogram feature amount H2 t and when the number of people is less than or equal to the predetermined number, generates the one-dimensional histogram feature amount H1 t (step S1507). Then, the process proceeds to step S1008.

[0134] After step S1011, the server 101 sets, by the learning unit 307, the state constituting the teacher signal as the target variable. The server 101, by the learning unit 307, according to the number of people acquired from the number measurement unit 309, when the number is more than a predetermined number, sets the shift feature amount 702-t as the explanatory variable of the inference model corresponding to the two-dimensional histogram feature amount H2 t and when the number is less than or equal to the predetermined number, sets the shift feature amount 702-t as the explanatory variable of the inference model corresponding to the one-dimensional histogram feature amount H1 t . The server 101 performs learning by the learning unit 307 using the set target variable and explanatory variable, and generates, by machine learning, an inference model for inferring the state of the target (person) within the frames 400-(δ + 1) to 400-t by the shift feature amount group 702 (step S1012).

[0135] <Figure 16 Inference Device> Figure 16 is a flowchart showing an example of a state inference processing procedure by the client 102 (inference device) according to the third embodiment. After step S1103, the client 102 calculates, by the number measurement unit 319, the number of movement vectors 420 acquired from the vector information generation unit 312 as the number of people (step S1603).

[0136] After step S1111, when the number of people acquired by the client 102 from the number measurement unit 309 is more than a predetermined number, the client 102, by the inference unit 317, uses the two-dimensional histogram feature amount H2 t included in the shift feature amount 702-t as the input of the inference model, and when the number is less than or equal to the predetermined number, uses the one-dimensional histogram feature amount H1 t included in the shift feature amount 702-t as the input of the inference model to infer the state of the target (person) within the frames 400-(δ + 1) to 400-t by the shift feature amount group 702. (step S1612).

[0137] Thus, according to the third embodiment, according to the number of detected movement vectors 420 of people, the generated histogram feature amount H tBy switching, an inference model is generated, and by switching the inference model according to the number of detected human movement vectors 420, state inference can be realized with high precision according to the number of detected people.

Example

[0138] Example 4 will be described centering on the differences from Example 1, Example 2, and Example 3. Regarding the points common to Examples 1 to 3, the same reference numerals are given and the description thereof is omitted.

[0139] <Functional configuration example of analysis system 100 in FIG. 17> FIG. 17 is a block diagram showing a functional configuration example of an analysis system 100 according to Example 4. In Example 4, the server 101 (learning device) does not use a person detection unit 300, a tracking processing unit 301, and a vector information generation unit 302. Then, the time-series learning data output from the teacher signal DB 104 to the speed distribution calculation unit 303 and the direction distribution calculation unit 304 includes the human movement vector 420 at each time.

[0140] <Time-series learning data in FIG. 18> FIG. 18 is an explanatory diagram showing an example of time-series learning data output from the teacher signal DB 104. The time-series image data group 1800 is composed of t frames 1800-1 to 1800-t.

[0141] The time-series learning data group 1801 is a set of learning data 1801-1 to 1801-t output from the teacher signal DB 104. The learning data 1801-t is the set of velocity vectors at time t (n 1 t , n 2 t , …, n i t ), and the set of direction vectors at time t (d 1 t , d 2 t , …, d i t ) and, thereby, the set of movement vectors at time t (V 1 t , V2 t , …, V i t ) constitutes. In the time-series learning data group 1801, the learning data 1801-t at each time t is information derived from the frame 1800-t at the same time t in the time-series image data group 1800. For example, the set of motion vectors (V 1 t , V 2 t , …, V i t ) in the t-th learning data 1801-t is the motion vector V of the person (trajectory ID = 1 to i) in the t-th frame 1800-t of the time-series image data group 1800 1 t , V 2 t , …, V i t .

[0142] Thus, according to Example 4, by using the learning data 1801-t which is only numerical information instead of the image data (frame 1800-t), the learning data size can be suppressed and the learning time can be shortened.

Example

[0143] Example 5 will be described centering on the differences from Examples 1 to 4. Regarding the points common to Examples 1 to 4, the same reference numerals are given and the description thereof is omitted.

[0144] <Figure 19 Functional configuration example of the analysis system 100> Figure 19 is a block diagram showing a functional configuration example of the analysis system 100 according to Example 5. In Example 5, an event occurrence location output unit 321 is newly added to the terminal 106. The inference result receiving unit 320 indicates that it is in a state where additional control is required according to the inference results received from one or more clients 102, and outputs the inference results and the location information of the client 102 from which they are derived to the event occurrence location output unit 321.

[0145] Next, the event occurrence location output unit 321 outputs in a displayable manner the states that require additional control and the locations from which the inference results are derived, in accordance with the inference results of a plurality of clients 102 and the location information acquired from the inference result receiving unit 320.

[0146] As described above, according to the fifth embodiment, when comprehensively processing the inference results and location information from a plurality of clients 102, it is possible to grasp the locations where states requiring additional control occur and quickly perform additional control.

Embodiment

[0147] The sixth embodiment will be mainly described with reference to the differences from the first to fifth embodiments. Regarding the points common to the first to fifth embodiments, the same reference numerals are used and the descriptions thereof are omitted.

[0148] In the sixth embodiment, the inclusion feature amounts generated by the histogram feature amount generation unit 305 and the histogram feature amount generation unit 315 include the maximum speed n i t fastest and the maximum speed n i t fastest of the movement vector V i t fastest in the direction d i t fastest In addition to the histogram feature amount H t other information can also be included as feature amounts. Here, the other information is, for example, the position coordinates of the detected person, the speed n of the movement vector V i t the acceleration which is the derivative of the speed n i t the variance or standard deviation of the direction d of the movement vector V at each time t i t in the direction d i t and so on.

[0149] The learning unit 307 includes other information as explanatory variables in the shift feature amount 702-t, and learns an inference model with the corresponding known state, the occurrence location of the state, and the occurrence time as target variables. The inference unit 317 infers the state, the occurrence location of the state, and the occurrence time by inputting the shift feature amount 702-t including other information into the inference model.

[0150] Thus, according to Example 6, the acceleration of the movement vector V i t and information such as the position coordinates of the detected person are also input into the inference model, so that the state can be inferred, and the occurrence location and occurrence time of the state can be accurately inferred.

Example

[0151] Example 7 will be described centering on the differences from Examples 1 to 6. Regarding the points common to Examples 1 to 6, the same reference numerals are given and the description thereof is omitted.

[0152] <Figure 20 Shift Feature Amount Group> FIG. 20 is an explanatory diagram showing an example of a shift feature amount group according to Example 7. In Example 7, the shift feature amount generation unit 306 and the shift feature amount generation unit 316 combine the inclusion feature amount 701-t of each acquired frame with the inclusion feature amount 701-t of the current frame 400-t and the inclusion feature amounts 701-(t-δ 1 frames before, δ 2 frames before,..., δ m frames before) of the inclusion feature amounts 701-(t-δ 1 ) to 701-(t-δ m ) and output them as the shift feature amount 2003-t.

[0153] δ 1 ~δ m are integers of 1 or more that satisfy δ 1 <δ 2 <,..., <δ m and m is an integer of 1 or more.

[0154] Note that from the first to δ mFrames up to the -th frame 400-1 to 400-δ m For, δ m Since there is no previous frame, the generation process of the shift feature amount 2003-t is δ m+1 For frames after the -th frame 400-δ m+1 Is executed.

[0155] Thus, according to Example 7, state inference can be realized with high accuracy using the feature amounts of a plurality of past frames.

[0156] In addition, in the above-described Examples 1 to 7, the normal state and the abnormal state shown in FIG. 9 are assumed, but the combination of the normal state and the abnormal state is not limited to such an assumption. For example, a case may be assumed in which a person moves in a specific direction in the normal state, and in the abnormal state, the moving speed of the person is slower than the moving speed in the normal state and the moving direction becomes random. Specifically, for example, in the normal state where a crowd is moving in a certain direction toward the venue of an outdoor event, a case is assumed in which the flow rate of the crowd decreases due to traffic congestion or riot and the crowd becomes chaotic.

[0157] In such a case, the inclusion feature amount 701-t generated from the t-th frame 400-t is the minimum speed n i t slowest And the minimum speed n i t slowest Of the movement vector V i t slowest Of the direction d i t slowest And the histogram feature amount H t And may be composed of. Also, in the abnormal state, the maximum speed n i t fastest Or the minimum speed n i t slowest Not limited to, a specific speed n different from the normal state i t And the direction d at the specific speed i tAssuming this, the inclusion feature quantity 701-t may be generated.

Example

[0158] Example 8 will be described focusing on the differences from Examples 1 to 7. For points common to Examples 1 to 7, the same reference numerals are given and the description thereof is omitted. Example 8 is an example of learning and inferring a state indicating whether a person is looking at a specific direction (for example, a sign on the upper part of an opening / closing door of a vehicle). This inference model has three-dimensional position information of the sign to be visually observed.

[0159] <Functional configuration example of analysis system 100 in FIG. 21> FIG. 21 is a block diagram showing a functional configuration example of the analysis system 100 according to Example 8. In Example 8, the person detection unit 310, the vector information generation unit 302, the speed distribution calculation unit 313, the direction distribution calculation unit 314, and the histogram feature quantity generation unit 315 are not used in the client 102. On the other hand, the face detection unit 2101, the eye detection unit 2102, the face direction calculation unit 2103, and the gaze vector generation unit 2104 are newly added to the client 102.

[0160] The face detection unit 2101 generates face landmarks for the acquired frame and outputs the frame and the face landmarks to the tracking processing unit 311. A landmark is a feature point for extracting features of a face, such as the positions of eyes and nose. When outputting the frame and the face landmarks to the tracking processing unit 311, the face detection unit 2101 attaches current processing information indicating whether the currently executing process is a learning process or an inference process.

[0161] The tracking processing unit 311 generates a trajectory ID for the face landmarks acquired from the face detection unit 2101. The trajectory ID corresponds to the three-dimensional position of the face. If the face landmarks are from the same person between the acquired frames, the tracking processing unit 311 sets the trajectory ID to the same value. The tracking processing unit 311 outputs the trajectory ID and the face landmarks associated with the trajectory ID to the eye detection unit 2102.

[0162] Based on the facial landmarks obtained from the tracking processing unit 311, the eye detection unit 2102 identifies the positions of a person's eyes from the frame, and outputs the trajectory ID, the frame, and the image data of the person's eyes to the face direction calculation unit 2103. The image data of the eyes is used to identify whether the eyes are open or closed. That is, if the eyes are open, it is determined that the person is looking at something, and if the eyes are closed, it is determined that the person is not looking at anything. Also, the image data of the eyes is used to identify the position of the black eyes when the eyes are open. The position of the black eyes is used to generate the gaze vector.

[0163] Based on the frame obtained from the face detection unit 2101, the face direction calculation unit 2103 calculates the direction in which the face is facing, and outputs the direction in which the face is facing, the image data of the person's eyes, and the trajectory ID to the gaze vector generation unit 2104. The direction of the face is movement information indicating in which direction the face is moving, and is used to generate the gaze vector together with the position of the black eyes.

[0164] Using the direction in which the face is facing and the image data of the person's eyes obtained from the face direction calculation unit 2103, the gaze vector generation unit 2104 generates a gaze vector representing the direction of the gaze, and outputs the gaze vector, the direction in which the face is facing, the image data of the person's eyes, and the landmarks as the inclusion feature amount of each frame to the shift feature amount generation unit 316.

[0165] The shift feature amount generation unit 316 generates a shift feature amount by combining the inclusion feature amount of the current frame and the inclusion feature amount of the past frame for the gaze vector, the direction in which the face is facing, the image data of the person's eyes, and the landmarks obtained from the gaze vector generation unit 2104. If the current processing information is learning processing, the shift feature amount generation unit 316 outputs the calculated shift feature amount to the learning unit 307, and if the current processing information is inference processing, the shift feature amount generation unit 316 outputs the calculated shift feature amount to the inference unit 317.

[0166] The learning unit 307 uses the acquired shift feature quantity group 702, which is the shift feature quantity from time (δ + 1) to time t, as an explanatory variable, and the states of the learning data in frames 400 - (δ + 1) to 400 - t as objective variables to perform learning, and generates an inference model by machine learning to infer the state of the target (person) within frames 400 - (δ + 1) to 400 - t by the shift feature quantity group 702 (whether looking in a specific direction (for example, the direction with signage)). The learning unit 307 outputs the inference model generated by learning to the inference unit 317.

[0167] The inference unit 317 inputs the acquired shift feature quantity group 702, which is the shift feature quantity from time (δ + 1) to time t, into the inference model acquired from the learning unit 307, and infers the state of the target (person) within frames 400 - (δ + 1) to 400 - t by the shift feature quantity group 702 (whether looking in a specific direction (for example, the direction with signage)). The inference unit 317 outputs the inference result to the inference result output unit 318.

[0168] <Figure 22 Learning Process> Figure 22 is a flowchart showing a detailed processing procedure example of the learning process by the client 102 (learning device) according to Example 8.

[0169] (Step S2201) After step S1000, the client 102 detects a person's face from the face detection unit 2101.

[0170] (Step S2202) After step S2201, the client 102 detects landmarks for the detected person's face from the face detection unit 2101.

[0171] (Step S2203) After step S2202, the client 102 extracts the image data of the person's eyes using the acquired landmarks from the eye detection unit 2102.

[0172] (Step S2204) After step S2203, the client 102 calculates the direction in which the face is facing from the acquired eye image data by the face direction calculation unit 2103.

[0173] (Step S2205) After step S2204, the client 102 generates a gaze vector representing the direction of the gaze using the direction in which the face is facing and the human eye image data acquired by the gaze vector calculation unit 333.

[0174] (Step S2209) After step S2205, the client 102 accumulates, along the time series, the gaze vector, the direction in which the face is facing, the human eye image data, and the face landmarks in the frame as the inclusion feature amount of the frame by the shift feature amount generation unit 316.

[0175] <Figure 23 Inference Process> Figure 23 is a flowchart showing an example of a state inference process procedure by the client 102 (inference device) according to Example 8.

[0176] (Step S2301) After step S1100, the client 102 detects a human face from the face detection unit 2101.

[0177] (Step S2302) After step S2301, the client 102 detects landmarks for the detected human face from the face detection unit 2101.

[0178] (Step S2303) After step S2302, the client 102 extracts human eye image data using the acquired landmarks from the eye detection unit 2102.

[0179] (Step S2304) After step S2303, the client 102 calculates the direction in which the face is facing from the acquired image data by the face direction calculation unit 2103.

[0180] (Step S2305) After step S2304, the client 102 generates a gaze vector representing the direction of the gaze using the direction in which the acquired face is facing and the image data of the person's eyes from the gaze vector calculation unit 333.

[0181] (Step S2309) After step S2305, the client 102 accumulates the gaze vector, the direction in which the face is facing, the image data of the person's eyes, and the face landmarks in each frame as the inclusion feature amounts of each frame in time series from the shift feature amount generation unit 316.

[0182] Thus, according to the eighth embodiment, it is possible to infer whether a person who is a subject is looking at a signboard by using information such as a gaze vector.

[0183] Thus, according to the first to eighth embodiments described above, it is possible to improve the learning accuracy and inference accuracy of the inference model for inferring the state in space.

[0184] In addition, the analysis systems of the first to eighth embodiments described above can also be configured as follows (1) to (11).

[0185] (1) An analysis system 100 having a processor 201 that executes a program and a storage device 202 that stores the program, wherein the processor 201 repeatedly performs an acquisition process of repeatedly acquiring inclusion feature quantities (701-1 to 701-t) including information regarding the movement of the inference target (person) and information regarding the direction of the inference target; and for each of the inclusion feature quantities (701-1 to 701-t) repeatedly acquired by the acquisition process, from the inclusion feature quantity (701-(δ + 1)) at the first time (δ + 1) to the inclusion feature quantity (701-t) at the second time (t), by combining the inclusion feature quantity 701-(t - δ) at the third time δ time before, a generation process of generating a group of shift feature quantities (702) from the shift feature quantity (702-(δ + 1)) at the first time (δ + 1) to the shift feature quantity (702-t) at the second time (t); and an output process of outputting the group of shift feature quantities (702) generated by the generation process to an inference model that infers the state of the inference target during the period from the first time (δ + 1) to the second time (t).

[0186] Thereby, during learning, the inference model can be learned using the group of shift feature quantities (702) as explanatory variables and the group of states from the state at the first time (δ + 1) to the state at the second time (t) as objective variables. Therefore, it is possible to improve the learning accuracy of the inference model for inferring the state.

[0187] Also, during inference, by inputting the group of shift feature quantities (702) into the inference model, it is possible to infer the state during the period from the first time (δ + 1) to the second time (t). Therefore, it is possible to improve the inference accuracy of the inference model for inferring the state.

[0188] (2) In the analysis system 100 of (1) above, the information regarding the movement of the inference target includes the moving speed of the inference target, and the information regarding the direction of the inference target includes the moving direction in which the inference target moves.

[0189] This enables learning and inference in a state that takes into account the movement tendency of the object to be inferred.

[0190] (3) In the analysis system 100 of (2) above, the moving speed is a specific moving speed of a specific object to be inferred among the multiple moving speeds of multiple objects to be inferred, and the moving direction is a specific moving direction in which the specific object to be inferred moves at the specific moving speed.

[0191] This enables learning and inference in a state that takes into account the specific moving direction in which an object moves at a specific moving speed.

[0192] (4) In the analysis system 100 of (3) above, the specific moving speed is the highest moving speed among the multiple moving speeds, and the specific moving direction is the moving direction at the highest moving speed at which the specific object to be inferred moves.

[0193] This enables learning and inference in a state that takes into account the moving direction in which an object moves at the highest moving speed.

[0194] (5) In the analysis system 100 of (2) above, the information regarding the movement of the object to be inferred is a first one-dimensional histogram indicating the number of people by moving speed, and the information regarding the direction of the object to be inferred is a second one-dimensional histogram indicating the number of people by moving direction. In the acquisition process, the processor acquires the specific moving speed, the specific moving direction, the first one-dimensional histogram, and the second one-dimensional histogram as the inclusion feature amount.

[0195] This enables learning and inference in a state that takes into account the number distribution of moving speeds and the number distribution of moving directions.

[0196] (6) In the analysis system 100 of (2) above, the information regarding the movement of the object of inference is a first one-dimensional histogram indicating the number of people by moving speed, the information regarding the direction of the object of inference is a second one-dimensional histogram indicating the number of people by moving direction, and in the acquisition process, the processor acquires, as the inclusion feature amount, the specific moving speed, the specific moving direction, and a two-dimensional histogram obtained by integrating the first one-dimensional histogram and the second one-dimensional histogram.

[0197] This enables learning and inference in a state considering the number distribution of moving speeds and moving directions.

[0198] (7) In the analysis system 100 of (2) above, the information regarding the movement of the object of inference is a first one-dimensional histogram indicating the number of people by moving speed, the information regarding the direction of the object of inference is a second one-dimensional histogram indicating the number of people by moving direction, and the processor executes a determination process for determining whether to generate a two-dimensional histogram obtained by integrating the first one-dimensional histogram and the second one-dimensional histogram based on the number of movement vectors indicating that the object of inference moves at the moving speed in the moving direction. In the acquisition process, when it is determined by the determination process that the two-dimensional histogram is not generated, the processor 201 acquires, as the inclusion feature amount, the specific moving speed, the specific moving direction, the first one-dimensional histogram, and the second one-dimensional histogram. When it is determined by the determination process that the two-dimensional histogram is generated, the processor acquires, as the inclusion feature amount, the specific moving speed, the specific moving direction, and the two-dimensional histogram.

[0199] This can improve the learning accuracy and inference accuracy of the state.

[0200] (8) In the analysis system 100 of (1) above, in the generation process, the processor combines the inclusion feature amounts at a plurality of third times a plurality of predetermined times before for each of the inclusion feature amounts from the inclusion feature amount at the first time to the inclusion feature amount at the second time among the plurality of inclusion feature amounts, thereby generating a group of shift feature amounts from the shift feature amount at the first time to the shift feature amount at the second time.

[0201] Thereby, state inference can be realized with high precision using the feature amounts of a plurality of past frames.

[0202] (9) In the analysis system 100 of (1) above, the information regarding the movement of the inference target includes the information regarding the movement of the face of the inference target, and the information regarding the direction of the inference target includes the line-of-sight direction at which the inference target is looking.

[0203] Thereby, learning and inference of a state considering the visual tendency as to whether the inference target is looking at the visual target become possible.

[0204] (10) In the analysis system 100 of (1) above, the processor 201 executes a learning process of learning the inference model, using the group of shift feature amounts (702) as explanatory variables and the group of states from the state at the first time (δ + 1) to the state at the second time (t) as objective variables.

[0205] Thereby, improvement in the learning accuracy of the inference model for inferring the state can be achieved.

[0206] (11) In the analysis system 100 of (1) above, the processor 201 executes an inference process of inferring the state in the period from the first time (δ + 1) to the second time (t) by inputting the group of shift feature amounts (702) into the inference model.

[0207] Thereby, improvement in the inference accuracy of the inference model for inferring the state can be achieved.

[0208] Note that the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Further, the configuration of another embodiment may be added to the configuration of one embodiment. Also, for a part of the configuration of each embodiment, addition, deletion, or replacement with other configurations may be made.

[0209] In addition, each of the above-described configurations, functions, processing units, processing means, etc. may be realized in hardware by designing a part or all of them, for example, by using an integrated circuit, or may be realized in software by the processor 201 interpreting and executing a program for realizing each function.

[0210] Information such as programs, tables, and files for realizing each function can be stored in a storage device such as a memory, a hard disk, an SSD (Solid State Drive), or a recording medium such as an IC (Integrated Circuit) card, an SD card, or a DVD (Digital Versatile Disc).

[0211] Also, the control lines and information lines show those considered necessary for explanation, and do not necessarily show all the control lines and information lines necessary for implementation. In reality, it may be considered that almost all the configurations are interconnected.

Explanation of Reference Numerals

[0212] 100 Analysis system 101 Server 102 Client 103 Sensor 104 Teacher signal DB 106 Terminal 201 Processor 202 Memory device 203 Input device 204 Output device 205 Communication IF 300, 310 Person detection unit 301, 311 Tracking processing unit 302, 312 Vector information generation unit 303, 313 Velocity distribution calculation unit 304, 314 Direction distribution calculation unit 305, 315 Histogram feature quantity generation unit 306, 316 Shift feature quantity generation unit 307 Learning unit 317 Inference unit 318 Inference result output unit 309, 319 Number of people measurement unit 320 Inference result reception unit 321 Incident location output unit 330 Face detection unit 331 Eye detection unit 332 Face direction calculation unit 333 Gaze vector generation unit

Claims

1. An analysis system having a processor that executes a program and a storage device that stores the program, wherein the processor performs an acquisition process of repeatedly acquiring an inclusion feature amount including information on the movement of the object to be inferred and information on the direction of the object to be inferred; a generation process of generating a group of shift feature amounts from the shift feature amount at the first time to the shift feature amount at the second time by combining the inclusion feature amounts at the third time before a predetermined time for each of the plurality of inclusion feature amounts repeatedly acquired by the acquisition process; an output process of outputting the group of shift feature amounts generated by the generation process to an inference model that infers the state of the object to be inferred during the period from the first time to the second time; An analysis system characterized by performing the above.

2. The analysis system according to claim 1, wherein the information on the movement of the object to be inferred includes the moving speed of the object to be inferred, the information on the direction of the object to be inferred includes the moving direction in which the object to be inferred moves, An analysis system characterized by the above.

3. The analysis system according to claim 2, wherein the moving speed is the specific moving speed of a specific object to be inferred among the plurality of moving speeds of the plurality of objects to be inferred, the moving direction is the specific moving direction in which the specific object to be inferred moves at the specific moving speed, An analysis system characterized by the above.

4. The analysis system according to claim 3, wherein the specific moving speed is the highest moving speed among the plurality of moving speeds, the specific moving direction is the moving direction at the highest moving speed at which the specific object to be inferred moves, An analysis system characterized by the above.

5. The analysis system according to claim 2, wherein the information on the movement of the object to be inferred is a first one-dimensional histogram indicating the number of people for each moving speed, the information on the direction of the object to be inferred is a second one-dimensional histogram indicating the number of people for each moving direction, in the acquisition process, the processor acquires the specific moving speed, the specific moving direction, the first one-dimensional histogram, and the second one-dimensional histogram as the inclusion feature amount, An analysis system characterized by the above.

6. The analysis system according to claim 2, wherein the information on the movement of the object to be inferred is a first one-dimensional histogram indicating the number of people for each moving speed, The information regarding the direction of the object to be inferred is a second one-dimensional histogram indicating the number of people for each moving direction, In the acquisition process, the processor acquires, as the inclusion feature amount, the specific moving speed, the specific moving direction, and a two-dimensional histogram obtained by integrating the first one-dimensional histogram and the second one-dimensional histogram. An analysis system characterized by the above.

7. The analysis system according to claim 2, The information regarding the movement of the object to be inferred is a first one-dimensional histogram indicating the number of people for each moving speed, The information regarding the direction of the object to be inferred is a second one-dimensional histogram indicating the number of people for each moving direction, The processor, executes a determination process for determining whether to generate a two-dimensional histogram obtained by integrating the first one-dimensional histogram and the second one-dimensional histogram based on the number of movement vectors indicating that the object to be inferred moves at the moving speed in the moving direction, In the acquisition process, when it is determined by the determination process that the two-dimensional histogram is not generated, the processor acquires, as the inclusion feature amount, the specific moving speed, the specific moving direction, the first one-dimensional histogram, and the second one-dimensional histogram, and when it is determined by the determination process that the two-dimensional histogram is generated, the processor acquires, as the inclusion feature amount, the specific moving speed, the specific moving direction, and the two-dimensional histogram. An analysis system characterized by the above.

8. The analysis system according to claim 1, In the generation process, the processor generates a group of shift feature amounts from the shift feature amount at the first time to the shift feature amount at the second time by combining the inclusion feature amounts at a plurality of third times before a plurality of predetermined times for each of the inclusion feature amounts from the inclusion feature amount at the first time to the inclusion feature amount at the second time among the plurality of inclusion feature amounts. An analysis system characterized by the above.

9. The analysis system according to claim 1, The information regarding the movement of the object to be inferred includes information regarding the movement of the face of the object to be inferred, The information regarding the direction of the object to be inferred includes the line-of-sight direction at which the object to be inferred looks, An analysis system characterized by the above.

10. The analysis system according to claim 1, The processor, A learning process for training the inference model, using the shift feature quantity group as an explanatory variable and the state group from the state at the first time to the state at the second time as an objective variable; An analysis system characterized by executing the above.

11. The processor is caused to: A acquisition process for repeatedly acquiring an inclusion feature quantity including information on the movement of the inference target and information on the direction of the inference target; A generation process for generating a shift feature quantity group from the shift feature quantity at the first time to the shift feature quantity at the second time by combining the inclusion feature quantity at the third time before a predetermined time for each of the inclusion feature quantities from the inclusion feature quantity at the first time to the inclusion feature quantity at the second time among the plurality of inclusion feature quantities repeatedly acquired by the acquisition process; An output process for outputting the shift feature quantity group generated by the generation process to an inference model for inferring the state of the inference target in the period from the first time to the second time; An analysis program characterized by causing the above to be executed.

Citation Information

Patent Citations

  • Suspicious behavior detection system

    JP2020149389A

Cited By

  • Analysis system and analysis program

    WO2025115364A1