Action recognition device, learning device, and action recognition method
The action recognition device enhances recognition efficiency by interpolating missing shape parts and performing component analysis to accurately recognize actions using an action classification model, addressing inefficiencies in existing technologies.
Patent Information
- Application Number
- JP2022095486
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-14
AI Technical Summary
Existing action recognition technologies do not fully utilize a single action classification model for recognizing multiple types of actions, especially when parts of the shape are missing, leading to inefficiencies in memory capacity and recognition speed.
An action recognition device that employs a processor and storage device to detect missing parts, interpolate missing information, and perform component analysis to generate component groups for accurate action recognition using an action classification model.
This approach reduces memory capacity and speeds up the recognition of multiple types of actions with high accuracy even when parts of the shape are missing.
Smart Images

Figure 0007713426000015 
Figure 0007713426000016 
Figure 0007713426000017
Abstract
Description
Technical Field
[0001] The present invention relates to an action recognition device, a learning device, and an action recognition method.
Background Art
[0002] As background art in this technical field, Patent Document 1 discloses an action recognition device that accurately recognizes a plurality of types of actions of a recognition target. This action recognition device can access a group of action classification models learned for each component group using a group of components obtained from the shape of a learning target by component analysis that generates statistical components by multivariate analysis, and the action of the learning target. It detects the shape of the recognition target from the analysis target data, and by component analysis, based on the shape of the recognition target, generates one or more components and the contribution rate of each component, and based on the cumulative contribution rate obtained from each contribution rate, determines an ordinal number indicating the dimension of each of the one or more components, selects a specific action classification model learned in the same component group as a specific component group including one or more components indicated by the determined ordinal number from the group of action classification models, and inputs the specific component group to the specific action classification model, thereby outputting a recognition result indicating the action of the recognition target.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, the above-mentioned Patent Document 1 is a technique for selecting a specific action classification model from a group of action classification models for action recognition, and does not consider the point of using one action classification model to the fullest for action recognition.
[0005] The object of the present invention is to reduce the memory capacity and speed up the action recognition when accurately recognizing a plurality of types of actions of a recognition target with a part of its shape missing.
Means for Solving the Problem
[0006] An action recognition device which is an aspect of the invention disclosed in the present application is an action recognition device having a processor that executes a program and a storage device that stores the program, and is accessible to an action classification model learned using a component group related to the learning target obtained from the shape of the learning target by component analysis that generates statistical components by multivariate analysis and the action of the learning target. The processor includes a detection process for detecting the shape of the recognition target from the analysis target data, a missing position information generation process for generating missing position information indicating the position of the missing part among the shapes of the recognition target detected by the detection process, an interpolation process for interpolating the missing part from non-missing information which is a part other than the missing part among the shapes of the recognition target including the missing part and updating the non-missing information after interpolation as the shape of the recognition target, a component analysis process for generating, by the component analysis, the same number of component groups related to the recognition target as the component groups related to the learning target based on the shape of the recognition target interpolated by the interpolation process, and an action recognition process for outputting a recognition result indicating the action of the recognition target by inputting the component group related to the recognition target generated by the component analysis process and the missing position information to the action classification model.
[0007] A learning device according to one aspect of the invention disclosed in the present application is a learning device having a processor that executes a program and a storage device that stores the program. The processor performs an acquisition process of acquiring teacher data including the shape and behavior of a learning target, a defect process of deleting the shape of the learning target acquired by the acquisition process, a defect position information generation process of generating defect position information indicating the position of a defect portion deleted from the shape of the learning target by the defect process, an interpolation process of interpolating from non-defect information which is a portion other than the defect portion deleted by the defect process in the shape of the learning target and updating the interpolated non-defect information as the shape of the learning target, a component analysis process of generating a component group related to the learning target based on the shape of the learning target interpolated by the interpolation process by component analysis that generates statistical components by multivariate analysis, and an action learning process of learning the behavior of the learning target and generating an action classification model for classifying the behavior of the learning target based on the component group related to the learning target generated by the component analysis process, the behavior of the learning target, and the defect position information.
Advantages of the Invention
[0008] According to a typical embodiment of the present invention, when recognizing a plurality of types of behaviors of a recognition target with a part of its shape missing with high accuracy, it is possible to reduce the storage capacity and speed up the behavior recognition. Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
MODE FOR CARRYING OUT THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings for explaining the embodiments, the same members are generally denoted by the same reference numerals, and repeated explanations thereof are omitted. In the following embodiments, it goes without saying that the constituent elements (including element steps, etc.) are not necessarily essential except in cases where it is particularly specified or considered to be clearly essential in principle. Also, when it is said that "comprising A", "consisting of A", "having A", or "including A", it goes without saying that other elements are not excluded except in cases where it is particularly specified that only those elements are present. Similarly, in the following embodiments, when referring to the shape, positional relationship, etc. of constituent elements, etc., it is assumed to include those that are substantially approximated or similar to the shape, etc., except in cases where it is particularly specified or considered to be clearly otherwise in principle.
[0011] Expressions such as "first", "second", "third", etc. in this specification and the like are attached for identifying constituent elements, and do not necessarily limit the number, order, or content thereof. Also, the numbers for identifying constituent elements are used for each context, and the numbers used in one context do not necessarily indicate the same configuration in other contexts. Further, it does not prevent a constituent element identified by a certain number from also having the functions of a constituent element identified by another number.
[0012] The positions, sizes, shapes, ranges, etc. of each configuration shown in the drawings and the like may not represent the actual positions, sizes, shapes, ranges, etc. in order to facilitate understanding of the invention. For this reason, the present invention is not necessarily limited to the positions, sizes, shapes, ranges, etc. disclosed in the drawings and the like.
Example
[0013] <Action Recognition System> FIG. 1 is an explanatory diagram showing a system configuration example of the action recognition system according to Example 1. The action recognition system 100 includes a server 101 and one or more clients 102. The server and the clients are communicably connected via a network 105 such as the Internet, a LAN (Local Area Network), or a WAN (Wide Area Network). The server 101 is a computer that manages the clients 102. The client 102 is a computer connected to a sensor 103 and acquires data from the sensor 103.
[0014] The sensor 103 detects analysis target data from the analysis environment. The sensor 103 is, for example, a camera that captures a still image or a moving image. Also, the sensor 103 may detect sound or odor. The teacher signal DB 104 is a database that holds combinations of learning data (human skeleton information) and action information (for example, human postures and movements such as "standing" and "falling") as teacher signals. The teacher signal DB 104 may be stored in the server 101, or may be connected to a computer communicable with the server 101 or the client 102 via the network 105.
[0015] The action recognition system 100 has a learning function using the teacher signal DB 104 and an action recognition function using the action classification model obtained by the learning function. The action classification model is a learning model for classifying the actions of recognition targets such as humans and animals. The learning function and the action recognition function may be implemented in either the server 101 or the client 102 as long as they are implemented in the action recognition system 100. For example, the server 101 may implement the learning function and the client 102 may implement the action recognition function. Also, the server 101 may implement the learning function and the action recognition function, and the client 102 may transmit data from the sensor 103 to the server 101 or receive the action recognition result by the action recognition function from the server 101.
[0016] In addition, client 102 may implement a learning function and a behavior recognition function, and server 101 may manage the behavior classification model and behavior recognition results from client 102. Note that a computer implementing the learning function is referred to as a learning device, and a computer implementing at least the behavior recognition function among the learning function and the behavior recognition function is referred to as a behavior recognition device. Also, in FIG. 1, a client-server type behavior recognition system 100 is taken as an example, but a stand-alone type behavior recognition device may also be used. In the first embodiment, for convenience of explanation, a behavior recognition system 100 in which server 101 implements a learning function (learning device) and client 102 implements a behavior recognition function (behavior recognition device) will be described as an example.
[0017] <Example of Computer Hardware Configuration> FIG. 2 is a block diagram showing an example of the hardware configuration of a computer (server 101, client 102). The computer 200 includes a processor 201, a storage device 202, an input device 203, an output device 204, and a communication interface (communication IF) 205. The processor 201, the storage device 202, the input device 203, the output device 204, and the communication IF 205 are connected by a bus 206. The processor 201 controls the computer 200. The storage device 202 serves as a working area for the processor 201. Also, the storage device 202 is a non-temporary or temporary recording medium that stores various programs and data. Examples of the storage device 202 include a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), and a flash memory. The input device 203 inputs data. Examples of the input device 203 include a keyboard, a mouse, a touch panel, a numeric keypad, and a scanner. The output device 204 outputs data. Examples of the output device 204 include a display, a printer, and a speaker. The communication IF 205 is connected to the network 105 and transmits and receives data.
[0018] <Learning Data> FIG. 3 is an explanatory diagram showing an example of learning data. The learning data 380 is composed of skeletal information 320 and joint angles 370 for each subject. The skeletal information 320 is detected based on the analysis target data acquired from the sensor 103. The joint angles 370 are calculated based on the skeletal information 320. The learning data 380 for one subject is composed of, for example, a combination of the skeletal information 320 and the joint angles 370 obtained from each of a plurality of time-series frames in which the subject is the object.
[0019] The skeletal information 320 has, for each of a plurality (18 in this example) of skeletal points 300 to 317, a name 321, an x coordinate value 322 on the x-axis, and a y coordinate value 323 on the y-axis orthogonal to the x-axis. The joint angles 370 also have, for each of a plurality (18 in this example) of skeletal points 300 to 317, a name 371. In the name 371, ∠a-b-c (a, b, c are the names 321 of the skeletal points) is the joint angle 370 of the skeletal point b formed by the line segment ab and the line segment bc. Note that the skeletal information 320 may include, for example, finger joints. Also, the joint angles 370 may include joint angles 370 other than these.
[0020] In FIG. 3, the coordinate values of the skeletal points 300 to 317 are two-dimensional position information (a combination of the x coordinate value and the y coordinate value), but may also be three-dimensional position information. Specifically, for example, a z coordinate value on the z-axis (for example, the depth direction) orthogonal to the x-axis and the y-axis may be added.
[0021] <Functional Configuration Example of Action Recognition System 100> FIG. 4 is a block diagram showing a functional configuration example of the action recognition system 100 according to the first embodiment. The server 101 includes a teacher signal acquisition unit 401, a defect occurrence unit 402, a defect control unit 421, a defect position information generation unit 422, a defect information interpolation unit 423, a skeletal information processing unit 403, a principal component analysis unit 404, and an action learning unit 406. The client 102 includes a skeletal detection unit 451, a defect position information generation unit 462, a defect information interpolation unit 461, a skeletal information processing unit 453, a principal component analysis unit 454, and an action recognition unit 457.
[0022] Specifically, these are realized by causing the processor 201 to execute a program stored in the memory device 202 shown in FIG. 2, for example. First, a functional configuration example on the server 101 side will be described.
[0023] The teacher signal acquisition unit 401 acquires one or more teacher signals for learning from the teacher signals acquired from the teacher signal DB 104, and outputs the selected teacher signals to the defect generation unit 402. The skeletal information 320 in the teacher signal is non-defective information without any missing skeletal points.
[0024] The defect control unit 421 sets the skeletal points to be defective according to a random number or a predetermined number, and outputs the defect generation information to the defect generation unit 402 and the defect position information generation unit 422. The number of skeletal points to be defective may be one, plural, or 0 indicating that no defect is to be made. Here, the predetermined number is, for example, a number indicating a skeletal position where there is a high possibility of detection omission when the skeletal detection unit 451 described later detects the skeletal information 320 of a person reflected in the analysis target data acquired from the sensor 103.
[0025] The defect generation unit 402 defects the skeletal points in the skeletal information 320 in the teacher signal acquired from the teacher signal acquisition unit 401 according to the defect generation information. The defect generation unit 402 updates the skeletal information 320 after the defect (including the case where the number of skeletal points to be defective is 0) as the skeletal information 320 in the teacher signal. Also, in order to enhance noise tolerance, noise may be added to shift the positions of the skeletal points that are not to be defective with respect to the skeletal information 320 when defecting the skeletal points, and the skeletal information 320 may be updated. The single or plural defective skeletal information 320 is denoted as "defect information".
[0026] Then, the defect generation unit 402 outputs a teacher signal including defect information, which is the name 321 of the defective skeletal points and the position information (x coordinate value 322, y coordinate value 323), to the defect information interpolation unit 423.
[0027] The missing information interpolation unit 423 interpolates the missing information from the skeletal information 320 that is not missing. The single or plural skeletal information 320 that is not missing is denoted as non-missing information.
[0028] Specifically, for example, the missing information interpolation unit 423 may interpolate the missing information from among the non-missing information, from the skeletal points that were connected to the skeletal points of the missing information or from the skeletal points located near the skeletal points of the missing information.
[0029] Also, the missing information interpolation unit 423 may substitute predetermined position information for the missing information. Further, the missing information interpolation unit 423 may interpolate using the missing information of the skeletal information 320 determined to include the missing information among the skeletal information 320 of another frame acquired thus far. In this way, the interpolation method of the missing information is not limited. The missing information interpolation unit 423 outputs the interpolated skeletal information 320 to the skeletal information processing unit 403.
[0030] The missing position information generation unit 422 outputs, as missing position information, the position information of the missing skeletal points from the missing occurrence information from the missing control unit 421 to the kinematic learning unit 406.
[0031] The skeletal information processing unit 403 processes the skeletal information 320 after interpolation by the missing information interpolation unit 423. Specifically, for example, the skeletal information processing unit 403 calculates the joint angle 370 and the movement amount between frames from among the acquired interpolated teacher signals and the skeletal information 320. Further, the skeletal information processing unit 403 excludes the absolute position information from the skeletal information 320 and executes normalization such that the size of the skeletal information 320 becomes constant. Then, the skeletal information processing unit 403 outputs the joint angle 370, the movement amount between frames, and the normalized skeletal information 320 to the principal component analysis unit 404.
[0032] FIG. 5 is a block diagram showing a detailed functional configuration example of the skeletal information processing units 403 and 453. The skeletal information processing units 403 and 453 include a joint angle calculation unit 501, a movement amount calculation unit 502, and a normalization unit 503.
[0033] The joint angle calculation unit 501 calculates the joint angle 370 from the interpolated skeleton information 320 among the acquired teacher signals, and outputs it to the principal component analysis unit 404 via the movement amount calculation unit 502 and the normalization unit 503.
[0034] The movement amount calculation unit 502 calculates the movement amount between frames from the interpolated skeleton information 320 among the acquired teacher signals, and outputs it to the principal component analysis unit 404 via the normalization unit 503.
[0035] The normalization unit 503 excludes the absolute position information from the interpolated skeleton information 320 among the acquired teacher signals, performs normalization so that the magnitude of the interpolated skeleton information 320 becomes constant, and outputs it to the principal component analysis unit 404.
[0036] Returning to FIG. 4, the principal component analysis unit 404 uses the normalized skeleton information 320, the joint angle 370, and the movement amount between frames among the teacher signals acquired from the skeleton information processing unit 403 as input data, performs principal component analysis to generate one or more principal components, and outputs them to the behavior learning unit 406. Note that among the skeleton information 320, the joint angle 370, and the movement amount between frames, at least the normalized skeleton information 320 may be the input data.
[0037] In principal component analysis, as shown in the following formula (1), the input data x i is multiplied by the coefficient w ij respectively, and the principal component y i is generated by adding them. The general formula of principal component analysis is shown in the following formula (2). The coefficient w ij is defined as shown in the following formula (3), and is determined so that the variance V(y i ) is maximized. i ) When defined as V(y i ), it is determined so that the variance V(y
[0038] However, if no constraint is imposed on the coefficient wij, the absolute value of the variance V(y i ) can be infinitely large, and the coefficient w ij cannot be uniquely determined. Therefore, it is desirable to impose the constraint of the following formula (4). Also, to eliminate information duplication, the newly generated principal component yk and the covariance with the principal component y generated so far k is desirably subject to the constraint of the following formula (5) that becomes 0.
[0039]
Equation
[0040] However, the above formula (4) and the above formula (5) attached as constraints are not limited to these, and other constraint conditions may be attached, or the constraints may be removed, and there is no problem in calculating the coefficient w ij . The variance V(yj) of the newly generated principal component y j is as shown in the following formula (6) when defined separately as λ j , and the sum of the variances V(x j ) of the input data x j and the sum of λ j are equal as shown in the following formula (7).
[0041]
Equation
[0042] Here, p is the number of the input data x j . The variance V(y j ) of the newly generated principal component y j is higher and reflects more of the original information. The principal components are called the first, second,..., m-th principal components in order from the principal component with the highest variance value. The ratio of the variance of the newly generated variable y j to the variance of the original data is called the contribution rate and is shown by the following formula (8). Also, the result of adding the contribution rates in descending order of the variance value (ascending order of the principal component ordinal m) from the contribution rate of the first principal component is called the cumulative contribution rate and is shown by the following formula (9).
[0043]
Equation
[0044] The contribution rate and the cumulative contribution rate are for the newly generated principal component y jIt becomes a measure such as how much the generated multiple principal components represent the information amount of the original data, and is generated together with the principal components. As an example of component analysis that generates statistical components in multivariate analysis, principal component analysis was applied. However, instead of principal component analysis, independent component analysis, which is also an example of component analysis, may be executed.
[0045] In the case of independent component analysis, the principal components become independent components. As an index indicating how much this independent component affects the input data xi, the contribution rate may be used. In independent component analysis, the sum of squares of the mixing coefficient matrix in the independent component analysis for each independent component becomes the strength of each independent component.
[0046] The strength of the independent component is the variance in the input data x of the independent component. i That is, since the variances of all the independent components obtained by independent component analysis are unified to 1, taking the sum of squares of the mixing coefficients results in the variance of the input data x. i And the value obtained by dividing the strength of the independent component by the sum of the strengths of all the independent components may be defined as the contribution rate of that independent variable.
[0047] The principal component analysis unit 404 controls the ordinal number k indicating the dimension of each of one or more components. Specifically, for example, the principal component analysis unit 404 determines up to which dimension to use the principal components used for learning in the behavior learning unit 406 among the generated principal components in descending order of the variance value, and outputs the principal components from the first principal component to the k-th principal component with the determined dimension k (k is an integer of 1 or more) as the ordinal number to the behavior learning unit 406 in descending order of the variance value. Specifically, for example, the principal component analysis unit 404 adopts the dimension k of the principal component that is equal to or greater than the threshold of the cumulative contribution rate and is closest to the threshold. Alternatively, the principal component analysis unit 404 adopts the dimension k of the principal component that is equal to or less than the threshold of the cumulative contribution rate and is closest to the threshold.
[0048] The behavioral learning unit 406 associates and learns the principal components obtained from the principal component analysis unit 404, the missing position information obtained from the missing position information generation unit 422, and the behavioral information in the teacher signal obtained from the teacher signal DB 104. Specifically, for example, the behavioral learning unit 406 uses, as an explanatory variable, data obtained by concatenating the principal component group from the first principal component to the k-th principal component obtained from the principal component analysis unit 404 and the missing position information, and uses, as an objective variable, the behavioral information in the teacher signal obtained from the teacher signal DB 104 to generate a behavior classification model by machine learning. The behavioral learning unit 406 outputs the behavior classification model generated as a result of the learning to the behavior recognition unit 457.
[0049] Next, a functional configuration example on the client 102 side will be described. The skeleton detection unit 451 detects the human skeleton information 320 reflected in the analysis target data obtained from the sensor 103 and outputs it to the missing information interpolation unit 461 and the missing position information generation unit 462. For the detection of the skeleton information 320, an NN (neural network) capable of estimating the human skeleton information 320 generated by machine learning may be used, or markers may be attached to the skeleton points of the person to be detected, and the skeleton information 320 may be detected from the marker positions reflected in the image. The method for detecting the skeleton information 320 is not limited. Among the skeleton information 320, the skeleton points whose x-coordinate value 322 and y-coordinate value 323 are not detected and are indefinite values are the missing information, and the skeleton points excluding the missing information among the skeleton information 320 are the non-missing information. For example, if only the x-coordinate value 322 and y-coordinate value 323 of the skeleton point 310 among the skeleton points 300 to 317 are indefinite values, the skeleton point 310 is the missing information, and the skeleton points 300 to 309 and 311 to 317 are the non-missing information.
[0050] The missing information interpolation unit 461 has the same function as the missing information interpolation unit 423. The missing information interpolation unit 461 performs the same processing as the missing information interpolation unit 423 on the skeleton information 320 from the skeleton detection unit 451 to interpolate the missing information (skeleton point 310 in the above example) from the non-missing information (skeleton points 300 to 309 and 311 to 317 in the above example), and outputs the interpolated skeleton information 320 to the skeleton information processing unit 453.
[0051] The missing position information generation unit 462 has the same function as the missing position information generation unit 422. The missing position information generation unit 462 executes the same processing as the missing position information generation unit 422 on the skeleton information 320 from the skeleton detection unit 451, determines whether there are skeleton points that cannot be obtained due to occlusion or the like among the skeleton information 320 detected by the skeleton detection unit 451. If there are skeleton points that cannot be obtained, it generates missing position information indicating the positions of those skeleton points as missing portions, and outputs it to the action recognition unit 457.
[0052] The skeleton information processing unit 453 has the same function as the skeleton information processing unit 403. The skeleton information processing unit 453 executes the same processing as the skeleton information processing unit 403 on the skeleton information 320 detected by the skeleton detection unit 451, and outputs the joint angle 370, the movement amount between frames, and the normalized skeleton information 320 to the principal component analysis unit 454.
[0053] The principal component analysis unit 454 has the same function as the principal component analysis unit 404. The principal component analysis unit 454 executes the same processing as the principal component analysis unit 404 on the output data from the skeleton information processing unit 453, generates one or more principal components, and outputs them to the action recognition unit 457.
[0054] The principal component analysis unit 454 determines the ordinal number k indicating the dimension of each of the components of 1 or more based on the cumulative contribution rate obtained from each contribution rate. Specifically, for example, the principal component analysis unit 454 determines the number of dimensions k indicating up to which dimension of the obtained principal components in descending order of variance should be output to the action recognition unit 457 from the obtained contribution rate and cumulative contribution rate. The number of dimensions k is the ordinal number k indicating the dimension of the principal component. For example, for the first principal component, the number of dimensions (ordinal number) k = 1, and for the second principal component, the number of dimensions (ordinal number) k = 2. The principal component analysis unit 454 outputs a group of principal components from the first principal component to the k-th principal component in descending order of variance to the action recognition unit 457.
[0055] The action recognition unit 457 recognizes the action of a person reflected in the analysis target data acquired from the sensor 103 based on the action classification model generated by the action learning unit 406, the principal component group from the first principal component to the k-th principal component, and the missing position information. Specifically, for example, the action recognition unit 457 inputs the principal component group (the first principal component to the k-th principal component) and the missing position information obtained from the analysis target data into the selected action classification model, and outputs a predicted value indicating the action of the person reflected in the analysis target data as the recognition result.
[0056] <Example of joint angle calculation> FIG. 6 is an explanatory diagram showing a detailed calculation method of the joint angle 370 executed by the joint angle calculation unit 501. The joint angle calculation unit 501 calculates the joint angle θ at the three skeleton points 600 to 602 to be connected. Regarding the skeleton information 620 of the skeleton points 600 to 602, they are respectively defined as position vectors O, A, and B with respect to the origin 630. The joint angle calculation unit 501 calculates the relative vector with the skeleton point 600 as the origin as shown in the following formulas (10) and (11), and since the following formula (12) holds for the calculated vectors, the joint angle θ is calculated by calculating the inverse cosine as shown in the following formula (13).
[0057]
Equation
[0058] <Example of calculation of movement amount between frames> FIG. 7 is an explanatory diagram showing an example of a detailed calculation method of the movement amount between frames executed by the movement amount calculation unit 502. In calculating the movement amount between frames, the movement amount calculation unit 502 uses the skeleton information 701 of the N-th frame and the skeleton information 702 of the (N - M)-th frame for the same subject. N and M are integers of 1 or more, and N > M. The value of M can be arbitrarily set. As shown in the following formulas (14) to (16), the movement amount calculation unit 502 calculates the distances of the same skeleton points 300 to 317 of the same person shown between each frame. The movement amount between frames of the 18 skeleton points 300 to 317 becomes the movement amount between frames for the person.
[0059] [Number]
[0060] However, the amount of movement between frames executed by the movement amount calculation unit 502 is not limited to this. As shown in the following formula (17), the movement amount calculation unit 502 calculates the distances between the same skeleton points 300 to 317 of the same person shown between each frame, and the value obtained by summing up the amounts of movement between frames of all 18 skeleton points 300 to 317 may be used as the amount of movement between frames for the person. [Number]
[0061] In addition, the movement amount calculation unit 502 may use the center-of-gravity skeleton information 711 and the center-of-gravity skeleton information 712 that are the centers of gravity among the skeleton information 701 of the nth frame and the skeleton information 702 of the (n - m)th frame. Specifically, for example, as shown in the following formulas (18) to (19), the movement amount calculation unit 502 calculates the center of gravity for each person, and as shown in the following formula (20), calculates the amount of movement between frames for the person with respect to the calculated center of gravity.
[0062] [Number]
[0063] [Example of Normalization] FIG. 8 is an explanatory diagram showing a detailed method for normalizing the skeleton information 320 executed by the normalization unit 503. First, the normalization unit 503 calculates the center of gravity from (a) all or part of the skeleton information 320, and (b) converts it to relative coordinates with the center of gravity as the origin. After that, the normalization unit 503 divides (d) the position information of each skeleton point of the skeleton information 320 by the length L of the diagonal of the smallest rectangle surrounding the 18 skeleton points 300 to 317. When the skeleton information 320 obtained in (d) is used as the teacher signal, the position information of the skeleton points 300 to 317 after division is also incorporated.
[0064] For example, if learning for skeleton detection and action classification is performed for an action such as "a person with a height of 180 cm sits at point A" without the normalization unit 503 being executed, determinations such as "not sitting outside point A" and "a person other than 180 cm tall does not sit" may be made. In order to exclude such limitations and give generality to action classification, the normalization unit 503 normalizes the skeleton information 320 in order to remove the absolute position information in the image and the information regarding the size of the skeleton.
[0065] <Teacher signal held by the teacher signal DB104> FIG. 9 is an explanatory diagram showing a detailed example of the teacher signal held by the teacher signal DB104. In the person shown in the (a) image 900 that is the data to be analyzed, the combination of (b) the skeleton information 320A, the joint angles 370 (not shown), and (c) the action information 901 ("stand up") associated with the skeleton information 320A becomes the teacher signal. Similarly, in the person shown in the (a) image 910 that is the data to be analyzed, the combination of (b) the skeleton information 320B, the joint angles 370 (not shown), and (c) the action information 911 ("fall down") associated with the skeleton information 320B becomes the teacher signal.
[0066] <Dimensionality reduction by the principal component analysis unit 404> FIG. 10 is an explanatory diagram showing an example in which the principal components generated by the principal component analysis unit 404 with the teacher signal as input data are plotted on the principal component space. The legend shows the action information 1000 to 1004 included in the teacher signal.
[0067] In FIG. 10, (a) shows an example in which the first principal component is taken on the X-axis and the second principal component is taken on the Y-axis, and the information up to the second principal component is plotted on a two-dimensional plane. (b) shows an example in which the first principal component is taken on the X-axis, the second principal component is taken on the Y-axis, the third principal component is taken on the Z-axis, and the information up to the third principal component is plotted in a three-dimensional space.
[0068] In (a), it can be seen that standing 1000, sitting 1001, and falling 1004 can be separated even on the two-dimensional plane up to the second principal component, while walking 1002 and crouching 1003 are difficult to separate on the two-dimensional plane up to the second principal component. Here, in (b), when walking 1002 and crouching 1003 are plotted in the three-dimensional space including up to the third principal component, the possibility of separation may increase.
[0069] Therefore, if many principal components generated by the principal component analysis unit 404 are used, there is a possibility of highly accurate behavior classification. However, since the computational amount increases as the ordinal number k indicating the dimension of the principal component increases, it is necessary to determine how many principal components to consider from the perspective of accuracy and computational amount, and in what dimensional space to represent the behavior.
[0070] Therefore, the principal component analysis unit 404 changes the maximum ordinal number of the principal components used for learning in the behavior learning unit 406, and outputs a group of principal components from the first principal component to the principal component with the maximum ordinal number to the behavior learning unit 406. Specifically, for example, the required accuracy of behavior classification described above (for example, the ordinal number indicating the minimum required dimension of the principal component) or / and the allowable computational amount are set in advance, and the principal component analysis unit 404 changes the maximum ordinal number of the principal components used for learning in the behavior learning unit 406, and determines the ordinal number that maximally satisfies the required accuracy or / and the allowable computational amount.
[0071] For example, in the case where the required accuracy is the ordinal number "3" (the third principal component) indicating the dimension, the principal component analysis unit 404 determines the maximum ordinal number to be "3", and outputs a group of principal components from the first principal component to the third principal component to the behavior learning unit 406.
[0072] Also, when the allowable computational amount is set as a condition, the principal component analysis unit 404 sequentially obtains the computational amounts in ascending order from the first principal component, and determines the maximum ordinal number to be one less than the ordinal number (for example, "5") when the allowable computational amount is first exceeded (for example, "4"), and outputs a group of principal components from the first principal component to the fourth principal component with the maximum ordinal number k = 4 to the behavior learning unit 406.
[0073] Also, when the required accuracy is the ordinal number "3" (the third principal component) or higher indicating dimensions and the allowable calculation amount is set as a condition, if the cumulative calculation amount up to the third principal component is equal to or less than the allowable calculation amount, the principal component analysis unit 404 changes the maximum ordinal number from "3" to "4". Then, if the cumulative calculation amount up to the fourth principal component exceeds the allowable calculation amount, the principal component analysis unit 404 determines the maximum ordinal number k as "3" and outputs the principal component group from the first principal component to the third principal component to the behavior learning unit 406.
[0074] On the other hand, if the cumulative calculation amount up to the third principal component exceeds the allowable calculation amount, the principal component analysis unit 404 changes the maximum ordinal number from "3" to "2". Then, if the cumulative calculation amount up to the second principal component is equal to or less than the allowable calculation amount, the principal component analysis unit 404 determines the maximum ordinal number k as "2" and outputs the principal component group from the first principal component to the second principal component to the behavior learning unit 406.
[0075] Note that the principal component group output to the behavior learning unit 406 does not need to be limited to ascending order from the first principal component. For example, the principal component analysis unit 404 may extract a specific number of predetermined principal component groups. Also, the principal component analysis unit 404 may determine the principal component group to be output to the behavior learning unit 406 after excluding a specific principal component group. Thus, the principal component group output to the behavior learning unit 406 is not limited to the principal component group in ascending order from the first principal component.
[0076] Also, even in this case, when the allowable calculation amount is set as a condition, the principal component analysis unit 404 sequentially obtains the calculation amounts in ascending order of the ordinal numbers for the principal component group not limited to ascending order from the first principal component as described above, and outputs the principal component group up to the ordinal number one before the ordinal number when the allowable calculation amount is first exceeded to the behavior learning unit 406. For example, when the principal component group consists of the second principal component, the third principal component, and the fifth principal component, if the allowable calculation amount is not exceeded for the second principal component, nor for the second and third principal components, but is first exceeded for the second, third, and fifth principal components, the principal component analysis unit 404 may determine that the principal component group to be output to the behavior learning unit 406 is from the second principal component to the third principal component, which is one before the fifth principal component.
[0077] <Data Output to the Behavioral Learning Unit 406> FIG. 11 is an explanatory diagram showing data output from the principal component analysis unit 404 and the missing position information generation unit 422 to the behavioral learning unit 406. The principal component analysis unit 404 performs principal component analysis using the interpolated skeleton information 320, the joint angles 370, and the amount of movement between frames as input data to generate one or more principal components 1101.
[0078] The missing position information generation unit 422 generates missing position information 1100. The missing position information 1100 is, for example, indicated by variables prepared in the number of skeletons as flag information indicating whether each skeleton is missing among the skeleton information 320. For example, when the skeleton information 320 is composed of 18 skeleton points, the flag information is prepared with 18 variables, and for each skeleton point, if there is coordinate information of the skeleton point without missing, "1" is used as the flag information, and if it is missing, "0" is used as the flag information to generate the missing position information 1100.
[0079] Also, the missing position information 1100 may be generated as a multi-valued variable, where each value indicates which skeleton point is missing, and the data format of the missing position information 1100 is not limited. In this way, data 1102, which combines the principal component 1101 generated by the principal component analysis unit 404 and the missing position information 1100 generated by the missing position information generation unit 422, is output to the behavioral learning unit 406. Note that the data output from the principal component analysis unit 454 and the missing position information generation unit 462 to the behavior recognition unit 457 is also the same as the data 1102.
[0080] The behavioral learning unit 406 uses the data 1102 for learning, generates a behavior classification model, and outputs it to the behavior recognition unit 457.
[0081] FIG. 12 is an explanatory diagram showing a detailed method for the behavior learning unit 406 to learn behaviors and the behavior recognition unit 457 to classify behaviors. In FIG. 12, (a) shows the distribution of the plot points of the behaviors up to the second variable after dimensionality reduction, and (b) shows the behavior classification results considering the missing position information 1100. For each behavior in the principal component space, the behavior learning unit 406 generates boundary lines 1210 and boundary planes 1220 to classify each behavior by region. Any method such as the k-means method, support vector machine, decision tree, or random forest may be adopted for learning and classifying behaviors, and the behavior learning method is not limited.
[0082] The missing position information 1100 during the behavior learning of the behavior learning unit 406 is plotted on the principal component space and assigned to each behavior information. The plotted points for each behavior information (standing, sitting, walking, falling, raising a hand, squatting) are plotted points 1200 to 1205. Here, for simplicity, the plotted points 1200 to 1203 of the behaviors when no skeleton point is missing are represented as white points, and the plotted points 1204 and 1205 of the behaviors when skeleton point is missing are represented by being filled in.
[0083] The plotted point 1204 of the behavior with missing skeleton point is assigned the behavior "raising a hand", and the plotted point 1205 of the behavior with missing skeleton point is assigned the behavior information "squatting". Note that the plotted points of the same shape (1200, 1204), (1201, 1205) indicate that they show the same numerical information when the interpolated skeleton information 320 is subjected to principal component analysis.
[0084] Skeletal information 320 with deficiencies has its skeleton points interpolated by the deficiency information interpolation unit 423. However, since the interpolated skeleton points may contain a large amount of error compared to the detected skeleton points, there is a possibility of degrading the action classification accuracy when classifying actions in the same dimension as actions without deficiencies. Therefore, by classifying the skeletal information 320 by distinguishing between interpolated information and detected information, it is possible to improve the action classification accuracy. That is, by adding the deficiency position information 1100 to the principal component including the interpolated error, information of another dimension is added to the principal component, and by providing a boundary plane 1220 for classification on a new dimension, the accuracy of action classification is improved.
[0085] Note that the deficiency position information 1100 in FIG. 12 is represented by filling in the plot points 1204 and 1205, and thus is treated as 1-bit information. For example, the deficiency position information 1100 is represented by variables prepared for the number of skeleton points, and in this case, it is treated as multi-valued information. In this case, each action is plotted on a multi-dimensional space, and more detailed action classification can be realized by the deficiency position information 1100.
[0086] Note that the actions are held in the teacher signal DB104 as a combination of learning data (human skeletal information 320 and joint angles 370) and actions, and are not changed by the processing of the deficiency control unit 421 and the deficiency position information generation unit 422. The deficiency control unit 421 and the deficiency generation unit 402 deliberately generate deficiencies, and the deficiency position information generation unit 422 generates the deficiency position information 1100 of the generated deficiencies in order to correctly perform action classification on a new axis taking into account the deficiency position information 1100 even in the principal component including the error of interpolation due to deficiency.
[0087] The action recognition unit 457 recognizes actions using the action classification model learned and generated by the action learning unit 406. Specifically, for example, if there is missing information in the missing information interpolation unit 461 for the newly input skeletal information 320, the client 102 interpolates the missing information. The action recognition unit 457 applies principal component analysis to the interpolated skeletal information 320 and inputs the newly generated principal components and missing position information 1100 into the action classification model. As a result, the action recognition unit 457 determines to which region the newly input skeletal information 320 belongs according to the boundary line 1210 and the boundary plane 1220 set by the action classification model, and recognizes the action according to the determined region.
[0088] FIG. 13 is a graph showing the transition of the cumulative contribution rate used by the principal component analysis unit 404 when determining the number of dimensions. The cumulative contribution rate is a measure indicating how much of the information amount of the original data is represented by a plurality of newly generated principal components. Therefore, even if the number of principal components is increased and the number of dimensions at the time of action classification is increased, if there is no significant change in the cumulative contribution rate, no significant improvement in accuracy can be expected.
[0089] Therefore, the principal component analysis unit 404 determines the number of dimensions by using only the number of principal components necessary to exceed a predetermined cumulative contribution rate threshold. For example, when the predetermined cumulative contribution rate threshold is set to "0.8", since the conditions are satisfied if there are up to the second principal component, the number of dimensions k here is set to "2", and the first principal component and the second principal component are output to the action learning unit 406.
[0090] Note that the group of principal components output to the action learning unit 406 does not need to be limited to the principal components in ascending order from the first principal component. For example, the principal component analysis unit 404 may determine a combination of the ordinal number k of the principal component that does not exceed the predetermined cumulative contribution rate threshold and has the maximum cumulative contribution rate. Further, the principal component analysis unit 404 may select such a combination of the ordinal number k of the principal components from the group of principal components applied to the action classification model. In this way, the group of principal components output to the action learning unit 406 is not limited to the group of principal components in ascending order from the first principal component.
[0091] Note that the method for selecting the single or multiple principal component groups output by the principal component analysis unit 454 to the action recognition unit 457 shall be the same as that of the principal component analysis unit 404.
[0092] <Learning Process> FIG. 14 is a flowchart showing a detailed processing procedure example of the learning process by the server 101 (learning device) according to the first embodiment. The server 101 acquires, by the teacher signal acquisition unit 401, one or more teacher signals for learning from the teacher signal DB 104 among the acquired teacher signals (step S1400).
[0093] The server 101 sets the skeleton points to be missing according to a random number or a predetermined number by the missing control unit 421, and the missing generation unit 402 makes the skeleton points missing from the skeleton information 320 in the teacher signal acquired in step S1400 according to the setting result, and updates the missing skeleton information 320 as the skeleton information 320 in the teacher signal (step S1401).
[0094] The server 101 generates, by the missing position information generation unit 422, the position information of the skeleton points to be missing set by the missing control unit 421 as the missing position information (step S1410).
[0095] The server 101 interpolates the missing information (one or more pieces of missing skeleton information 320) from the non-missing information by the missing information interpolation unit 423 (step S1411). The teacher signal on which the missing information interpolation unit 423 has been executed is referred to as an updated teacher signal.
[0096] The server 101 executes skeleton information processing for each updated teacher signal by the skeleton information processing unit 403 (step S1402). Specifically, for example, the server 101 executes processing by the joint angle calculation unit 501, the movement amount calculation unit 502, and the normalization unit 503.
[0097] FIG. 15 is a flowchart showing a detailed processing procedure example of the skeleton information processing according to the first embodiment. The server 101 calculates the joint angle 370 from the skeleton information 320 in the update teacher signal for each update teacher signal by the joint angle calculation unit 501 (step S1501). Next, the server 101 calculates the amount of movement between frames from the skeleton information 320 in the update teacher signal for each update teacher signal by the movement amount calculation unit 502 (step S1501).
[0098] Then, the server 101 executes normalization in which the absolute position information is excluded from the skeleton information 320 for each update teacher signal by the normalization unit so that the size of the skeleton information 320 becomes constant (step S1403). As a result, for the update teacher signal, the joint angle 370, the amount of movement between frames, and the normalized skeleton information 320 are obtained. Then, the process proceeds to step S1403 in FIG. 13.
[0099] Returning to FIG. 14, the server 101 performs principal component analysis using the normalized skeleton information 320, the joint angle 370, and the amount of movement between frames as input data to generate one or more principal components. Among the generated principal components, it is determined how many dimensions of the principal components used for learning are to be used in descending order of the variance value, and the determined k-dimensional principal components (the first principal component to the k-th principal component) are selected in descending order of the variance value (step S1403).
[0100] In step S1407, for the skeleton information 320 that was made missing in step S1401, if there is still skeleton information 320 that has not been made missing (step S1407: No), the process returns to the process of step S1401, and the server 101 makes missing the skeleton points that have not been made missing so far (step S1401).
[0101] On the one hand, when all the skeleton information 320 is made missing (step S1407: Yes), the process proceeds to step S1408. However, the determination in the process of step S1407 is not limited to this. The server 101 may also determine whether to return to step S1401 or proceed to step S1408 according to a predetermined number of repetitions. Alternatively, the skeletons to be made missing may be predetermined, and the server 101 may also determine whether to return to step S1401 or proceed to step S1408 based on whether all the predetermined skeletons have been made missing.
[0102] In step S1408, for the teacher signal selected in step S1400, if there is a teacher signal that has not been selected yet (step S1408: No), the server 101 selects a teacher signal that has not been selected so far (step S1400). On the other hand, when all the teacher signals have been selected (step S1408: Yes), the process proceeds to step S1405. However, the determination in the process of step S1408 is not limited to this. The server 101 may also determine whether to return to step S1400 or proceed to the process of step S1405 according to a predetermined number of repetitions.
[0103] Note that the processes of step S1407 and step S1408 result in iterative processing, and the principal component analysis process of step S1403 occurs for the number of repetitions. However, the results of the principal component analysis processed in each iteration are additionally retained, and all the values are retained by the principal component analysis unit 404. When all the iterative processing is completed, all the values retained by the principal component analysis unit 404 are output to the behavior learning unit 406.
[0104] The server 101 uses the principal components input from the principal component analysis unit 404 and the missing position information input from the missing position information generation unit 422 as explanatory variables by the behavior learning unit 406, and uses the behavior information in the updated teacher signal as the target variable to perform learning, and as a result of the learning, generates a behavior classification model (step S1405).
[0105] <Behavior recognition process> FIG. 16 is a flowchart showing an example of an action recognition processing procedure by the client 102 (action recognition device) according to the first embodiment. The client 102 detects the skeleton information 320 of a person reflected in the analysis target data acquired from the sensor 103 by the skeleton detection unit 451 (step S1600). Next, the client 102 generates, by the missing position information generation unit 462, the position information of the skeleton points that could not be detected due to occlusion or the like among the detected skeleton information 320 as the missing position information (step S1601).
[0106] The client 102 interpolates the missing information from the non-missing information by the missing information interpolation unit 423 (step S1610).
[0107] Next, the client 102 executes skeleton information processing on the skeleton information 320 in which the missing information is interpolated in step S1610 in the same manner as the processing in step S1402 by the skeleton information processing unit 453 (step S1602). Specifically, for example, as shown in FIG. 15, the client 102 executes the processing by the joint angle calculation unit 501, the movement amount calculation unit 502, and the normalization unit 503 (steps S1501 to S1503).
[0108] Next, the client 102 performs principal component analysis on the skeleton information 320 normalized in step S1602, the joint angle 370, and the movement amount between frames as input data by the principal component analysis unit 454 to generate one or more principal components. The principal component analysis unit 454 determines the first principal component to the k-th principal component as the principal components to be used in the action recognition (step S1606) according to the ordinal number k determined in step S1403. (step S1603).
[0109] Next, the client 102, based on the action classification model learned in step S1305 by the action recognition unit 457, the principal components generated in step S1603, and the missing position information generated in step S1601, recognizes the actions of the person reflected in the analysis target data acquired from the sensor 103 (step S1606). The client 102 may transmit the recognition result of step S1606 to the server 101, or may also control the devices connected to the client 102 using the recognition result.
[0110] For example, when the analysis environment where the sensor 103 is deployed is a factory, the action recognition system 100 can be applied to work monitoring of workers in the factory and defect inspection of products using the recognition result. When the analysis environment is a train, the action recognition system 100 can be applied to monitoring of passengers in the train, monitoring of in-vehicle facilities, and detection of disasters such as fires using the recognition result.
[0111] Thus, according to the first embodiment, it is possible to accurately recognize a plurality of types of actions of the recognition target. Also, it is not necessary to generate a plurality of action classification models in order to accurately perform action recognition even in a partially missing shape. Thereby, it is possible to shorten the learning period and reduce the storage area for the action classification model. In particular, even when some of the skeleton points 300 to 317 are missing due to occlusion or the like, it is possible to accurately recognize a plurality of types of actions corresponding to the missing skeleton points without increasing the action classification model.
Embodiment
[0112] Embodiment 2 will be described centering on the differences from Embodiment 1. Regarding the points common to Embodiment 1, the same reference numerals are given and the description thereof is omitted.
[0113] FIG. 17 is a block diagram showing a functional configuration example of the action recognition system 100 according to Embodiment 2. In Embodiment 2, the teacher signal DB 104 holds the position information of the skeleton points to be missing by the missing control unit 421, and outputs the position information to the missing control unit 421.
[0114] When detecting the human skeleton information 320 reflected in the analysis target data acquired by the skeleton detection unit 451 from the sensor 103, the possibility of detection omission depends on the characteristics of the skeleton and the analysis target data. For example, the skeleton points at the ends of the human body such as the wrists and ankles are more likely to cause detection omissions. Also, for example, when the analysis target data is at a high angle, the probability of detection omission of the skeleton points of the lower body hidden by the upper body increases, and when shooting from a horizontal angle from the side of the passage, the detection omission of the skeleton points of the right half or the left half of the body is likely to occur.
[0115] When such characteristics are known in advance, it is desirable to cause such deficiencies and focus on learning to improve the action recognition accuracy. For this reason, the teacher signal DB104 holds the position information of the skeleton points to be deficient by the deficiency control unit 421. The position information of the skeleton points to be deficient by the deficiency control unit 421 is referred to as deficiency suggestion information. The deficiency control unit 421 acquires the deficiency suggestion information from the teacher signal DB104 and outputs it to the deficiency generation unit 402. The deficiency generation unit 402 causes a deficiency in the skeleton points according to the position information of the skeleton points to be deficient. Thereby, accurate action recognition can be achieved even in a situation where specific deficiencies are likely to occur in advance.
[0116] As described above, according to the second embodiment, in a situation where specific deficiencies are likely to occur in advance, the information is held in the teacher signal DB104, and by using the held information to intentionally cause deficiencies and perform learning, accurate action recognition can be achieved.
Embodiment
[0117] Embodiment 3 will be described mainly focusing on the differences from Embodiment 1 and Embodiment 2. Regarding the points common to Embodiments 1 to 3, the same reference numerals are given and the description thereof is omitted.
[0118] FIG. 18 is an explanatory diagram showing a detailed example of the teacher signal held by the teacher signal DB104 according to Example 3. In Example 3, the teacher signal DB104 holds environmental information 1801 and 1811 for use in learning as explanatory variables in the behavior learning unit 406. The environmental information 1801 and 1811 is information regarding the environment from which the analysis target data is acquired. For example, there is angle information where the sensor 103, which is the acquisition source of the analysis target data, is installed, or shape information of an object existing around where the sensor 103 is installed. In addition, information on the acquisition time zone of the analysis target data or variables independently generated by the developer may be added, and the environmental information 1801 and 1811 is not limited to these.
[0119] FIG. 19 is an explanatory diagram showing data output from the principal component analysis unit 404 and the missing position information generation unit 422 according to Example 3 to the behavior learning unit 406. The environmental information 1900 (1801, 1811) held by the teacher signal DB104 is output to the principal component analysis unit 404 along with the skeleton information 320. The environmental information 1900 is not the subject of the principal component analysis process, and is output to the behavior learning unit 406 as an explanatory variable together with the principal component 1101 and the missing position information 1100 generated by the principal component analysis unit 404.
[0120] FIG. 20 is a block diagram showing a functional configuration example of the behavior recognition system 100 according to Example 3. In Example 3, the teacher signal acquisition unit 401 is changed to the teacher signal acquisition unit 2000, and the environmental information detection unit 2001 is added.
[0121] The teacher signal acquisition unit 2000 acquires one or more teacher signals for learning regarding the teacher signal including environmental information from the teacher signal DB104, and outputs the selected teacher signal to the missing occurrence unit 402. Note that the environmental information 1900 is input to the behavior learning unit 406 and used for generating the behavior classification model.
[0122] The environmental information detection unit 2001 detects the environmental information 1900 defined in the teacher signal DB104 from the analysis target data acquired from the sensor 103, and outputs it to the skeleton detection unit 451. Note that the environmental information 1900 may depend on the installation position of the sensor 103 or the like. In this case, instead of the analysis target data, the environmental information 1900 at the time of installing the sensor 103 may be held in the environmental information detection unit 2001 in advance and output to the skeleton detection unit 451, and the method for detecting the environmental information 1900 is not limited.
[0123] <Action recognition processing> FIG. 21 is a flowchart showing an example of an action recognition processing procedure by the client 102 (action recognition device) according to the third embodiment. In the third embodiment, in FIG. 16, prior to skeleton detection (step S1600), environmental information detection processing (step 3300) is executed. The client 102 detects the environmental information 1900 defined in the teacher signal DB104 from the analysis target data acquired from the sensor 103 by the environmental information detection unit 2001 (step S2100). After that, steps S1600 to S1606 are executed. The environmental information 1900 is used as an explanatory variable in the action recognition (step S1606) by the action learning unit 406.
[0124] As described above, according to the third embodiment, since the teacher signal DB104 holds the environmental information 1900 and adds it as an explanatory variable during action learning, in addition to the missing position information in the principal component, the environmental information 1900 is added as information in another dimension, and a boundary for classification on a new dimension is set. Therefore, it is possible to improve the action classification accuracy.
Embodiment
[0125] The fourth embodiment will be described mainly with respect to the differences from the first to third embodiments. Note that the same reference numerals are given to the points common to the first to third embodiments, and the description thereof is omitted.
[0126] FIG. 22 is a block diagram showing a functional configuration example of the skeleton information processing unit according to Example 4. In Example 4, the skeleton information processing units 403 and 453 include a mutual information normalization unit 2204. The mutual information normalization unit 2204 normalizes the value ranges of the skeleton information 320, joint angles 370, and amount of movement between frames to be within a certain range before outputting them to the principal component analysis unit 404.
[0127] The value ranges of the skeleton information 320 and the amount of movement between frames depend on the resolution of the data to be analyzed. On the other hand, the value range of the joint angle 370 is in the range from 0 to 2π, or from 0 degrees to 360 degrees. When there are large differences in the value ranges of the data to be subjected to principal component analysis, there may be a bias for each data type in the influence on the principal components of the original data.
[0128] To eliminate this bias, the mutual information normalization unit 2204 performs normalization to set the value range of the data applied to the principal components within a certain range. For example, the mutual information normalization unit 2204 normalizes the value range of the original data to be from 0 to 2π for the skeleton information 320 according to the following formulas (21) to (22), and for the amount of movement between frames according to the following formula (23).
[0129]
Equation
[0130] However, the normalization method executed by the mutual information normalization unit 2204 is not limited to this. The mutual information normalization unit 2204 may, for example, normalize the value range of the joint angle 370 to be constant according to the magnitude of the resolution of the data to be subjected to principal component analysis.
[0131] FIG. 23 is a flowchart showing a detailed processing procedure example of the skeleton information processing unit according to Example 4. In Example 4, in the skeleton information processing (steps S1402 and S1602), after normalization (step S1503), the client 102 executes mutual information normalization (step S2304). In the mutual information normalization (step S2304), the possible value ranges of the skeleton information 320 normalized by the normalization unit, the joint angles 370, and the amount of movement between frames are normalized to be constant.
[0132] Thus, according to Example 4, by uniformly normalizing the possible value ranges of the original data (skeleton information 320, joint angles 370, and amount of movement between frames) for which principal component analysis is performed, the bias in the influence on the principal components by specific data having a wide value range is eliminated, and multiple types of actions can be discriminated with high accuracy.
Example
[0133] Example 5 will be described centering on the differences from Examples 1 to 4. For points common to Examples 1 to 4, the same reference numerals are given and the description thereof is omitted.
[0134] FIG. 24 is a block diagram showing a functional configuration example of the action recognition system 100 according to Example 5. In Example 5, the principal component analysis units 404 and 454 are changed to dimensionality reduction units 2400 and 2401. Dimensionality reduction is a process of reducing the number of original variables or the number of original dimensions while maintaining the original amount of information as much as possible, and is a concept that includes component analysis such as principal component analysis and independent component analysis in Examples 1 to 4.
[0135] The dimensionality reduction unit 2400 uses, as input data, the normalized skeleton information 320, the joint angles 370, and the amount of movement between frames among the teacher signals acquired from the skeleton information processing unit 403, executes dimensionality reduction to generate one or more variables, and outputs them to the action learning unit 406.
[0136] As methods of dimensionality reduction performed by the dimensionality reduction unit 2400, there are methods such as SNE (Stochastic Neighbor Embedding), t-SNE (t-Distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), Isomap, LLE (Locally Linear Embedding), Laplacian Eignmap, LargeVis, and diffusion maps. The dimensionality reduction unit 2400 may perform dimensionality reduction by combining principal component analysis or independent component analysis with t-SNE or UMAP. Hereinafter, each method of dimensionality reduction and the method of dimensionality reduction performed by combining each method will be described.
[0137] The process of SNE will be described using the following formulas (24) to (28).
[0138]
Equation
[0139] x i and x j For the similarity of two x coordinate values 322 (input data), when x i is given, the conditional probability p j for selecting x j|i as a neighbor is set. The conditional probability p j|i is shown in the above formula (24). At this time, it is assumed that x j is selected based on a normal distribution centered on x i . Next, the similarity of two y coordinate values 323 (principal components) of y i and y j after dimensionality reduction is also set as the conditional probability q j|i i and x j shown in the above formula (25), similar to the similarity of xbefore dimensionality reduction. However, the variance of the coordinate values after dimensionality reduction is fixed at 1 / √2 to simplify the formula.
[0140] If y for dimensionality reduction is generated so as to maintain the distance relationship before and after dimensionality reduction, it is possible to perform dimensionality reduction while maintaining the amount of information as much as possible. In order to perform dimensionality reduction while suppressing the reduction of the amount of information, the dimensionality reduction unit 2400 sets p j|i =q j|i and performs processing accordingly. Kullback-Leibler divergence, which is a measure indicating how similar two probability distributions are, is used for dimensionality reduction.
[0141] The formula adapting the probability distributions before and after dimensionality reduction using Kullback-Leibler divergence as a loss function is shown in the above formula (26). The dimensionality reduction unit 2400 minimizes the above formula (26), which is a loss function, by means of stochastic gradient descent. This gradient uses the above formula (27) obtained by differentiating the loss function with respect to y i to vary y i . The update formula at the time of this variation is shown in the above formula (28).
[0142] As described above, while varying y i , the above formula (28) is updated to obtain y i for which the above formula (27) is minimized, thereby performing dimensionality reduction to obtain a new variable. However, in the case of SNE, unlike principal component analysis, due to the characteristics of the processing, the number of dimensions (variables) after reduction becomes two or three types. Therefore, when performing dimensionality reduction by SNE, the predetermined number of dimensions (variables) is output to the behavior learning unit 406 in advance.
[0143] However, in SNE, it is difficult to minimize the loss function, and there is a problem that the skeleton points specified by the x coordinate value 322 and the y coordinate value 323 become dense in an attempt to maintain equidistance during dimensionality reduction. t-SNE is a solution method for this problem.
[0144] The processing of t-SNE will be described using the following formulas (29) to (33).
[0145]
Equation
[0146] To simplify the minimization of the loss function, the loss function is symmetrized. In the symmetrization process of the loss function, as shown in the above formula (29), the distance between x i and x j is represented by the joint probability distribution p ij . p j|i is the same as the above formula (24) and can be expressed by the above formula (30). Also, the distance between y i and y j after dimensionality reduction is represented by the joint probability distribution q ij shown in the above formula (31).
[0147] The distance between points after dimensionality reduction assumes a Student's t-distribution. The Student's t-distribution is characterized by a higher probability of values deviated from the mean compared to the normal distribution. This characteristic makes it possible to allow the distribution of long distances for the distances between data after dimensionality reduction.
[0148] In t-SNE, the dimensionality reduction unit 2400 performs dimensionality reduction by minimizing the loss function shown in the above formula (32) using pij and q ij obtained from the above formulas (29) to (31). The dimensionality reduction unit 2400 uses the stochastic gradient descent method shown in the above formula (33) for minimizing the loss function, similar to SNE.
[0149] By obtaining y i for which the above formula (33) is minimized, the dimensionality reduction unit 2400 performs dimensionality reduction to obtain new variables. Similar to SNE, due to the characteristics of the process, the number of dimensions (variables) after reduction in t-SNE becomes two or three types. Therefore, when performing dimensionality reduction by t-SNE, the predetermined number of dimensions (variables) is output to the behavior learning unit 406 in advance.
[0150] t-SNE can accurately perform dimensionality reduction because it preserves the local structure of the high dimensions before dimensionality reduction and captures the global structure as much as possible. However, there is a problem that the calculation time increases according to the number of dimensions before dimensionality reduction. As a method for solving this problem of the calculation time of dimensionality reduction, there is UMAP. The process of UMAP will be described using the following formulas (34) to (36).
[0151] [Mathematics]
[0152] Among the entire set A of possible values, there is a high-dimensional set X (the above formula (34)). When any data is taken out from A, let μ be the membership function that outputs, within the range of 0 to 1, the degree to which it is included in the set X. For the input X shown in the above formula (1), prepare Y shown in the above formula (2). Y is a set of m (<p) points existing in a space of a lower dimension compared to X and is a set of data after dimensionality reduction. Then, with ν as the membership function of Y, the dimensionality reduction unit 2400 performs dimensionality reduction by determining Y such that the above formula (36) becomes the minimum, and obtains a new variable.
[0153] When performing dimensionality reduction by UMAP, the dimensionality reduction unit 2400 may output a predetermined number of dimensions (variables) to the behavior learning unit 406 in the same manner as SNE or t-SNE, or may output, as the necessary number of dimensions, the number of dimensions (variables) such that the membership function ν after dimensionality reduction is equal to or greater than a predetermined value range, to the behavior learning unit 406.
[0154] Describe the processing of Isomap. The dimensionality reduction unit 2400 calculates the shortest distance to the data in the neighborhood for any data, and performs dimensionality reduction by representing the calculated distance as a geodesic distance matrix by multidimensional scaling (MDS), and obtains a new variable. When performing dimensionality reduction by Isomap, the dimensionality reduction unit 2400 outputs a predetermined number of dimensions (variables) to the behavior learning unit 406.
[0155] Explain LLE using the following formulas (37) to (43).
[0156] [Mathematics]
[0157] x iPoints near are approximately represented by the linear combination in the above formula (37). Here, by minimizing the above formula (39) under the constraint of the above formula (38), the approximation value of x before dimensionality reduction is determined. i Next, for y after dimensionality reduction i , in order to maintain the linear adjacency relationship of x i as much as possible even after dimensionality reduction, the dimensionality reduction unit 2400 minimizes the above formula (40). This solution is obtained as shown in the above formula (42) by extracting the eigenvectors of the above formula (41) from the second smallest eigenvalue v i to the (d + 1)-th v d , and the dimensionality reduction unit 2400 obtains y after dimensionality reduction i as shown in the above formula (43).
[0158] When performing dimensionality reduction by LLE, the dimensionality reduction unit 2400 outputs a predetermined number of dimensions (variables) to the behavior learning unit 406 in advance, and the behavior learning unit 406 may determine the number of variables to be used according to the predetermined number of dimensions.
[0159] The processing of Laplacian eigenmaps is described using the following formulas (44) to (49).
[0160]
Equation
[0161] Each edge x i x j of the neighborhood graph generated by the data before dimensionality reduction is assigned to the above formula (44) or the above formula (45). The graph Laplacian of the above formula (46) is introduced for the assigned weights, and the eigenvectors of the graph Laplacian (the above formula (47)) are obtained as shown in the above formula (48) by extracting from the second smallest eigenvalue v i to the (d + 1)-th v d , and the dimensionality reduction unit 2400 obtains the value y after dimensionality reduction i as shown in the above formula (49).
[0162] When performing dimensionality reduction using the Laplacian eigenmap, the dimensionality reduction unit 2400 outputs a predetermined number of dimensions (variables) to the behavior learning unit 406.
[0163] The processing of LargeVis will be described. LargeVis is a method that improves the calculation time of t-SNE. In t-SNE, since the distance between data points is obtained, the calculation time increases according to the number of data. In LargeVis, the dimensionality reduction unit 2400 divides the data into regions using a K-NN graph from neighboring data, and performs dimensionality reduction on each data model divided by region using the same method as t-SNE.
[0164] When performing dimensionality reduction by LargeVis, the dimensionality reduction unit 2400 outputs a predetermined number of dimensions (variables) to the behavior learning unit 406.
[0165] The diffusion map will be described using the following equations (50) to (55).
[0166]
Equation
[0167] x before dimensionality reduction i and each edge x of the neighborhood graph composed of xj in the neighborhood i x j is assigned a weight W ij and this is normalized to create the N×N transition probability matrix P shown in the above equation (50). p t (x i x j ) represents the probability of reaching x i after t steps starting from x j by a random walk on the graph represented by P. From the properties of the transition matrix, p t (x i x j ) converges to the stationary distribution φ0(x j ) as t→∞. At this time, the point x i x jDefine the diffusion distance by the above formula (51). Let the eigenvalue of the transition probability matrix P be the above formula (52) and the eigenvector be the above formula (53). At this time, the above formula (54) holds. λ i Since the absolute value of λ is 1 or less, the dimensionality reduction unit 2400 takes eigenvectors up to an appropriate dimension d(t) smaller than N, performs dimensionality reduction as in the above formula (55), and obtains a new variable.
[0168] When performing dimensionality reduction by diffusion maps, the predetermined number of dimensions (variables) is output to the behavior learning unit 406 in advance.
[0169] The dimensionality reduction unit 2400 may be implemented by combining the principal component analysis, independent component analysis, t-SNE, UMAP, Isomap, LLE, Laplacian eigenmaps, LargeVis, diffusion maps, etc. described so far. For example, the dimensionality reduction unit 2400 performs dimensionality reduction up to 10 dimensions using principal component analysis on high-dimensional data with 36 dimensions or 36 variables, and then uses UMAP for dimensionality reduction up to 2 dimensions. The combination of methods used for dimensionality reduction is not limited. By combining various methods in this way during dimensionality reduction, composite effects can be expected on performance and calculation time.
[0170] Also, these dimensionality reduction methods are not limited to the scope described in Example 5. For example, simply adding, subtracting, multiplying, or dividing high-dimensional information, or convolving according to a predetermined coefficient may be possible. As long as it is a method that generates lower-dimensional data or a smaller number of variables from high-dimensional data or multivariate variables like the methods described in Example 5, the dimensionality reduction method is not limited.
[0171] The dimensionality reduction unit 2401 has the same function as the dimensionality reduction unit 2400. The dimensionality reduction unit 2401 performs the same processing as the dimensionality reduction unit 2400 on the output data from the skeleton information processing unit 453 to generate a smaller number of single or multiple new variables compared to before dimensionality reduction. Also, the dimensionality reduction unit 2401 outputs new variables determined from the contribution rate and cumulative contribution rate generated together with the principal components to the behavior recognition unit 457.
[0172] Thus, according to the fifth embodiment, by changing the dimensionality reduction method, it becomes possible to effectively reduce the dimensionality according to the data obtained from the skeleton information processing unit 403 or shorten the calculation time, and it is possible to accurately discriminate complex behaviors.
Embodiment
[0173] The sixth embodiment will be described mainly focusing on the differences from the first to fifth embodiments. Regarding the points common to the first to fifth embodiments, the same reference numerals are given and the description thereof is omitted.
[0174] FIG. 25 is a block diagram showing a functional configuration example of the behavior recognition system 100 according to the sixth embodiment. In the sixth embodiment, the behavior learning unit 406 and the behavior recognition unit 457 are changed to a behavior learning unit 2500 and a behavior recognition unit 2501. A detailed method for the behavior learning unit 2500 and the behavior recognition unit 2501 to classify behaviors will be described with reference to FIGS. 26 to 28.
[0175] FIG. 26 is an explanatory diagram showing a decision tree which is a basic method for the behavior learning unit 2500 and the behavior recognition unit 2501 to classify behaviors. A behavior classification method using a decision tree will be described. In the decision tree, for each behavior in the newly generated variable space after dimensionality reduction, using the behavior (plot points 1200 to 1203) which is a variable with a previously given behavior type, (a) the boundary line 2610 is generated.
[0176] (a) A method for generating the boundary line 2610 will be described. The decision tree classifies actions step by step so that the impurity of the variable group 2621 input from the population 2620 including actions (plot points 1200 to 1203) is minimized. In the first step, on the second variable axis, the variable group 2621 is classified into the variable group 2622 and the variable group 2623. In the second step, on the first variable axis, the variable groups 2622 and 2623 are classified into the variable groups 2624 to 2627. In this way, the discriminant obtained in the process of classification to minimize the impurity is used to generate the (a) boundary line 2610. Note that there is no limitation on which axis to classify actions at each step, and there is no limitation such as classifying the action classification on each axis a specified number of times such as once.
[0177] FIG. 27 is an explanatory diagram showing a detailed development method of classification by a decision tree. There are a level-wise 2700 for growing the decision tree for each level (depth) and a leaf-wise 2701 for growing the decision tree for each leaf (data group after branching) in the decision tree. The method of stacking classifiers like a decision tree for learning is called ensemble learning.
[0178] FIG. 28 is an explanatory diagram showing ensemble learning and the methods used by the action learning unit 2500 and the action recognition unit 2501 to classify actions. Ensemble learning includes bagging 2801 that uses classification trees like decision trees in parallel and boosting 2802 that updates the learning result by inheriting the previous result. The random forest of Example 1 is a method that employs bagging 2801 for decision trees, and the action learning unit 2500 and the action recognition unit 2501 of Example 6 are classification methods that use boosting 2802.
[0179] When the action learning unit 2500 learns actions and the action recognition unit 2501 classifies actions, each decision tree may be grown level-wise and the variables input by boosting that stacks a plurality of decision trees may be classified, or each decision tree may be grown leaf-wise and the variables input by boosting that stacks a plurality of decision trees may be classified.
[0180] When adopting boosting, which grows each decision tree level by level and stacks multiple decision trees, as an action classification method, it may be implemented using the software library xgboost. On the other hand, when adopting boosting, which grows each decision tree leaf by leaf and stacks multiple decision trees, as an action classification method, it may be implemented using the software library LightGBM. However, the implementation method is not limited to these.
[0181] Thus, according to Example 6, by using boosting as an action classification method and stacking multiple decision trees, complex actions can be discriminated with high accuracy.
Example
[0182] Example 7 will be described centering on the differences from Examples 1 to 6. Regarding the points common to Examples 1 to 6, the same reference numerals are given and the description thereof is omitted.
[0183] FIG. 29 is a block diagram showing a functional configuration example of the action recognition system 100 according to Example 7. In Example 7, the dimensionality reduction unit 2400, the action learning unit 406, the dimensionality reduction unit 2401, and the action recognition unit 457 are changed to the dimensionality reduction unit 2900, the action learning unit 2901, the dimensionality reduction unit 2903, and the action recognition unit 2904.
[0184] The dimensionality reduction unit 2900 performs dimensionality reduction by any of the methods of Examples 1 to 6 according to a predetermined number of dimensions, and outputs the variables after dimensionality reduction to the action learning unit 2901.
[0185] The action learning unit 2901 generates a boundary line for action classification by machine learning from the given action types together with the obtained variables after dimensionality reduction, and generates an action classification model. At this time, the action classification accuracy, which indicates how accurately actions can be predicted for the generated action classification model, is calculated.
[0186] The behavior learning unit 2901 may calculate the behavior classification accuracy using the variables used for generating the behavior classification model. The behavior learning unit 2901 may calculate the behavior classification accuracy using some of the variables obtained from the dimensionality reduction unit 2900, without using some of them for generating the behavior classification model and using the variables not used for behavior classification generation. However, the method for calculating the behavior classification accuracy is not limited to these. If the calculated behavior classification accuracy is higher than a predetermined accuracy, the behavior learning unit 2901 outputs the generated behavior classification model to the behavior recognition unit 2904. At this time, the behavior learning unit 2901 also outputs to the dimensionality reduction unit 2900 the obtained number of dimensions and the fact that the behavior classification accuracy has passed.
[0187] On the other hand, if the calculated behavior classification accuracy is lower than a predetermined accuracy, the behavior learning unit 2901 outputs to the dimensionality reduction unit 2900 the fact that the behavior classification accuracy has failed. However, if the behavior classification model is generated with all the configurable number of dimensions (variables) and the behavior classification accuracy fails for all of them, the behavior learning unit 2901 outputs to the behavior recognition unit 2904 the behavior classification model with the highest behavior classification accuracy among the behavior classification models generated so far, and outputs to the behavior learning unit 2901 the number of dimensions (variables) used at the time of output together with the all learning completion information. Note that without determining the behavior classification accuracy for judging pass or fail, the behavior learning unit 2901 may perform learning with all the configurable number of dimensions, calculate the behavior classification accuracy, and then determine the behavior classification model according to the calculated behavior classification accuracy, and judge the determined behavior classification model as passed.
[0188] According to the pass / fail information and all learning completion information obtained from the behavior learning unit 2901, when the dimensionality reduction unit 2900 obtains pass or all learning completion information, it outputs the obtained dimensionality information to the dimensionality reduction unit 2903. When it is determined as failed, it changes the number of dimensions used for dimensionality reduction and executes dimensionality reduction again, and outputs the generated variables to the behavior learning unit 2901.
[0189] The dimensionality reduction unit 2903 performs dimensionality reduction on the data obtained from the skeleton information processing unit 453 according to the number of dimensions k (variable) obtained from the dimensionality reduction unit 2900 using the dimensionality reduction methods of Examples 1 to 6, and outputs the generated variable to the action recognition unit 2904.
[0190] The action recognition unit 2904 executes action recognition using the variable input from the dimensionality reduction unit 2903 using the action classification model determined to be qualified by the action learning unit 2901.
[0191] Note that the action classification accuracy calculated by the action learning unit 2901 may be regarded as the contribution rate described in Example 1. For example, associate the obtained variable after dimensionality reduction with the action classification accuracy calculated using it, and set the calculated action classification accuracy as the contribution rate to the original information of the variable after dimensionality reduction used in the calculation. The dimensionality reduction unit 2900 determines which variable after dimensionality reduction to use for control according to the thus regarded contribution rate.
[0192] <Learning process> FIG. 30 is a flowchart showing a detailed processing procedure example of the learning process by the server 101 (learning device) according to Example 7. The server 101 determines the number of dimensions k by the dimensionality reduction unit 2900. At this time, when performing dimensionality reduction for the first time, a predetermined number of dimensions k is determined, and in the case of dimensionality reduction after the second time, a number of dimensions k that has not been determined so far is determined. The dimensionality reduction unit 2900 performs dimensionality reduction according to the determined number of dimensions k and generates a new variable (the number of dimensions k after dimensionality reduction) (S3001).
[0193] In step S3002, the server 101 makes a pass / fail determination on the action classification accuracy obtained from the action learning unit 2901. If it is qualified, the learning process ends, and if it is unqualified, it returns to step S3001.
[0194] As described above, according to Example 7, by changing the number of dimensions in accordance with the target action classification accuracy and repeating dimensionality reduction, complex actions can be discriminated with high accuracy.
[0195] Also, the action recognition device and the learning device of the above-described Examples 1 to 7 can also be configured as follows [1] to
[14] .
[0196] [1] An action recognition device (client 102) having a processor 201 that executes a program and a storage device 202 that stores the program can access an action classification model learned using a component group related to a learning target obtained from the shape (skeleton information 320) of the learning target by component analysis (principal component analysis or independent component analysis) that generates statistical components in multivariate analysis, and the action of the learning target. The processor 201 performs a detection process of detecting the shape (skeleton information 320) of the recognition target from the analysis target data obtained from the sensor 103, a missing position information generation process of generating missing position information indicating the position of the missing part among the skeleton information 320 detected by the detection process, an interpolation process of interpolating the missing part and updating the non-missing information after interpolation with the missing part as the shape of the recognition target, a component analysis process of generating the same number of component groups related to the recognition target as the component groups related to the learning target based on the shape of the recognition target interpolated by the interpolation process by the component analysis, and an action recognition process of outputting a recognition result indicating the action of the recognition target by inputting the component group related to the recognition target generated by the component analysis process and the missing position information to the action classification model.
[0197] Thereby, compared with the case of selecting a specific action classification model from a plurality of action classification models by component analysis and executing the action recognition process, it is possible to reduce the usage amount of the storage device 202 and speed up the process until the recognition result is output. In addition, it is possible to accurately recognize a plurality of types of actions of a recognition target with a part of the shape missing.
[0198] [2] In the action recognition device of [1] above, in the action recognition process, the processor inputs the environmental information indicating the environment of the acquisition source of the analysis target data, the component group related to the recognition target, the action to be learned, and the missing position information into the action classification model, and outputs a recognition result indicating the action of the recognition target.
[0199] Thereby, since an action classification model considering environmental information is prepared, it is possible to accurately recognize an action according to the environment of the recognition target in the recognition target.
[0200] [3] In the action recognition device of [1] above, the action classification model is learned using the ascending component group of the first variable related to the learning target obtained from the shape of the learning target by dimensionality reduction (principal component analysis or independent component analysis or SNE (Stochastic Neighbor Embedding) or t-SNE (t-Distributed Stochastic Neighbor Embedding) or UMAP (Uniform Manifold Approximation and Projection) or Isomap or LLE (Locally Linear Embedding) or Laplacian Eignmap or LargeVis or diffusion map) that generates statistical components by multivariate analysis, and the action of the learning target. In the generation process, the processor generates an ascending component group of the first variable related to the recognition target having the same number as the ascending component group of the first variable related to the learning target based on the shape of the recognition target interpolated by the interpolation process by the dimensionality reduction.
[0201] Thereby, compared with the case of selecting a specific action classification model from a plurality of action classification models by dimensionality reduction and executing the action recognition process, it is possible to reduce the usage amount of the storage device 202 and speed up the process until the recognition result is output.
[0202] [4] In the action recognition device of [1] above, each action classification model in the action classification model group is learned using a component group related to the learning target obtained from the shape of the learning target and the angles of a plurality of vertices (joint angles 370) that make up the shape, and the action of the learning target. The processor 201 executes a calculation process for calculating the angles of a plurality of vertices (joint angles 370) that make up the shape of the recognition target based on the shape of the recognition target. In the component analysis process, the processor 201 generates a component group related to the recognition target based on the shape of the recognition target and the angles of the vertices of the recognition target calculated by the calculation process.
[0203] Thereby, according to the change in shape caused by the angle of the vertex, a plurality of types of actions of the recognition target can be recognized with high accuracy.
[0204] [5] In the action recognition device of [1] above, each action classification model in the action classification model group is learned using a component group related to the learning target obtained from the shape of the learning target and the movement amount of the learning target, and the action of the learning target. The processor 201 executes a calculation process for calculating the movement amount of the recognition target based on a plurality of shapes of the recognition target at different times. In the component analysis process, the processor 201 generates a component group related to the recognition target based on the shape of the recognition target and the movement amount of the recognition target calculated by the calculation process.
[0205] Thereby, according to the temporal change in shape caused by movement, a plurality of types of actions of the recognition target can be recognized with high accuracy.
[0206] [6] In the action recognition device of [1] above, the processor 201 executes a first normalization process for normalizing the size of the shape of the recognition target. In the component analysis process, the processor 201 generates a component group related to the recognition target based on the shape of the recognition target after the first normalization by the first normalization process.
[0207] This makes it possible to suppress misrecognition by improving the versatility of action classification.
[0208] [7] In the action recognition device of [1] above, the processor 201 executes a second normalization process for normalizing the value range that the shape and vertex angles of the recognition target can take. In the component analysis process, the processor 201 generates a component group related to the recognition target based on the shape and vertex angles (joint angles 370) of the recognition target after the second normalization by the second normalization process.
[0209] This makes it possible to suppress the bias in the value ranges of different data types such as shape and angle, and to improve the accuracy of action recognition.
[0210] [8] In a learning device having a processor 201 that executes a program and a storage device 202 that stores the program, the processor 201 performs an acquisition process of acquiring teacher data including the shape and action of a learning target, a deletion process of deleting the shape of the learning target acquired by the acquisition process, a deletion position information generation process of generating deletion position information indicating the position of the deletion part deleted from the shape of the learning target by the deletion process, an interpolation process of interpolating from non-deleted information which is a part other than the deletion part deleted from the shape of the learning target by the deletion process, and updating the interpolated non-deleted information as the shape of the learning target, and a component analysis process of generating a component group related to the learning target based on the shape of the learning target interpolated by the interpolation process by component analysis (principal component analysis or independent component analysis) for generating statistical components by multivariate analysis, and an action learning process of generating an action classification model for learning the action of the learning target and classifying the action of the learning target based on the component group related to the learning target generated by the component analysis process, the action of the learning target, and the deletion position information.
[0211] This eliminates the need to select a specific action classification model from a plurality of action classification models in the action recognition device.
[0212] [9] In the learning device of [8] above, the processor 201 executes a calculation process of calculating the angles (joint angles 370) of a plurality of vertices constituting the shape of the learning target based on the shape of the learning target, and in the component analysis process, the processor 201 generates a component group related to the learning target based on the shape of the learning target and the angles of the vertices of the learning target calculated by the calculation process.
[0213] As a result, since it is possible to prepare a behavior classification model according to the change in shape caused by the angle of the vertex, it is possible to accurately recognize a plurality of types of behaviors according to the change in shape caused by the angle of the vertex of the recognition target.
[0214]
[10] In the learning device of [8] above, the processor 201 executes a calculation process of calculating the movement amount of the learning target based on a plurality of shapes of the learning target at different times, and in the component analysis process, the processor 201 generates a component group related to the learning target based on the shape of the learning target and the movement amount of the learning target calculated by the calculation process.
[0215] As a result, since it is possible to prepare a behavior classification model according to the temporal change in shape caused by movement, it is possible to accurately recognize a plurality of types of behaviors according to the temporal change in shape caused by movement.
[0216]
[11] In the learning device of [8] above, the processor 201 executes a first normalization process of normalizing the size of the shape of the learning target, and in the component analysis process, the processor 201 generates a component group related to the learning target based on the shape of the learning target after the first normalization by the first normalization process.
[0217] As a result, by improving the generality of behavior classification learning, it is possible to suppress mislearning.
[0218]
[12] In the learning device of [9] above, the processor 201 executes a second normalization process for normalizing the value ranges that the shape and vertex angles of the learning target can take. In the component analysis process, the processor 201 generates a component group related to the learning target based on the shape and vertex angles of the learning target after the second normalization by the second normalization process.
[0219] Thereby, the bias in the value ranges in different data types such as shape and angle can be suppressed, and the high-precision of the behavior classification learning can be achieved.
[0220]
[13] In the learning device of [8] above, in the missing process, the processor makes a specific part of the shape of the learning target missing.
[0221] Thereby, the shape of the learning target that is likely to be missing can be learned, and the high-precision of the behavior classification learning can be achieved.
[0222]
[14] In the learning device of [8] above, when the processor 201 acquires the shape of the learning target in the acquisition process, it acquires the environmental information at the time of acquiring the shape of the learning target, and generates a behavior classification model based on the component group, the behavior of the learning target, the missing position information, and the environmental information.
[0223] Thereby, a behavior classification model corresponding to the shape of the learning target and the environmental information can be generated, and the high-precision of the behavior classification learning adapted to various environments can be achieved.
[0224] Note that the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Further, the configuration of another embodiment may be added to the configuration of one embodiment. Also, for a part of the configuration of each embodiment, addition, deletion, or replacement with another configuration may be made.
[0225] In addition, each of the above-described configurations, functions, processing units, processing means, etc. may be realized in hardware by designing a part or all of them, for example, by using an integrated circuit, or may be realized in software by the processor 201 interpreting and executing a program for realizing each function.
[0226] Information such as programs, tables, files, etc. for realizing each function can be stored in a storage device such as a memory, a hard disk, an SSD (Solid State Drive), or a recording medium such as an IC (Integrated Circuit) card, an SD card, or a DVD (Digital Versatile Disc).
[0227] Also, the control lines and information lines show those considered necessary for explanation, and do not necessarily show all the control lines and information lines necessary for implementation. In practice, it may be considered that almost all the configurations are interconnected.
Explanation of Reference Numerals
[0228] 100 Action Recognition System 101 Server 102 Client 103 Sensor 104 Teacher Signal DB 201 Processor 202 Memory Device 320 Skeleton Information 401, 2000 Teacher Signal Acquisition Unit 402 Defect occurrence part 421 Defect control part 403, 453 Skeleton information processing part 404, 454 Principal component analysis part 406, 2500, 2901 Behavior learning part 422, 462 Defect position information generation part 423, 461 Defect information interpolation part 451 Skeleton detection part 455 Dimension determination part 457, 2501, 2904 Behavior recognition part 501 Joint angle calculation part 502 Movement amount calculation part 503 Normalization part 2001 Environment information detection part 2204 Mutual information normalization part 2400, 2401, 2900, 2903 Dimension reduction part
Claims
1. An action recognition device having a processor that executes a program and a storage device that stores the program, which is accessible to an action classification model learned using a component group related to the learning target obtained from the shape of the learning target by component analysis that generates statistical components by multivariate analysis and the action of the learning target, wherein the processor performs a detection process for detecting the shape of the recognition target from the analysis target data, a missing position information generation process for generating missing position information indicating the position of a missing portion among the shapes of the recognition target detected by the detection process, an interpolation process for interpolating the missing portion from non-missing information, which is a portion other than the missing portion among the shapes of the recognition target including the missing portion, and updating the non-missing information after interpolation as the shape of the recognition target, a component analysis process for generating, by the component analysis, a component group related to the recognition target having the same number as the component group related to the learning target based on the shape of the recognition target interpolated by the interpolation process, and an action recognition process for outputting a recognition result indicating the action of the recognition target by inputting the component group related to the recognition target generated by the component analysis process and the missing position information to the action classification model. An action recognition device characterized by performing the above.
2. The action recognition device according to claim 1, wherein in the action recognition process, the processor inputs environmental information indicating the environment where the analysis target data is acquired, the component group related to the recognition target, the action of the learning target, and the missing position information to the action classification model, and outputs a recognition result indicating the action of the recognition target. An action recognition device characterized by the above.
3. The action recognition device according to claim 1, wherein the action classification model is learned using an ascending component group from a first variable related to the learning target obtained from the shape of the learning target by dimensionality reduction that generates statistical components by multivariate analysis and the action of the learning target, and in the component analysis process, the processor generates, by the dimensionality reduction, an ascending component group related to the recognition target having the same number as the ascending component group related to the learning target from the first variable based on the shape of the recognition target interpolated by the interpolation process. An action recognition device characterized by the above.
4. The action recognition device according to claim 1, The action classification model is learned using a component group related to the learning target obtained from the shape of the learning target and the angles of a plurality of vertices constituting the shape, and the action of the learning target. The processor executes a calculation process for calculating the angles of a plurality of vertices constituting the shape of the recognition target based on the shape of the recognition target, In the component analysis process, the processor generates a component group related to the recognition target based on the shape of the recognition target and the angles of the vertices of the recognition target calculated by the calculation process. An action recognition device characterized by the above.
5. The action recognition device according to claim 1, wherein the action classification model is learned using a component group related to the learning target obtained from the shape of the learning target and the amount of movement of the learning target, and the action of the learning target. The processor executes a calculation process for calculating the amount of movement of the recognition target based on a plurality of shapes of the recognition target at different times, In the component analysis process, the processor generates a component group related to the recognition target based on the shape of the recognition target and the amount of movement of the recognition target calculated by the calculation process. An action recognition device characterized by the above.
6. The action recognition device according to claim 1, The processor executes a first normalization process for normalizing the size of the shape of the recognition target, In the component analysis process, the processor generates a component group related to the recognition target based on the shape of the recognition target normalized by the first normalization process. An action recognition device characterized by the above.
7. The action recognition device according to claim 1, The processor executes a second normalization process for normalizing the value range that the shape and vertex angles of the recognition target can take, In the component analysis process, the processor generates a component group related to the recognition target based on the shape and vertex angles of the recognition target after the second normalization by the second normalization process. An action recognition device characterized by the above.
8. A learning device having a processor that executes a program and a storage device that stores the program, The processor performs an acquisition process for acquiring teacher data including the shape and action of the learning target, a deletion process for deleting the shape of the learning target acquired by the acquisition process. A missing position information generation process that generates missing position information indicating the position of a missing portion that has been removed from the shape of the learning target by the missing process; An interpolation process that interpolates from non-missing information, which is a portion other than the missing portion that has been removed by the missing process in the shape of the learning target, and updates the interpolated non-missing information as the shape of the learning target; A component analysis process that generates a component group related to the learning target based on the shape of the learning target interpolated by the interpolation process by component analysis that generates statistical components by multivariate analysis; An action learning process that learns the action of the learning target and generates an action classification model that classifies the action of the learning target based on the component group related to the learning target generated by the component analysis process, the action of the learning target, and the missing position information; A learning device characterized by executing the above.
9. The learning device according to claim 8, wherein the processor executes a calculation process of calculating the angles of a plurality of vertices constituting the shape of the learning target based on the shape of the learning target, and in the component analysis process, the processor generates a component group related to the learning target based on the shape of the learning target and the angles of the vertices of the learning target calculated by the calculation process. A learning device characterized by the above.
10. The learning device according to claim 8, wherein the processor executes a calculation process of calculating the movement amount of the learning target based on a plurality of shapes of the learning target at different times, and in the component analysis process, the processor generates a component group related to the learning target based on the shape of the learning target and the movement amount of the learning target calculated by the calculation process. A learning device characterized by the above.
11. The learning device according to claim 8, wherein the processor executes a first normalization process of normalizing the size of the shape of the learning target, and in the component analysis process, the processor generates a component group related to the learning target based on the shape of the learning target after the first normalization by the first normalization process. A learning device characterized by the above.
12. The learning device according to claim 9, The processor executes a second normalization process for normalizing the value range that the shape and vertex angles of the learning target can take. In the component analysis process, the processor generates a component group related to the learning target based on the shape and vertex angles of the learning target after the second normalization by the second normalization process. A learning device characterized by the above.
13. The learning device according to claim 8, In the defect process, the processor causes a specific part of the shape of the learning target to be defective. A learning device characterized by the above.
14. The learning device according to claim 8, In the acquisition process, the processor acquires environmental information when acquiring the shape of the learning target. In the behavior learning process, the processor learns the behavior of the learning target based on the component group related to the learning target, the behavior of the learning target, the defect position information, and the environmental information, and generates a behavior classification model for classifying the behavior of the learning target. A learning device characterized by the above.
15. A behavior recognition method executed by a behavior recognition device having a processor that executes a program and a storage device that stores the program, Accessible to a behavior classification model learned using a component group related to the learning target obtained from the shape of the learning target by component analysis that generates statistical components by multivariate analysis and the behavior of the learning target, The processor, A detection process for detecting the shape of the recognition target from the analysis target data, A defect position information generation process for generating defect position information indicating the position of the defective part among the shapes of the recognition target detected by the detection process, An interpolation process for interpolating the defective part from non-defective information that is a part other than the defective part among the shapes of the recognition target including the defective part, and updating the non-defective information after interpolation as the shape of the recognition target, A component analysis process for generating a component group related to the recognition target having the same number as the component group related to the learning target based on the shape of the recognition target interpolated by the interpolation process by the component analysis, A behavior recognition process for outputting a recognition result indicating the behavior of the recognition target by inputting the component group related to the recognition target generated by the component analysis process and the defect position information to the behavior classification model. A behavior recognition method characterized by executing the above.
Citation Information
Patent Citations
Image retrieving apparatus, image retrieving method, and setting screen used therefor
JP2019091138A
Information system and program
JP2021179865A
Behavior recognition apparatus, learning apparatus, and behavior recognition method
JP2022043974A