Information Processing Systems

The system addresses misrecognition in video-based behavior estimation by generating features from both effective and ineffective video sections, ensuring accurate recognition through differentiated scoring.

JP7754289B2Active Publication Date: 2025-10-15NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024511022
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-10-15
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

Existing methods for estimating movement sections in video data can lead to misrecognition due to the inclusion of data not assumed by the movement estimation model, causing inaccuracies in behavior recognition.

Method used

An information processing system that generates features from both effective and ineffective sections of video data, using a learning model to distinguish and assign appropriate scores to these sections, thereby improving recognition accuracy.

Benefits of technology

The system suppresses erroneous recognition by assigning high scores to effective sections and low scores to ineffective sections, enhancing the accuracy of behavior recognition even during transitions or difficult-to-determine data inputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007754289000001
    Figure 0007754289000001
  • Figure 0007754289000002
    Figure 0007754289000002
  • Figure 0007754289000003
    Figure 0007754289000003
Patent Text Reader

Abstract

This information processing device 100 comprises: a feature quantity generation unit for generating a first feature quantity based on video data of a first section, which is a section in one part, and a second feature quantity based on video data of a second section, which is a section other than the first section, from section video data divided into prescribed time sections; and a training unit that, when generating a learning model for outputting an action that corresponds to a feature quantity based on video data in response to input of the feature quantity, trains the learning model for an action that corresponds to the first feature quantity and for an action that corresponds to the second feature quantity.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] Patent Document 1 describes a method for extracting feature quantities of the behavior of a moving object from a video consisting of spatiotemporal information and estimating the behavior. Specifically, Patent Document 1 describes that a motion section from the start to the end of a motion in the video is estimated, and the behavior is estimated based on feature quantities of the video of the motion section. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-179728 [Non-patent literature]

[0004] [Non-Patent Document 1] Sijie Yan, Yuanjun Xiong, Dahua Lin, “Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition”, In AAAI, 2018. [Non-patent document 2] Max Jaderberg, Karen Simonyan, Andrew Zisserman, Koray Kavukcuoglu, “Spatial Transformer Networks”, In NIPS 2015. Summary of the Invention [Problem to be solved by the invention]

[0005] However, the technology described in Patent Document 1 above has a problem in that misrecognition may occur if the estimation of the movement section is inaccurate. For example, if the estimated movement section includes a transition from one movement to another, data not assumed by the movement estimation model is mixed in, making movement estimation difficult and causing misrecognition.

[0006] Therefore, an object of the present invention is to provide an information processing system that can solve the above-mentioned problem that misrecognition may occur when recognizing behavior from video. [Means for solving the problem]

[0007] An information processing system according to one embodiment of the present invention includes: a feature generating unit that generates a first feature based on video data of a first section that is a part of section video data, which is video data divided into predetermined time sections, and a second feature based on video data of a second section that is a section other than the first section; a learning unit that learns an action corresponding to the first feature amount and an action corresponding to the second feature amount when generating a learning model that outputs an action corresponding to the feature amount in response to an input of the feature amount based on video data; Equipped with The structure is as follows.

[0008] Furthermore, an information processing method according to one aspect of the present invention includes: generating a first feature amount that is a feature amount based on video data of a first section that is a part of section video data, which is video data divided into predetermined time sections, and a second feature amount that is a feature amount based on video data of a second section that is a section other than the first section; When generating a learning model that outputs an action corresponding to a feature amount in response to an input of the feature amount based on video data, the learning model learns an action corresponding to the first feature amount and an action corresponding to the second feature amount. The structure is as follows.

[0009] Furthermore, a program according to one aspect of the present invention includes: In the information processing device, generating a first feature amount that is a feature amount based on video data of a first section that is a part of section video data, which is video data divided into predetermined time sections, and a second feature amount that is a feature amount based on video data of a second section that is a section other than the first section; When generating a learning model that outputs an action corresponding to a feature amount in response to an input of the feature amount based on video data, the learning model learns an action corresponding to the first feature amount and an action corresponding to the second feature amount. Execute the process, The structure is as follows. [Effects of the Invention]

[0010] With the above-described configuration, the present invention can suppress erroneous recognition when recognizing actions from video. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram showing the overall configuration of a behavior recognition system according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the configuration of the learning device disclosed in FIG. [Figure 3] FIG. 2 is a block diagram showing the configuration of the estimation device disclosed in FIG. 1. [Figure 4] FIG. 2 is a diagram illustrating a process performed by the learning device disclosed in FIG. [Figure 5] FIG. 2 is a diagram illustrating a process performed by the learning device disclosed in FIG. [Figure 6] FIG. 2 is a diagram illustrating a process performed by the learning device disclosed in FIG. [Figure 7] 2 is a flowchart showing the operation of the learning device disclosed in FIG. [Figure 8] 2 is a flowchart showing the estimation operation disclosed in FIG. 1. [Figure 9] FIG. 10 is a block diagram showing the configuration of an estimation device according to a second embodiment of the present invention. [Figure 10] 10 is a flowchart showing the operation of the estimation device according to the second embodiment of the present invention. [Figure 11] FIG. 10 is a block diagram showing the hardware configuration of an information processing system according to a third embodiment of the present invention. [Figure 12] FIG. 10 is a block diagram showing the configuration of an information processing system according to a third embodiment of the present invention. [Figure 13] 10 is a flowchart showing the operation of the information processing system according to the third embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] <Embodiment 1> A first embodiment of the present invention will be described with reference to Fig. 1 to Fig. 8. Fig. 1 to Fig. 3 are diagrams for explaining the configuration of a behavior recognition system, and Fig. 4 to Fig. 8 are diagrams for explaining the processing operation of the behavior recognition system.

[0013] [composition] The behavior recognition system 1 of the present invention generates a behavior estimation model through machine learning to recognize human behavior from video data, and then uses the generated behavior estimation model to recognize human behavior from new video data. For example, the behavior recognition system 1 can be used for safety management at construction sites to recognize whether or not a worker has performed a safety confirmation action, such as pointing and checking. Specifically, the behavior recognition system 1 can record in a database, from video footage captured by a surveillance camera or the like installed at the construction site, when, where, and how many times a worker performed a safety confirmation action, such as pointing and checking, and issue an alert to sites where safety confirmation actions are not being performed. The behavior recognition system 1 can also be used for labor management at construction sites, and by recording which tasks a worker captured in video performed, how many times, and for what duration, it can be determined whether the tasks are being performed as expected. In this embodiment, the behavior recognition system 1 will be described using examples of behaviors such as pointing, walking, and crouching by a person as recognized behaviors, but any behavior may be recognized, and the system may be used to recognize the behavior of any object, not just a person.

[0014] As shown in FIG. 1, the behavior recognition system 1 includes a learning device 10, a storage device 20, and an estimation device 30. The learning device 10 is a device that learns a learning model used to estimate behavioral information from time-series data (also referred to as learning data) based on time-series data used to learn the learning model. The storage device 20 is a device that can refer to and write data, and is a device that stores the learning data, parameters of the deep learning model, and the like. The estimation device 30 is a device that, when input data is input from an external device, configures an output (estimation) unit by referring to the learned parameters stored in the storage device 20, and generates information about the behavior of the target to be estimated. Each device will be described in detail below.

[0015] The storage device 20 is configured with one or more information processing devices each including a calculation device and a storage device. The storage device 20 has a learning data storage unit 21 and a parameter storage unit 22, as shown in FIG.

[0016] The training data storage unit 21 is a device that stores training data for performing the training process of the training device 10. The training data stored in the training data storage unit 21 will now be described with reference to FIG. 4. As shown by reference character D1 in FIG. 4, the training data is video data consisting of multiple frames that are consecutive in time series. The training data is generated as time-series clips, which are section video data divided into predetermined time segments by cutting out the data using a window of a predetermined width Sw. Then, as shown by reference character D2 in FIG. 4, frames are cut out while sliding the window at a sliding interval St, and time-series clips are sequentially generated and used as training data. Here, this method is referred to as the sliding window method. As an example, the video data has a frame rate of 60 FPS, and time-series clips are sequentially created by sliding the window at a width Sw=120 and a sliding interval St=1. As will be described later, inference data, which is video data input to the estimation device 30 from an external device, has a similar configuration.

[0017] Furthermore, the learning data is stored in association with correct information, which is behavior information (correct behavior) that should be estimated from the learning data. Here, the correct information includes identification information of the correct behavior. For example, in the case of time-series data of skeletal information extracted from a sequence of images showing a person walking, the correct information associated with the target learning data includes identification information indicating that the person is walking.

[0018] The parameter storage unit 22 is a device that stores parameters obtained by learning the learning model. Here, the learning model may be a learning model based on a neural network, or may be another type of learning model such as a support vector machine, or a combination of these. For example, if the learning model is a neural network such as a convolutional neural network, the parameters include the layer structure, the neuron structure of each layer, the number and filter size of filters in each layer, and the weight of each element of each filter. Before learning is performed, the parameter storage unit 22 stores initial values ​​of parameters to be applied to the learning model, and the parameters are updated each time learning is performed by the learning device 10, as described below.

[0019] The learning device 10 is composed of one or more information processing devices each equipped with a calculation device and a storage device. As shown in FIG. 2, the learning device 10 includes a feature extraction unit 11, an action interval detection unit 12, an intra-interval feature extraction unit 13, an extra-interval feature extraction unit 14, a classification unit 15, and a learning unit 16. The functions of the feature extraction unit 11, the action interval detection unit 12, the intra-interval feature extraction unit 13, the extra-interval feature extraction unit 14, the classification unit 15, and the learning unit 16 can be realized by the calculation device executing a program for realizing each function stored in the storage device. Each component will be described in detail below.

[0020] The feature extraction unit 11 (feature generation unit) acquires training data with the above-described window width from the training data storage unit 21 and converts the acquired training data with the window width into a feature F. Here, the feature F has a width in the time direction, as shown in an example in FIG. 6. The feature F may be, for example, three-dimensional data in the time direction, the direction of a person's skeleton points, and the dimension direction of vector amounts calculated at each time and position, as calculated by a neural network as described in Non-Patent Document 1, or it may be two-dimensional data in the dimension direction of time-direction features compressed by taking the maximum value or average value in the direction of the skeleton points. Furthermore, the time direction may be per frame, or may be compressed by convolution processing of a neural network.

[0021] For example, when the feature F has a sliding window width Sw=120, 18 skeleton points, and a vector quantity calculated at each time and position of 256 dimensions, and is compressed to half its size in the time direction by convolution processing, it becomes three-dimensional data of 60 × 18 × 256. In this case, the feature extraction unit 11 configures a feature extractor by applying parameters stored in the parameter storage unit 22 to a learning model that is trained to output the feature F from the input learning data. Then, the feature extraction unit 11 supplies the feature F obtained by inputting the learning data to the feature extractor to the intra-interval feature extraction unit 13 and the extra-interval feature extraction unit 14, respectively.

[0022] The behavior interval detection unit 12 (interval detection unit) acquires training data from the training data storage unit 21, detects intervals (first intervals) that are important for estimating behavioral information from the acquired training data, and outputs interval information S. An important interval for estimating behavioral information here refers to an interval that is useful as a criterion for estimating behavior and that later reduces the error when compared with correct answer information linked to the training data. That is, the behavior interval detection unit 12 outputs, as interval information S, an activity interval corresponding to the correct answer information of the training data. At this time, the behavior interval detection unit 12 configures an behavior interval detector by applying parameters stored in the parameter storage unit 22 to a learning model trained to output the interval information S from the input training data. Here, the behavior interval detector may estimate the behavior interval of the recognition target in the time direction of the feature F using a neural network that estimates transformation parameters of the feature, as described in Non-Patent Document 2, or may use an interval detector that prepares correct answer values ​​for the intervals and directly learns interval detection.

[0023] As an example, when the behavior interval detection unit 12 determines that a pointing behavior is the behavior interval of the recognition target, it outputs the interval during which such behavior is performed as interval information S through the above-described learning, and as a result, it detects intervals during which other behaviors are performed as outside the behavior interval (second interval). For example, in a window W such as that shown by reference symbol D4 in FIG. 5, the frame interval indicated in gray is detected as the behavior interval of the recognition target, and the frame interval indicated by reference symbol Da is detected as outside the behavior interval. Then, the behavior interval detection unit 12 supplies the obtained interval information S to the intra-interval feature extraction unit 13 and the extra-interval feature extraction unit 14, respectively.

[0024] The intra-section feature extraction unit 13 (feature generation unit) and the out-of-section feature extraction unit 14 (feature generation unit) use the feature F supplied from the feature extraction unit 11 and the section information S supplied from the action section detection unit 12 to generate an intra-section feature F1 and an out-of-section feature F2, respectively. Specifically, as shown in FIG. 6, the intra-section feature extraction unit 13 extracts a section corresponding to the section information S from the feature F as described in Non-Patent Document 2, and applies resizing or warping processes to adjust the time direction so that it is always a constant size regardless of the size of the section information S, thereby generating an intra-section feature F1 (first feature). Like the feature F, the intra-section feature F1 may be three-dimensional data in the time direction, the skeleton position direction, and the dimension direction of a vector quantity calculated at each time and each position, or it may be two-dimensional data in the time and dimension directions. The intra-section feature extraction unit 13 supplies the generated intra-section feature F1 to the recognition unit 15.

[0025] 6, the outside-section feature extraction unit 14 extracts features corresponding to the section corresponding to the section information S from the features F as described above, and then generates an outside-section feature F2 (second feature) from the features of the section not extracted, that is, the section detected as being outside the action section. At this time, if there are multiple sections not extracted that are separated in the time direction as shown in FIG. 5, the outside-section feature extraction unit 14 connects these in the time direction to generate one outside-section feature F2. Then, the outside-section feature extraction unit 14 supplies the generated outside-section feature F2 to the identification unit 15.

[0026] Here, in the above example, a case has been illustrated in which a feature F is generated from training data of a window width and then an intra-interval feature F1 and an extra-interval feature F2 are generated, but the method of generating the intra-interval feature F1 and the extra-interval feature F2 is not limited to the above method. For example, the intra-interval feature extraction unit 13 may generate the intra-interval feature F1 from a portion of the training data that corresponds to the interval information S, and the extra-interval feature extraction unit 14 may generate the extra-interval feature F2 from a portion of the training data that does not correspond to the interval information S.

[0027] The classification unit 15 generates information about the target behavior based on the intra-interval feature F1 supplied from the intra-interval feature extraction unit 13 and the extra-interval feature F2 supplied from the extra-interval feature extraction unit 14. The classification unit 15 configures a behavior information output unit by applying parameters stored in the parameter storage unit 22 to a learning model trained to output intra-interval behavior information If1 and extra-interval behavior information If2 from the input intra-interval feature F1 and extra-interval feature F2, respectively. The intra-interval behavior information If1 and extra-interval behavior information If2 are, for example, behavior identification information corresponding to the target learning data or estimated behavior score values, and are vectors with the same number of dimensions as the number of defined behavior categories. If the intra-interval feature F1 and extra-interval feature F2 are three-dimensional data, they are collapsed to two dimensions by taking the maximum or average value in the skeleton point direction, and classification processing is performed for each dimension in the time direction. Then, averaging processing is performed in the time direction. Therefore, the output of the classification unit 15 is a vector quantity with the same number of dimensions as the number of behavior categories to be recognized. The identification unit 15 supplies the learning unit 16 with in-interval behavior information If1 and out-interval behavior information If2 obtained by inputting the in-interval feature amount F1 and the out-interval feature amount F2 to the behavior information output device, respectively.

[0028] The learning unit 16 acquires correct answer information corresponding to the training data input to the feature extraction unit 11 from the training data storage unit 21. The learning unit 16 then trains the feature extraction unit 11, the activity section detection unit 12, and the classification unit 15 based on the acquired correct answer information and the in-section activity information If1 and the out-of-section activity information If2 supplied from the classification unit 15. At this time, the learning unit 16 calculates a loss value L1 calculated from the error between the activity information indicated by the in-section activity information If1 and the correct answer information, and a loss value L2 calculated from the out-of-section activity information If2, calculates a loss from these values, and updates each parameter of the feature extraction unit 11, the activity section detection unit 12, and the classification unit 15 based on this loss. At this time, the loss value L1 may be calculated using any loss function used in machine learning, such as softmax cross-entropy error or mean square error. The loss value L2 is calculated so that the value is uniform across all activity categories. The loss value L2 may be the average of the negative logarithms of the average values ​​of the softmax values ​​of each category of the out-of-interval action information If2, or may be constrained to output 0 for all action categories. For example, if the actions to be recognized are pointing, walking, and crouching, and the correct action in the input learning data is pointing, the error L1 will be small if the value of the dimension corresponding to the pointing action in the vector quantity of the in-interval action information If1 is largest, and the error L2 will be small as the values ​​of all dimensions of the vector quantity of the out-of-interval action information If2 are more uniform. The learning unit 16 determines each parameter so as to minimize these losses.

[0029] To prevent the results of the section information S from becoming too large or too small, a penalty may be imposed on results detected that exceed a certain threshold and added to the loss. For example, if the magnitude of the section information S in the time direction of the data is 0.9 (1 is the entire data), the threshold is set to 0.7, and for inference results that exceed 0.7, the squared value obtained by subtracting the threshold from the inference result is added to the loss. For small results, the same processing is performed as for results below the threshold. The algorithm for determining the above parameters to minimize the loss may be any learning algorithm used in machine learning, such as gradient descent or backpropagation. The learning unit 16 stores the determined parameters of the feature extraction unit 11, action section detection unit 12, and identification unit 15 in the parameter storage unit 22.

[0030] As described above, the learning device 10 has a function of detecting intervals within time-series data that are useful for activity estimation using the activity interval detection unit 12. In the learning process performed by the learning device 10, the intra-interval feature extraction unit 13 and the extra-interval feature extraction unit 14 extract the intra-interval feature F1 and the extra-interval feature F2 from the feature F output by the feature extraction unit 11. After that, when the intra-interval feature F1 is passed through the classification unit 15, the learning device 10 progresses learning so that the score value corresponding to the correct activity class is maximized. Furthermore, when the extra-interval feature F2 is passed through the classification unit 15, the learning device 10 progresses learning so that the score values ​​of all activity classes are uniform and no class stands out. In this way, by simultaneously handling information not only within intervals but also outside intervals, the learning device 10 can ensure the operation of the activity estimation model so that, if sufficient data is available, the score values ​​of all activity classes do not stand out significantly when data from intervals with low importance is input during activity estimation. For example, when a person's pointing, walking, and crouching behaviors are to be recognized, when transitioning from walking to pointing, the walking speed must be slowed down to prepare for the pointing behavior. If a candidate behavior section approaches a section where behaviors are transitioning, previous models have not taken this section into account, resulting in a false positive, such as an output result in which the score value of a certain behavior class is out of sync. In contrast, the present invention learns to detect sections useful for behavior estimation, and continues learning so that score values ​​outside the section are all uniform, thereby reducing false positives.

[0031] Here, the action interval detection unit 12 performs interval detection based on the learning data received from the learning data storage unit 21, but it may also perform interval detection using the feature amount F output by the feature extraction unit 11 as input.

[0032] Next, the configuration of the estimation device 30 will be described. The estimation device 30 is configured with one or more information processing devices each including a calculation device and a storage device. As shown in FIG. 3, the estimation device 30 includes a feature extraction unit 31, a classification unit 35, and an output unit 36. The functions of the feature extraction unit 31, the classification unit 35, and the output unit 36 ​​can be realized by the calculation device executing a program for realizing each function stored in the storage device. Each component will be described in detail below.

[0033] The feature extraction unit 31 (target feature generation unit) acquires time-series data input from an external device and converts the acquired time-series data into feature F (target feature). Here, the time-series data input from the external device is inference data to be used for behavior identification, and is video data (target video data) similar to the learning data described above. In other words, as shown in FIG. 4, the inference data is video data consisting of multiple frames (image sequences) that are consecutive in time series, and is a time-series clip cut out within a window of a predetermined width Sw.

[0034] However, the inference data input to the estimation device 30 may also be data extracted from an image sequence, such as skeletal information. Furthermore, the external device that inputs the inference data may be a camera if an image sequence is used as input, or may be a device that stores the generated image sequence or information extracted from the image sequence if the input is used.

[0035] The feature extraction unit 31 then configures a feature extractor based on the parameters obtained by the learning process by the learning device 10, by referring to the parameters stored in the parameter storage unit 22. The feature extraction unit 31 then supplies the feature amount F obtained by inputting the inference data to the feature extractor to the identification unit 35.

[0036] The identification unit 35 generates behavioral information Ifa from the feature F supplied from the feature extraction unit 31. The identification unit 35 performs estimation for each dimension in the time direction of the feature F. At this time, the identification unit 35 configures a behavioral information output device by referring to the parameters stored in the parameter storage unit 22. Then, the identification unit 35 supplies the behavioral information Ifa obtained by inputting the feature F to the behavioral information output device to the output unit 36.

[0037] The output unit 36 ​​outputs the identification information of the behavior to be extracted to an external device based on the behavior information Ifa. Since the behavior information Ifa is output for each dimension in the time direction, the score values ​​are compressed in the time direction by averaging or summing them, and a vector quantity with the same number of dimensions as the number of behavior classes to be recognized is output. This vector quantity becomes the identification information of the behavior at the center time of this window. Using this behavior identification information, when input time-series data divided at regular window intervals using a sliding window method is used as input data, the output unit 36 ​​outputs the identification information of the behavior and its start and end points by defining the time when the score values ​​of the behavior identification information based on the behavior information Ifa in each window are arranged in time series, as the start point and the time when the score exceeds a certain threshold as the end point.

[0038] Note that the estimation process executed by the estimation device 30 described above does not perform interval detection, which is performed in the learning process executed by the learning device 10. This is because the behavior estimation model reacts and the score value increases only in characteristic behavioral intervals during the learning process, and therefore, when scanning time-series data using a sliding window method to arrange score values ​​in time series and determine start and end points by threshold judgment, it is not necessary to perform interval detection within the model.

[0039] [Operation] Next, the operation of the above-mentioned behavior recognition system 1 will be described mainly with reference to the flowcharts of Figures 7 and 8. First, the operation of the learning device 10 in the learning mode will be described with reference to the flowchart of Figure 7.

[0040] The feature extraction unit 11 acquires training data from the training data storage unit 21 (step S1). At this time, the feature extraction unit 11 acquires training data that has not yet been used for training (i.e., not acquired in step S1) from the training data stored in the training data storage unit 21. Then, the feature extraction unit 11 refers to the parameters stored in the parameter storage unit 22 to configure a feature extractor, thereby generating a feature F from the training data acquired in step S1 (step S2).

[0041] Next, the behavior section detection unit 12 generates the section information S from the training data by configuring a behavior section detector with reference to the parameters stored in the parameter storage unit 22 (step S3). Then, the intra-section feature extraction unit 13 and the extra-section feature extraction unit 14 generate the intra-section feature F1 and the extra-section feature F2, respectively, from the feature F and the section information S (step S4). Next, the identification unit 15 configures a behavior information output unit with reference to the parameters stored in the parameter storage unit 22, thereby generating the behavior information If1 from the intra-section feature F1 generated by the intra-section feature extraction unit 13 and the behavior information If2 from the extra-section feature F2 generated by the extra-section feature extraction unit 14 (step S5).

[0042] Then, the learning unit 16 calculates a loss based on the in-interval action information If1 and the out-interval action information If2 generated by the classification unit 15 and the correct answer information associated with the target learning data and stored in the learning data storage unit 21 (step S6). Furthermore, the learning unit 16 updates the parameters used by the feature extraction unit 11, the action interval detection unit 12, and the classification unit 15 based on the loss calculated in step S6 (step S7). At this time, the learning unit 16 stores the parameters used by the feature extraction unit 11, the action interval detection unit 12, and the classification unit 15 in the parameter storage unit 22.

[0043] Next, the learning device 10 determines whether a learning termination condition is satisfied (step S8). At this time, the learning device 10 may determine the learning termination condition by, for example, determining whether a predetermined number of loops has been reached, determining whether learning has been performed on a predetermined number of learning data, determining whether the loss has fallen below a predetermined threshold, or determining whether a change in loss has fallen below a predetermined threshold. Note that step S8 may be a combination of the above examples, or may be any other determination method. Then, if the learning termination condition is satisfied (Yes in step S8), the learning device 10 ends the flowchart. On the other hand, if the learning termination condition is not satisfied (No in step S8), the learning device 10 returns to step S1. At this time, the learning device 10 retrieves unused learning data from the learning data storage unit 21 in step S1 and performs the processes from step S2 onwards.

[0044] In this way, the learning device 10 learns a learning model used to estimate behavior from the learning data, and records the parameters of the learned learning model in the storage device 20.

[0045] Next, the operation of the estimating device 30 in the inference mode will be described with reference to the flowchart in Fig. 8. The estimating device 30 repeatedly executes the processing of the flowchart shown in Fig. 8 every time input data is input to the estimating device 30. As described above, it is assumed that the input data is video data, which is time-series data, scanned in a sliding window manner.

[0046] The feature extraction unit 31 acquires input data supplied from an external device (step S11). Then, the feature extraction unit 31 configures a feature extractor by referring to the parameters stored in the parameter storage unit 22, thereby generating a feature amount F from the input data acquired in step S11 (step S12). Next, the identification unit 35 configures a behavior information output unit by referring to the parameters stored in the parameter storage unit 22, thereby generating behavior information Ifa from the feature amount F (step S3). Then, the output unit 36 ​​outputs behavior identification information and its start and end points to the external device based on the behavior information Ifa generated by the identification unit 35 (step S14).

[0047] In this way, the estimation device 30 references the stored learned parameters, constructs an inference model, uses this model to infer behavior from the video data that is the inference target, and outputs the inference results.

[0048] As described above, the behavior recognition system of this embodiment simultaneously learns to detect behavior sections within video data when learning the behavior recognition model, distinguishes between sections within the video data that are effective for behavior recognition and sections that are not effective, and performs behavior recognition with a high score for effective sections and a low score for ineffective sections. Therefore, even when data that is difficult to determine, such as a change in behavior, is input within a video section that is a candidate for behavior recognition, behavior information for such data can be output with a low reliability, thereby suppressing false detection and enabling more accurate behavior recognition.

[0049] <Embodiment 2> Next, a second embodiment of the present invention will be described with reference to Fig. 9 and Fig. 10. Fig. 9 is a diagram for explaining the configuration of an estimation device, and Fig. 10 is a diagram for explaining the operation of the estimation device.

[0050] [composition] The behavior recognition system 1 of the present invention differs from the above-described first embodiment in the configuration of the estimation device 30. The following mainly describes the configuration that differs from the first embodiment.

[0051] 9, the estimation device 30 in this embodiment includes a feature extraction unit 31, an action interval detection unit 32, an intra-interval feature extraction unit 33, a classification unit 35, and an output unit 36. Note that the functions of the feature extraction unit 31, the action interval detection unit 32, the intra-interval feature extraction unit 33, the classification unit 35, and the output unit 36 ​​can be realized by a computing device executing a program for realizing each function stored in a storage device.

[0052] The action interval detection unit 32 (target interval detection unit) and the intra-interval feature extraction unit 33 (target feature generation unit) in this embodiment are components added to the estimation device 30 of embodiment 1, and have the same functions as the action interval detection unit 12 and the intra-interval feature extraction unit 13 provided in the learning device 10 of embodiment 1. In other words, the action interval detection unit 32 generates interval information S from the inference data in the same manner as described above, and the intra-interval feature extraction unit 13 generates an intra-interval feature F1 from the feature F generated from the inference data and the interval information S.

[0053] The identification unit 35 in this embodiment generates intra-section action information Ifa based on the intra-section feature F1, and the output unit 36 ​​outputs identification information of the action to be extracted to an external device based on the intra-section action information Ifa. As described above, in this embodiment, compared to the first embodiment, action sections are detected within the inference data, so the start and end of an action can be detected without scanning the input time-series data using a sliding window method, and identification information of the action in that section can be obtained using the output of the identification unit 35. For example, the action class corresponding to the dimension with the largest score value in the intra-section action information Ifa is output.

[0054] When the input time series data is scanned using a sliding window method, the behavior section detection unit 32 detects behavior sections and checks whether the target behavior is included in the data. If the length of the section exceeds a threshold, the identification information of the behavior to be extracted is output based on the in-section behavior information Ifa. Since the in-section behavior information Ifa is output for each dimension in the time direction within the window, the score values ​​are compressed in the time direction by averaging or summing them, and a vector quantity with the same number of dimensions as the number of behavior classes to be recognized is output. This vector quantity becomes the identification information of the behavior at the center time of the window. If the length of the section does not exceed the threshold, a fixed value is prepared (for example, all values ​​are set to 0) for the number of behavior classes to be recognized, and this is used as the identification information of the behavior. Using this identification information of the behavior, when the input time series data is divided into fixed window intervals using a sliding window method, the score values ​​of the identification information of the behavior based on the behavior information Ifa for each window are arranged in chronological order, and the time when a certain threshold is exceeded is defined as the start point, and the time when the threshold is exceeded is defined as the end point, and the identification information of the behavior and its start and end points are output.

[0055] Here, the action interval detection unit 32 performs interval detection based on input data input from an external device, but it may also input the feature value F output by the feature extraction unit 31 and perform interval detection based on this.

[0056] [Operation] Next, the operation of the estimation device 30 in this embodiment will be described with reference to the flowchart in Fig. 10. The estimation device 30 repeatedly executes the process of the flowchart shown in Fig. 10 every time video data to be estimated, which is input data, is input to the estimation device 30. The input data may be time-series data input as is, or may be data scanned using a sliding window method and input.

[0057] The feature extraction unit 31 acquires input data supplied from an external device (step S21). Then, the feature extraction unit 31 generates a feature quantity F from the input data acquired in step S21 by configuring a feature extractor with reference to the parameters stored in the parameter storage unit 22 (step S22).

[0058] Next, the action section detection unit 32 generates section information S from the training data by configuring an action section detector with reference to the parameters stored in the parameter storage unit 22 (step S23). Then, the intra-section feature extraction unit 33 generates an intra-section feature F1 from the feature F and the section information S (step S24).

[0059] Next, the identification unit 35 configures the behavior information output device with reference to the parameters stored in the parameter storage unit 22, thereby generating behavior information Ifa from the intra-section feature F1 generated by the intra-section feature extraction unit 33 (step S25). Then, the output unit 36 ​​outputs the behavior identification information, start point, and end point to an external device based on the behavior information Ifa output by the identification unit 35 (step S26).

[0060] As described above, in this embodiment, the estimation device 30 includes the behavioral interval detection unit 32, and therefore, when the input data is the entire time-series data, interval detection can be performed without threshold processing. When scanning the input time-series data using the sliding window method, it is possible to determine whether the target behavior is included within the window based on the length of the interval. Even if the interval detection is off and data such as behavioral transitions is mixed in within the interval, false detection is unlikely to occur because the system is trained to produce uniform scores during training.

[0061] <Embodiment 3> Next, a third embodiment of the present invention will be described with reference to Fig. 11 to Fig. 13. Fig. 11 to Fig. 12 are block diagrams showing the configuration of an information processing system in the third embodiment, and Fig. 13 is a flowchart showing the operation of the information processing system. Note that this embodiment shows an outline of the configuration of the information processing system and information processing method described in the above embodiments.

[0062] First, the hardware configuration of the information processing system 100 in this embodiment will be described with reference to Fig. 11. The information processing system 100 is configured with a general information processing device, and is equipped with the following hardware configuration, for example. ·CPU(Central Processing Unit)101(Arithmetic unit) ROM (Read Only Memory) 102 (storage device) RAM (Random Access Memory) 103 (storage device) Programs 104 loaded into RAM 103 A storage device 105 for storing a group of programs 104 A drive device 106 that reads and writes from a storage medium 110 external to the information processing device A communication interface 107 that connects to a communication network 111 outside the information processing device Input / output interface 108 for inputting and outputting data Bus 109 connecting each component

[0063] The information processing system 100 can be equipped with a feature generation unit 121 and a learning unit 122 shown in FIG. 12 by having the CPU 101 acquire and execute the program group 104. The program group 104 is stored in advance in, for example, the storage device 105 or the ROM 102, and is loaded into the RAM 103 and executed by the CPU 101 as needed. The program group 104 may be supplied to the CPU 101 via the communication network 111, or may be stored in advance in the storage medium 110, with the drive device 106 reading out the programs and supplying them to the CPU 101. However, the feature generation unit 121 and the learning unit 122 described above may be constructed using electronic circuits dedicated to realizing such means.

[0064] 11 shows an example of the hardware configuration of the information processing device that is the information processing system 100, and the hardware configuration of the information processing device is not limited to the above-described case. For example, the information processing device may be configured with a part of the above-described configuration, such as not including the drive device 106.

[0065] The information processing system 100 then executes the information processing method shown in the flowchart of FIG. 13 using the functions of the feature generating unit 121 and the learning unit 122, which are constructed by the program as described above.

[0066] As shown in FIG. 13, the information processing system 100 A first feature amount is generated based on video data of a first section, which is a part of section video data that is video data divided into predetermined time sections, and a second feature amount is generated based on video data of a second section, which is a section other than the first section (step S101); When generating a learning model that recognizes an action corresponding to a feature amount in response to an input of the feature amount based on video data, the learning model learns an action corresponding to the first feature amount and an action corresponding to the second feature amount (step S102). The following process is executed.

[0067] With the above configuration, the present invention generates feature amounts for sections in video data that are effective for action recognition and feature amounts for sections that are ineffective, and learns actions corresponding to the feature amounts for the effective sections and actions corresponding to the feature amounts for the ineffective sections. In this case, for example, learning is performed so that correct actions are assigned high scores in effective sections, and so that multiple actions are assigned low scores in ineffective sections. As a result, even when difficult-to-determine data, such as transitions in actions, is input within a video section that is a candidate for action recognition, action information for such data can be output with low reliability, thereby suppressing false detection and enabling more accurate action recognition.

[0068] The above-described program can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can supply the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.

[0069] Although the present invention has been described above with reference to the above-described embodiments, the present invention is not limited to the above-described embodiments. Various modifications that are understandable to those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. Furthermore, at least one or more of the functions of the feature generation unit 121 and the learning unit 122 described above may be executed by an information processing device installed and connected anywhere on a network, that is, may be executed by so-called cloud computing.

[0070] <Additional Notes> A part or all of the above-described embodiments can be described as follows: The following provides an overview of the configurations of the information processing system, information processing method, and program according to the present invention. However, the present invention is not limited to the following configurations. (Appendix 1) a feature generating unit that generates a first feature based on video data of a first section that is a part of section video data, which is video data divided into predetermined time sections, and a second feature based on video data of a second section that is a section other than the first section; a learning unit that learns an action corresponding to the first feature amount and an action corresponding to the second feature amount when generating a learning model that outputs an action corresponding to the feature amount in response to an input of the feature amount based on video data; An information processing system comprising: (Appendix 2) 10. The information processing system of claim 1, the learning unit learns so that a correct action set for the section video data corresponds to the first feature amount, and learns so that a plurality of actions correspond to the second feature amount; Information processing system. (Appendix 3) 10. The information processing system according to claim 2, the learning unit learns so that a plurality of actions correspond equally to the second feature amount; Information processing system. (Appendix 4) 4. The information processing system according to claim 2, the learning unit, when generating a learning model that outputs a correspondence degree of an action corresponding to an input feature based on video data in response to the input feature, learns to increase the correspondence degree of the correct action to the first feature and learns to decrease the correspondence degrees of each of the plurality of actions to the second feature. Information processing system. (Appendix 5) An information processing system according to any one of Supplementary Notes 1 to 4, the feature generation unit generates one of the first feature and one of the second feature so as to have a component in a time direction; Information processing system. (Appendix 6) 6. The information processing system according to claim 5, the feature generation unit generates the first feature so that a size in a time direction of the first feature becomes a preset size. Information processing system. (Appendix 7) 7. The information processing system according to claim 5, when the second section is a plurality of sections separated in the time direction, the feature generation unit generates the second feature based on video data of the second section obtained by connecting the plurality of sections into one. Information processing system. (Appendix 8) An information processing system according to any one of Supplementary Notes 1 to 7, the feature generation unit generates a feature of the section video data based on the section video data, and generates the first feature and the second feature based on the feature, the first section, and the second section. Information processing system. (Appendix 9) An information processing system according to any one of Supplementary Notes 1 to 8, a section detection unit that detects the first section and the second section of the section video data based on the section video data, Information processing system. (Appendix 10) An information processing system according to any one of Supplementary Notes 1 to 9, an object feature generation unit that generates object feature values, which are feature values ​​of the object video data that is a target of behavior classification, based on the object video data; an identification unit that identifies an action of the target video data based on an action output from the learning model by inputting the target feature amount into the learning model; An information processing system comprising: (Appendix 11) 11. The information processing system of claim 10, a target section detection unit that detects the first section of the target video data based on the target video data; the target feature generation unit generates a first target feature, which is a feature corresponding to the first section, based on the target feature and the first section of the target video data; the identification unit inputs the first target feature into the learning model and identifies the behavior of the target video data based on the behavior output from the learning model; Information processing system. (Appendix 12) generating a first feature amount that is a feature amount based on video data of a first section that is a part of section video data, which is video data divided into predetermined time sections, and a second feature amount that is a feature amount based on video data of a second section that is a section other than the first section; When generating a learning model that outputs an action corresponding to a feature amount in response to an input of the feature amount based on video data, the learning model learns an action corresponding to the first feature amount and an action corresponding to the second feature amount. Information processing methods. (Appendix 13) 13. The information processing method according to claim 12, further comprising: learning so that a correct action set for the section video data corresponds to the first feature amount, and learning so that a plurality of actions correspond to the second feature amount; Information processing methods. (Appendix 14) 14. The information processing method according to claim 12 or 13, generating target features that are features of the target video data based on the target video data that is a target of behavior identification; inputting the target feature amount into the learning model to identify the behavior of the target video data based on the behavior output from the learning model; Information processing methods. (Appendix 15) In the information processing device, generating a first feature amount that is a feature amount based on video data of a first section that is a part of section video data, which is video data divided into predetermined time sections, and a second feature amount that is a feature amount based on video data of a second section that is a section other than the first section; When generating a learning model that outputs an action corresponding to a feature amount in response to an input of the feature amount based on video data, the learning model learns an action corresponding to the first feature amount and an action corresponding to the second feature amount. A computer-readable storage medium that stores a program for executing processing. [Explanation of symbols]

[0071] 1. Behavior Recognition System 10 Learning Device 11 Feature Extraction Unit 12 Action section detection unit 13 Interval feature extraction unit 14. Out-of-interval feature extraction unit 15 Identification unit 16 Learning Department 20 Storage device 21 Learning data storage unit 22 Parameter storage section 30 Estimation device 31 Feature Extraction Unit 32 Action section detection unit 33 Interval feature extraction unit 35 Identification unit 36 Output section 100 Information Processing Systems 101 CPU 102 ROM 103 RAM 104 Programs 105 Storage device 106 Drive device 107 Communication Interface 108 Input / Output Interface 109 Bus 110 Storage medium 111 Communication Network 121 Feature Generation Unit 122 Learning Department

Claims

1. a feature generating unit that generates a first feature based on video data of a first section that is a part of section video data, which is video data divided into predetermined time sections, and a second feature based on video data of a second section that is a section other than the first section; a learning unit that learns an action corresponding to the first feature amount and an action corresponding to the second feature amount when generating a learning model that outputs an action corresponding to the feature amount in response to an input of the feature amount based on video data; Equipped with the feature generation unit generates one each of the first feature and the second feature so as to have a component in the time direction, and further, when the second section is a plurality of sections separated in the time direction, generates the second feature based on video data of the second section obtained by connecting the plurality of sections into one. Information processing system.

2. 2. The information processing system according to claim 1, the learning unit learns so that a correct action set for the section video data corresponds to the first feature amount, and learns so that a plurality of actions correspond to the second feature amount; Information processing system.

3. 3. The information processing system according to claim 2, the learning unit learns so that a plurality of actions correspond equally to the second feature amount; Information processing system.

4. 4. The information processing system according to claim 2, the learning unit, when generating a learning model that outputs a correspondence degree of an action corresponding to an input feature based on video data in response to the input feature, learns to increase the correspondence degree of the correct action to the first feature and learns to decrease the correspondence degrees of each of the plurality of actions to the second feature. Information processing system.

5. 5. An information processing system according to claim 1, the feature generation unit generates the first feature so that a size in a time direction of the first feature becomes a preset size. Information processing system.

6. 6. An information processing system according to claim 1, the feature generation unit generates a feature of the section video data based on the section video data, and generates the first feature and the second feature based on the feature, the first section, and the second section. Information processing system.

7. The information processing device When generating a first feature amount that is a feature amount based on video data of a first section that is a part of section video data, which is video data divided into predetermined time sections, and a second feature amount that is a feature amount based on video data of a second section that is a section other than the first section, one first feature amount and one second feature amount are generated so as to have a time direction component, and further, when the second section is a plurality of sections divided in the time direction, one second feature amount is generated based on video data of the second section that is obtained by concatenating the plurality of sections into one; When generating a learning model that outputs an action corresponding to a feature amount in response to an input of the feature amount based on video data, the learning model learns an action corresponding to the first feature amount and an action corresponding to the second feature amount. Information processing methods.

8. In the information processing device, When generating a first feature amount that is a feature amount based on video data of a first section that is a part of section video data, which is video data divided into predetermined time sections, and a second feature amount that is a feature amount based on video data of a second section that is a section other than the first section, one first feature amount and one second feature amount are generated so as to have a time direction component, and further, when the second section is a plurality of sections divided in the time direction, one second feature amount is generated based on video data of the second section that is obtained by concatenating the plurality of sections into one; When generating a learning model that outputs an action corresponding to a feature amount in response to an input of the feature amount based on video data, the learning model learns an action corresponding to the first feature amount and an action corresponding to the second feature amount. A program for executing a process.

Citation Information

Patent Citations

  • State detection method, state detection device, and state detection program

    JP2016158954A

  • Behavior recognition device, learning device, and method and program

    JP2019040465A

  • Data dividing device, data dividing method, and program

    JP2020021421A

  • Video processing device and method thereof

    JP2021179728A