Behavior location identification device, control method, and program

The behavior localization device uses fisheye cameras to generate person clips, extract feature maps, and perform class activation mapping to accurately identify and localize behaviors in fisheye images, addressing the challenges of distortion and complexity in existing systems.

JP7823753B2Active Publication Date: 2026-03-04NEC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-06
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing systems struggle to accurately localize activities in video-captured images, particularly in fisheye views, due to distortion and complexity in identifying behaviors near the optical axis of fisheye lenses.

Method used

A behavior localization device and method that utilizes a fisheye camera to capture top-view images, generates person clips, extracts feature maps, calculates behavior scores, and performs class activation mapping to identify and localize behaviors using 3D convolutional neural networks and class activation techniques.

Benefits of technology

Accurately localizes behaviors in fisheye images by generating person clips, extracting feature maps, and performing class activation mapping, enhancing the precision of behavior identification and localization in surveillance and monitoring applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007823753000004
    Figure 0007823753000004
  • Figure 0007823753000005
    Figure 0007823753000005
  • Figure 0007823753000006
    Figure 0007823753000006
Patent Text Reader

Abstract

The behavior location identification device (2000) acquires a target clip (10) and detects a person (40) from the target clip (10). The behavior location identification device (2000) generates a person clip (60) from the target clip (10) for each person (40) detected from the target clip (10), and extracts a feature map from each of the person clips (60). The behavior location identification device (2000) calculates a behavior score for each predefined behavior class based on the feature map extracted from the person clip (60), calculates a behavior score indicating the reliability of the behavior of the behavior class included in the target clip (10), and performs class activation mapping for each of the person clips (60) to locate each behavior whose behavior class has a behavior score equal to or greater than a threshold.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE This disclosure relates generally to localizing video-captured activities. [Background technology]

[0002] Activity localization is the process of locating an activity captured in an image, i.e., determining which activity is performed at which location in the image. Patent document 1 discloses a system that uses neural networks to detect activities in videos for loss prevention in the retail industry. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] U.S. Patent Publication No. 2021 / 0248377 [Non-patent literature]

[0004] [Non-Patent Document 1] Junnan Li, Jianquan Liu, Yongkang Wong, Shoji Nishimura, Mohan Kankanhalli, "Weakly-Supervised Multi-Person Action Recognition in 360° Videos," [online], February 9, 2020, [Retrieved November 10, 2021], Internet<URL: https: / / arxiv.org / pdf / 2002.03266.pdf> Summary of the Invention [Problem to be solved by the invention]

[0005] It is an object of the present disclosure to provide novel techniques for localizing video-captured activity. [Means for solving the problem]

[0006] The present disclosure provides an apparatus for identifying behavioral locations, the apparatus including at least one memory configured to store instructions and at least one processor, the processor being configured to execute the instructions to acquire target clips, each of which is a sequence of target images, each of which is a fisheye image of one or more people captured from a generally overhead perspective, detect one or more people from the target clips, generate a person clip from the target clip for each person detected from the target clip, each person clip being a sequence of person images that is a partial region of the target image and including the detected person corresponding to the person clip, extract a feature map from each person clip, and calculate, for each predetermined behavioral class, an behavior score indicating a reliability of behavior of the behavioral class included in the target clip based on the feature map extracted from the person clip, the behavioral class being a type of behavior, and localize each behavior of the behavioral class having the behavior score equal to or greater than a threshold by performing class activation mapping on each of the person clips.

[0007] The present disclosure further provides a computer-implemented control method, the control method including the steps of: acquiring target clips, each of which is a sequence of target images, each of which is a fisheye image of one or more people captured from a generally overhead viewpoint; detecting one or more people from the target clips; generating a person clip from the target clip for each of the people detected in the target clips, where each person clip is a sequence of person images that is a partial region of the target image and includes the detected person corresponding to the person clip; extracting a feature map from each of the person clips; and calculating, for each predetermined behavior class, a behavior score indicating a reliability of behavior of the behavior class included in the target clip based on the feature map extracted from the person clip, where the behavior class is a type of behavior; and performing class activation mapping on each of the person clips to locate each behavior of the behavior class having the behavior score equal to or greater than a threshold.

[0008] The present disclosure further provides a non-transitory computer-readable storage medium having a program stored thereon, the program causing a computer to execute the steps of: acquiring a target clip, which is a sequence of target images, each of which is a fisheye image of one or more people captured from a generally overhead viewpoint; detecting one or more people from the target clip; generating a person clip from the target clip for each person detected from the target clip, where each person clip is a sequence of person images that is a partial region of the target image and includes the detected person corresponding to the person clip; extracting a feature map from each person clip; and calculating, for each predetermined behavioral class based on the feature map extracted from the person clip, a behavior score indicating a reliability of behavior of the behavioral class included in the target clip, where the behavioral class is a type of behavior, and performing class activation mapping on each of the person clips to locate each behavior of the behavioral class having the behavior score equal to or greater than a threshold. [Effects of the Invention]

[0009] According to the present disclosure, novel techniques are provided for localizing video-captured activity. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a diagram illustrating an overview of a behavior location identification device according to a first embodiment. [Figure 2] FIG. 10 is a diagram showing an example of a person clip. [Figure 3] FIG. 2 is a block diagram showing an example of a functional configuration of a behavior position identification device. [Figure 4] FIG. 2 is a block diagram showing an example of a hardware configuration of a behavior position identification device. [Figure 5] 10 is a flowchart illustrating an exemplary flow of processing performed by a behavior location identification device. [Figure 6]FIG. 10 is a diagram illustrating a case where a feature map represents spatiotemporal features of a target person and the surrounding area. [Figure 7] FIG. 10 illustrates an example of a method for calculating a behavior score based on a feature map. [Figure 8] FIG. 10 is a diagram illustrating an example of a class activation map. [Figure 9] 10 is a flowchart illustrating an exemplary flow of processing performed by a position specifying unit. [Figure 10] FIG. 10 shows a first example in which a target clip is modified to show the results of behavioral location. [Figure 11] FIG. 10 shows a second example in which a target clip is modified to show the results of behavioral location. [Figure 12] FIG. 10 is a diagram illustrating an overview of a behavior location identification device according to a second embodiment. [Figure 13] FIG. 10 is a block diagram showing an example of the functional configuration of a behavior position identification device according to a second embodiment. [Figure 14] 10 is a flowchart showing a flow of processing performed by a behavior location identification device of the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Embodiments according to the present disclosure will be described below with reference to the drawings. The same elements are assigned the same reference numerals throughout the drawings, and redundant descriptions will be omitted as necessary. Furthermore, unless otherwise specified, predetermined information (e.g., predetermined values ​​or predetermined threshold values) is pre-stored in a storage device accessible by a computer that uses the information.

[0012] Embodiment 1 <Summary> Fig. 1 shows an overview of a behavior location identification device 2000 according to embodiment 1. Note that the overview shown in Fig. 1 shows an example of the operation of the behavior location identification device 2000 in order to make it easier to understand the behavior location identification device 2000, and is not intended to limit or narrow the range of operations that the behavior location identification device 2000 can perform.

[0013] The activity localization device 2000 operates on a target clip 10, which is part of a video 20 and is formed by a sequence of video frames 22. Each video frame 22 in the target clip 10 is called a "target image 12." The video 20 is a sequence of video frames 22 generated by a fisheye camera 30.

[0014] The fisheye camera 30 includes a fisheye lens and is installed as a top-view camera that captures images of the target location from a perspective above. Thus, each video frame 22 (and each target image 12) is a fisheye top-view image of the target location captured from a perspective above.

[0015] The target location can be any location. For example, when the fisheye camera 30 is used as a surveillance camera, the target location is a location to be monitored, such as a facility or its surroundings. The facility to be monitored may be a station, a stadium, or the like.

[0016] It may be preferable that the field of view of the fisheye camera 30 is substantially close to 360 degrees in the horizontal direction. However, the field of view of the fisheye camera 30 does not need to be exactly 360 degrees in the horizontal direction. The fisheye camera 30 may also be installed so that the optical axis of its fisheye lens is approximately parallel to the vertical direction. However, the optical axis of the fisheye lens of the fisheye camera does not need to be exactly parallel to the vertical axis.

[0017] The activity location identification device 2000 detects one or more people 40 from the target clip 10, detects one or more actions taken by the detected people 40, and locates the detected actions (i.e., identifies which type of action occurred in which area of ​​the target clip 10). There may be various activity classes, such as "walking," "drinking," "putting on a jacket," and "operating a phone." Specifically, the activity location identification device 2000 may operate as follows.

[0018] The behavior location identification device 2000 acquires a target clip 10 and detects one or more people 40 from the target clip 10. Then, the behavior location identification device 2000 generates a person clip 60 for each person 40 detected from the target clip 10. Hereinafter, a person corresponding to a person clip 60 will be referred to as a "target person." The person clips 60 of a target person are a sequence of person images 62, each of which includes the target person and is cropped from a corresponding target image 12. The cropping positions (the positions in the target image 12 from which the person images 62 are cropped) and the dimensions of the person images 62 in a single person clip 60 are the same as each other.

[0019] 2 shows an example of a person clip 60. In this example, the target clip 10 includes three people 40-1 to 40-3. Therefore, person clips 60-1 to 60-3 are generated for the people 40-1 to 40-3, respectively.

[0020] The behavior location identification device 2000 extracts a feature map from each person clip 60. The feature map represents the spatiotemporal features of the person clip 60. Then, the behavior location identification device 2000 calculates a behavior score for each predefined behavior class based on the feature map extracted from the person clip 60, and generates a behavior score vector 50. The behavior score of a behavior class represents the reliability of one or more behaviors of the behavior class included in the target clip 10 (in other words, the reliability that one or more people 40 in the target clip 10 are performing the behavior of the behavior class). The behavior score vector 50 is a vector whose elements are the behavior scores of each predefined behavior class.

[0021] Assume that there are three predefined behavioral classes A1, A2, and A3. In this case, the behavior score vector is a three-dimensional vector v=(c1, c2, c3), where c1 represents the behavior score of behavioral class A1 (i.e., the reliability of one or more behaviors in behavioral class A1 included in target clip 10), c2 represents the behavior score of behavioral class A2 (i.e., the reliability of one or more behaviors in behavioral class A2 included in target clip 10), and c3 represents the behavior score of behavioral class A3 (i.e., the reliability of one or more behaviors in behavioral class A3 included in target clip 10).

[0022] The behavior score vector 50 does not indicate what type of behavior is performed in which area of ​​the target clip 10. Based on the feature map obtained from the target clip 10 and the behavior score vector 50, the behavior localization device 2000 localizes the behavior in the target clip 10 (i.e., identifies what type of behavior occurred in which area of ​​the target clip 10). This type of localization is called "behavior localization."

[0023] The behavior location identification device 2000 uses class activation mapping to perform behavior location for the target clip 10. Specifically, for each person clip 60 and for each behavior class detected from the target clip 10, the behavior location identification device 2000 performs class activation mapping to identify which area in the person image 62 of that person clip 60 includes which type of behavior. This allows each of the detected behaviors of the person 40 to be located.

[0024] <Examples of effects> As described above, the behavior location identification device 2000 provides a novel technique for locating behaviors captured on video. Specifically, the behavior location identification device 2000 generates a person clip 60 for each person 40 detected from a target clip 10, extracts a feature map for each person clip 60, and calculates a behavior score vector indicating each behavior score of predefined behavior classes. Then, the behavior location identification device 2000 locates each behavior in the target clip 10 by identifying the behavior class of the behavior included in each person clip 60 using class activation mapping.

[0025] The behavior location identification device 2000 will be described in detail below.

[0026] <Example of functional configuration> 3 is a block diagram showing an example of the functional configuration of the behavior position identification device 2000. The behavior position identification device 2000 includes an acquisition unit 2020, a person clip generation unit 2040, a feature extraction unit 2060, a score calculation unit 2080, and a position identification unit 2100.

[0027] The acquisition unit 2020 acquires a target clip 10. The person clip generation unit 2040 generates a person clip 60 for each person 40 detected in the target clip 10. The feature extraction unit 2060 extracts a feature map from each person clip 60. The score calculation unit 2080 calculates a behavior score for each predefined behavior class and generates a behavior score vector 50 using the feature map. The location identification unit 2100 performs class activation mapping based on the behavior score vector 50 and the feature map to locate the detected behavior in the target clip 10.

[0028] <Example of hardware configuration> The behavior location identification device 2000 may be realized by one or more computers. Each of the one or more computers may be a dedicated computer manufactured for realizing the behavior location identification device 2000, or may be a general-purpose computer such as a personal computer (PC), a server machine, or a mobile device.

[0029] The behavior location identification device 2000 may be realized by installing an application on a computer. The application is realized by a program that causes the computer to function as the behavior location identification device 2000. In other words, the program implements the functional units of the behavior location identification device 2000.

[0030] Fig. 4 is a block diagram showing an example of the hardware configuration of a computer 1000 that realizes the behavior location identification device 2000. In Fig. 4, the computer 1000 has a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output (I / O) interface 1100, and a network interface 1120.

[0031] The bus 1020 is a data transmission path through which the processor 1040, memory 1060, storage device 1080, input / output interface 1100, and network interface 1120 transmit and receive data to and from each other. The processor 1040 is a processor such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage element such as a random access memory (RAM) or a read-only memory (ROM). The storage device 1080 is an auxiliary storage element such as a hard disk, a solid state drive (SSD), or a memory card. The input / output interface 1100 is an interface between the computer 1000 and peripheral devices such as a keyboard, a mouse, or a display device. The network interface 1120 is an interface between the computer 1000 and a network. The network may be a local area network (LAN) or a wide area network (WAN). In some implementations, the computer 10 00 is connected to the fisheye camera 30 via this network. The storage device 1080 may store the above-mentioned programs. Processor The program 1040 executes the program to realize each functional unit of the behavior location identification device 2000 .

[0032] The hardware configuration of the computer 1000 is not limited to that shown in Fig. 4. For example, as described above, the behavior location identification device 2000 may be realized by a plurality of computers. In this case, the computers may be connected to each other via a network.

[0033] One or more of the functional components of the behavior location identification device 2000 may be implemented in the fisheye camera 30. In this case, the fisheye camera 30 functions as the whole or part of the computer 500. For example, if all of the functional components of the behavior location identification device 2000 are implemented in the fisheye camera 30, the fisheye camera 30 may analyze the target clip 10 it has generated, detect behavior from the target clip 10, locate the detected behavior within the target clip 10, and output information indicating the result of the behavior location identification for the target clip 10. The fisheye camera 30 that functions as described above may be a network camera, an IP (internet protocol) camera, or an intelligent camera.

[0034] <Processing flow> 5 is a flowchart showing an exemplary flow of processing performed by the behavior location identification device 2000. The acquisition unit 2020 acquires a target clip 10 (S102). The person clip generation unit 2040 detects one or more people 40 from the target clip 10 (S104). The person clip generation unit 2040 generates a person clip 60 for each detected person 40 (S106). The feature extraction unit 2060 extracts a feature map from each person clip 60 (S108). The score calculation unit 2080 generates a behavior score vector 50 based on the feature map (S110). The position identification unit 2100 identifies the position of the behavior detected from the target clip 10 (S112).

[0035] <Getting target clip 10: S102> The acquiring unit 2020 acquires a target clip 10 (S102). As described above, the target clip 10 is a part of the video 20 generated by the fisheye camera 30. The number of target images 12 in the target clip 10 may be predetermined. Since multiple sequences of a predetermined number of video frames 22 may be extracted from different parts of the video 20, the behavior location identifying device 2000 may treat each of these sequences as a target clip 10. In this way, the behavior location identifying device 2000 may perform behavior location identification for each of the different parts of the video 20.

[0036] There are various ways to acquire the subject clips 10. For example, the acquisition unit 2020 may acquire the video 20 and divide it into multiple sequences of a predetermined number of video frames 22, thereby acquiring multiple subject clips 10. The video 20 may be acquired by accessing a memory unit on which the video 20 is stored, or by receiving the video 20 transmitted from another computer, such as a fisheye camera 30.

[0037] In another example, the acquisition unit 2020 may periodically acquire one or more video frames 22. Then, the acquisition unit 2020 generates a target clip 10 from a predetermined number of acquired video frames 22. The acquisition unit 2020 can generate multiple target clips 10 by repeating the above-described flow. The video frames 22 may be acquired in a manner similar to the method for acquiring the video 20 described above.

[0038] <Detection of Person 40: S104> The person clip generator 2040 detects one or more people 40 from the target clip 10 (S104). There are various well-known methods for detecting people from a sequence of fisheye images, and the person clip generator 2040 may be configured to detect the people 40 from the target clip 10 using one of these methods. For example, the person 40 is detected from the target clip 10 using a machine learning-based model (hereinafter referred to as a "person detection model"). The person detection model is configured to receive a fisheye clip as input, and is pre-trained to output, in response to the input of the fisheye clip, a person in the input fisheye clip and a region (e.g., a bounding box) containing the person for each fisheye video frame.

[0039] The person clip generation unit 2040 may detect the person 40 from the entire area of ​​the target clip 10, or may detect the person 40 from a partial area of ​​the target clip 10. In the latter case, for example, the person 40 is detected from a circular area in each of the target images 12 of the target clip 10. An example of this latter case will be described later as a second embodiment of the behavior position identification device 2000.

[0040] <Creating Person Clip 60: S106> The person clip generator 2040 generates a person clip 60 for each person 40 detected from the target clip 10 (S106). The person clip 60 of a target person is a sequence of images called "person images 62," each of which is a partial area of ​​the corresponding target image 12 and includes the target person. Hereinafter, the area cut out from the target image 12 to generate the person image 62 will be referred to as a "cutout area." Note that the dimensions (width and height) of the cutout area may be defined in advance.

[0041] The person clip generator 2040 may operate as follows for each target person (person 40 detected from the target clip 10): The person clip generator 2040 rotates the target images 12 by the same angle relative to each other so that the target person appears approximately upright in each of the target images 12. For example, if the person clip generator 2040 detects a bounding box for each person 40, the rotation angle of the target image 12 for the target person may be determined based on the orientation of the bounding box of the target person in a representative image (e.g., the first image) of the target images 12.

[0042] Next, the person clip generator 2040 identifies a cutout position (i.e., a position of a cutout region) where the target person is included in the cutout region in each of the rotated target images 12. Next, the person clip generator 2040 cuts out an image from the identified cutout region of the rotated target image 12 to generate a person image. 6 2, thereby generating a person clip 60 for the person of interest.

[0043] <Feature extraction: S108> The feature extraction unit 2060 extracts a feature map from each of the person clips 60 (S108). The feature map represents spatiotemporal features of the person clips 60. The feature map may be extracted using a neural network, such as a 3D convolutional neural network (CNN), which is configured to acquire clips and is pre-trained to output spatiotemporal features of the input clips as a feature map in response to the input clips.

[0044] The feature map may represent the spatiotemporal characteristics of the entire region of the person clip 60, or may represent the characteristics of a portion of the region of the person clip 60. In the latter case, the feature map may represent the spatiotemporal characteristics of the target person and the region around it.

[0045] FIG. 6 illustrates a case in which a feature map represents spatiotemporal features of a target person and their surrounding area. In FIG. 6, a network 70 receives a person clip 60 as input and outputs a feature map 80. The network 70 can be any type of 3D CNN, such as a 3D ResNet. The feature extraction unit 2060 then performs pooling, such as max pooling, on the bounding boxes of the target person across all person images 62 to obtain a binary mask. The feature extraction unit 2060 then resizes the binary mask to the same width and height as the feature map 80 and multiplies the resized mask by the feature map 80 to obtain a feature map 90. Hereinafter, the feature map output from the feature extraction unit 2060 will be referred to as the "feature map 90," regardless of whether it is masked with a binary mask or not.

[0046] <Calculation of behavior score: S110> The score calculation unit 2080 calculates the behavior score for the predefined behavior classes, thereby calculating the behavior score vector 50 (S110). The feature map 90 obtained from the person clip 60 is used to calculate the behavior score. In some implementations, the score calculation unit 2080 may operate as shown in FIG. 7. FIG. 7 shows an example of a method for calculating the behavior score based on the feature map 90. In this example, the person clip 60-1 from the target clip 10 is used to calculate the behavior score vector 50. from 60-N are obtained, and feature maps 90-1 to 90-N are obtained from the person clips 60-1 to 60-N, respectively.

[0047] First, the score calculation unit 2080 performs pooling, for example, average pooling, on each of the feature maps 90. Then, each of the pooling results is input to the fully connected layer 200. The fully connected layer 200 is configured to obtain the pooling results on the feature maps 90, and is pre-trained to output, for each pre-defined behavior class, a vector called an "intermediate vector 210" that represents the confidence of the behavior of that behavior class included in the person clip 60 corresponding to the feature map 90.

[0048] Using the fully connected layer 200, intermediate vectors 210 are obtained for all of the feature maps 90. Then, the score calculation unit 2080 aggregates the intermediate vectors 210 into a behavior score vector 50. The intermediate vectors 210 may be aggregated using a log-sum-exponential function such as that disclosed in Non-Patent Document 1. Note that the behavior score vector 50 may be scaled so that the maximum range of each element is [0, 1].

[0049] <Class Activation Mapping: S112> The location identification unit 2100 performs class activation mapping such as CAM (class activation mapping), Grad-CAM, and SmoothGrad to identify the location of the behavior detected in the target clip 10. Class activation mapping is a method for finding one or more regions in the input image that are associated with the prediction result (in the case of the behavior location identification device 2000, the behavior score of the behavior class). The location identification unit 2100 performs class activation mapping for each behavior class detected in the target clip 10 and for each person clip 60 to identify where in the target clip 10 the behavior of the detected behavior class occurred.

[0050] The behavioral classes detected from the target clip 10 are those that each have a behavioral score greater than or equal to a predefined threshold Tc. Assume that there are three predefined behavioral classes A1, A2, and A3, and their behavioral scores are 0.8, 0.3, and 0.7, respectively. The threshold Tc is 0.6. In this case, the behavioral classes detected are behavioral classes A1 and A3 because their behavioral scores are greater than the threshold Tc.

[0051] FIG. 8 shows an example of a class activation map. This figure shows a class activation map generated for person clip P1. Person clip P1 includes the activity "operate phone." Assume the detected activities are "drink," "operate phone," and "walk." In this case, the localization unit 2100 generates a class activation map for each of these three detected activities. The darker an area in the class activation map is, the more relevant that area is to predicting the activity score of the activity class corresponding to that map. Additionally, the darker a location in the class activation map is, the greater the value of the cell in the map corresponding to that location.

[0052] In Figure 8, the class activation map for "operating the phone" has a large dark area, while the other two class activation maps do not. Therefore, it is possible to predict that the action included in person clip P1 is "operating the phone." Furthermore, by overlaying the action class map for "operating the phone" on person clip P1, it is possible to predict the area of ​​person clip P1 where the action of "operating the phone" is performed.

[0053] FIG. 9 is a flowchart showing an exemplary flow of processing performed by the position identification unit 2100. Steps S202 to S212 form a loop process L1 that is performed for each person clip 60. In step S202, the position identification unit 2100 determines whether the loop process L1 has been performed for all of the person clips 60. If the loop process L1 has been performed for all of the person clips 60, the processing shown in FIG. 9 ends. On the other hand, if the loop process L1 has not been performed for all of the person clips 60, the position identification unit 2100 selects a person clip 60 for which the loop process L1 has not yet been performed. The person clip 60 selected here is referred to as a "person clip Pi."

[0054] Steps S204 to S208 form a loop process L2 that is performed for each detected behavior class. In step S204, the position identification unit 2100 determines whether or not the loop process L2 has been performed for all of the detected behavior classes. If the loop process L2 has been performed for all of the detected behavior classes, S210 is performed next. On the other hand, if the loop process L2 has not been performed for all of the detected behavior classes, the position identification unit 2100 selects a behavior class from the detected behavior classes for which the loop process L2 has not yet been performed. The behavior class selected here is referred to as "behavior class Aj."

[0055] The position identification unit 2100 obtains the behavior score vector Pi from the person clip Pi. 5Using the behavioral score of Aj, indicated by 0, and the feature map 90, class activation mapping is performed on the person clip Pi (S206). As mentioned above, there are various types of class activation mapping, and any of them can be adopted.

[0056] For example, when Grad-CAM is employed, the class activation map can be generated in a similar manner to that disclosed in Non-Patent Document 1. Specifically, for each channel of the feature map 90 obtained from the person clip Pi, the position identification unit 2100 calculates the importance of the channel for predicting the behavior score of the behavior class Aj based on the gradient of the behavior score for the channel. This can be formulated as follows:

number

[0057] The importance of each channel calculated above is then used as the weight for each channel to generate a class activation map as a weighted combination of the channels of feature map 90 for person clip P1, which can be formulated as follows:

number

[0058] Step S208 is the end of the loop process L2, and therefore step S204 is performed next.

[0059] After the loop process L2 for the person clip Pi is completed, the position identification unit 2100 has class activation maps obtained from the person clip Pi for all detected behavior classes. In step S210, the position identification unit 2100 identifies the behavior class of the behavior performed by the target person in the person clip Pi and locates the behavior based on the obtained class activation maps. To do this, the position identification unit 2100 identifies one of the class activation maps corresponding to the behavior class of the behavior performed by the target person in the person clip Pi.

[0060] The class activation map calculated for behavior class Aj can be said to contain an area highly correlated with the behavior score of behavior class Aj only if the target person in person clip Pi behaves in behavior class Aj. Therefore, the class activation map most correlated with the behavior score (predicted behavior score) of the corresponding behavior class is the one for person clip P. i This corresponds to the class of actions taken in the

[0061] Specifically, for example, the location identification unit 2100 calculates the sum of the cells for each class activation map and identifies which class activation map has the largest sum. The location identification unit 2100 then identifies the behavior class corresponding to the class activation map with the largest sum as the behavior class of the behavior performed in the person clip 60. This can be formulated as follows:

number

[0062] <Output from behavior location identification device 2000> The behavior location identification device 2000 may output information called "output information" indicating the spatial and temporal behavior location results, i.e., information indicating what type of behavior is performed in what area of ​​the target clip 10 during what period. The information indicated by the output information may be of various types. In some implementations, the output information may include, for each behavior of a detected behavior class, a set of the behavior class of the behavior, the location where the behavior was performed (e.g., the location of the bounding box of the person 40 who performed the behavior), and the period where the behavior was performed (e.g., the frame number of the target clip 10). For example, the output information may include, for each person 40 detected in the target clip 10, the target clip 10 modified to indicate the bounding box of the person 40 including an annotation indicating the behavior class of the behavior performed by the person 40.

[0063] 10 shows a first example in which a target clip 10 is modified to show the results of activity localization. In this example, the activity classes "drinking," "operating phone," and "walking" are detected and localized. To show the results, a combination of bounding box 220 and annotation 230 is superimposed on the target clip 10. The bounding box 220 indicates the location of the detected person 40, and the corresponding annotation 230 indicates the activity class of the activity performed by this person 40.

[0064] In another implementation, the output information may include a modified target clip 10 with a class activation map of the detected behavior class superimposed thereon. Suppose a target person in person clip Pi is identified as having performed a behavior in behavior class Aj. In this case, the target image 12 corresponding to the person image 62 in person clip Pi is modified to superimpose the class activation map generated for the combination of person clip Pi and behavior class Aj. The location in the target image 12 where the class activation map is superimposed is the location where the corresponding person image 62 is cropped.

[0065] Figure 11 shows a second example in which the target clip 10 is modified to show the results of activity localization. This figure assumes the same situation as Figure 10. However, in Figure 11, instead of the bounding box 220, a map 240 is superimposed on the target clip 10. The map 240 superimposed on a person is a class activation map generated for that person's person clip 60 and corresponding to the activity class of the activity performed by that person.

[0066] <On Trainable Parameter Optimization> The behavior localization device 2000 has trainable parameters such as weights in the network 70 and the fully connected layer 200. These trainable parameters are pre-optimized by repeatedly updating them using multiple training data (in other words, the behavior localization device 2000 is pre-trained). The training data may include a combination of test clips of behavior score vectors 50 and ground truth data. The test clip is any clip generated by a top-view fisheye camera (preferably the fisheye camera 30) and containing one or more people. The ground truth data is a behavior score vector that indicates the maximum confidence (e.g., 1 when the behavior score vector 50 is scaled to [0, 1]) for each behavior class included in the test clip.

[0067] The trainable parameters are updated based on a loss representing the difference between the ground truth data and the behavior score vector 50 calculated by the behavior localization device 2000 in response to the input of the test clip. The loss may be calculated using any type of loss function, such as a cross-entropy loss function. The loss function may further include a regularization term. For example, since a target person may perform a single behavior in a person clip 60, it is preferable to penalize cases where the intermediate vector (the behavior score vector calculated from a single feature map 80) indicates high confidence for multiple behavior classes. This type of regularization term is disclosed in Non-Patent Document 1.

[0068] Embodiment 2 Fig. 12 shows an outline of the behavior location identification device 2000 of embodiment 2. Note that the outline shown in Fig. 12 shows an example of the operation of the behavior location identification device 2000 of embodiment 2 in order to make it easier to understand the behavior location identification device 2000 of embodiment 2, and is not intended to limit or narrow the range of operations that the behavior location identification device 2000 of embodiment 2 can perform.

[0069] In the target clip 10, people appear at different angles. Non-Patent Document 1 addresses this problem by converting video frames acquired from a fisheye camera into a panoramic image and analyzing the panoramic image to detect and localize activities.

[0070] However, converting the image into a panoramic image in this way can distort people located near the center (the optical axis of the fisheye lens when viewed from above). Regarding this problem, the method performed by the behavior location identification device 2000 is considered to be more effective than the method described in Non-Patent Document 1 in order to identify the location of behavior that occurred around the optical axis of the fisheye lens, because people located near the center are not distorted.

[0071] Therefore, the behavior location identification device 2000 of the second embodiment generates two different types of clips from the target clip 10 and calculates behavior scores for these two clips using two different methods. By doing so, it is possible to more accurately identify the location of the behavior within the target clip 10. Hereinafter, these two types of clips will be referred to as the "center clip 100" and the "panoramic clip 110," respectively. Furthermore, the method performed on the center clip 100 will be referred to as "fisheye processing," and the method performed on the panoramic clip 110 will be referred to as "panoramic processing."

[0072] The central clip 100 is a sequence of central regions of the target images 12. To generate the central clip 100, the behavior location identification device 2000 extracts a central region 14 from each of the target images 12. The central clip 100 is generated as a sequence of central regions 14. The central region 14 is a circular region with a predefined radius, whose center is located at a position corresponding to the optical axis of the fish-eye camera 30. The position corresponding to the optical axis of the fish-eye camera 30 may be detected by the method disclosed in Non-Patent Document 1. Hereinafter, each image (i.e., central region 14) included in the central clip 100 will be referred to as a "central image 1 0 It is called "2".

[0073] The panoramic clip 110 is a sequence of target images converted into a panoramic image. To generate the panoramic clip 110, the behavior locating device 2000 converts each target image 12 into a panoramic image. The panoramic clip 110 is generated as a sequence of these panoramic images. The target images 12 may be converted into panoramic images using the method disclosed in Non-Patent Document 1. Hereinafter, each image included in the panoramic clip 110 will be referred to as a "panoramic image 112."

[0074] The fisheye processing is performed by the behavior location identification device 2000 of the eleventh embodiment. clip The method for calculating a behavior score from the central clip 100 is similar to the method for calculating a behavior score for the central clip 10. Specifically, the fisheye processing includes steps of detecting one or more people 40 from the central clip 100, generating a person clip 60 for each person 40 detected from the central clip 100, extracting a feature map 90 from each of the person clips 60, and calculating a behavior score based on the feature map 90. Hereinafter, a vector indicating the behavior score calculated for the central clip 100 will be referred to as a "behavior score vector 130."

[0075] Panoramic processing is a method for calculating behavioral scores in a manner similar to that disclosed in Non-Patent Document 1. Specifically, panoramic processing can be performed as follows. The behavior locating device 2000 calculates a feature map (spatiotemporal features) of the panoramic clip 110 by inputting the panoramic clip 110 to a neural network, such as a 3D CNN, which can extract spatiotemporal features from a sequence of images as a feature map. Next, the behavior locating device 2000 calculates a binary mask by detecting people in the panoramic clip 110, resizes the binary mask to the same width and height as the feature map, and multiplies the binary mask by the feature map to obtain a masked feature map. The masked feature map is divided into multiple blocks. The behavior locating device 2000 pools each block and then inputs it to a fully connected layer, thereby obtaining a behavioral score for each block. The behavioral scores obtained for each block are aggregated into a single vector that indicates the behavioral score for the entire panoramic clip 110. Hereinafter, this vector will be referred to as the "behavior score vector 140."

[0076] As described above, the behavior location identifying device 2000 obtains the behavior score vector 130 as a result of the fisheye processing and the behavior score vector 140 as a result of the panoramic processing. The behavior location identifying device 2000 uses the behavior score vector 130 and the behavior score vector 140 to identify the target clip 1. Within 0 Localize each action.

[0077] In some implementations, the behavior location identifying device 2000 uses the behavior score vector 130 and the behavior score vector 140 separately as shown in Fig. 12. In this case, the behavior location identifying device 2000 uses the behavior score vector 130 to identify the behavior location for the center clip 100 in the same way as the behavior location identifying device 2000 of the eleventh embodiment uses the behavior score vector 50 to identify the behavior location for the target clip 10. As a result, the behavior score vector 140 is used to identify the behavior location for the center clip 100. 00, the behavior detected is located. Furthermore, the behavior location identification device 2000 uses the behavior score vector 140 to identify the behavior of the panoramic clip 110 in a manner similar to the method disclosed in Non-Patent Document 1. In this way, the behavior detected from the panoramic clip 110 is located. Next, the behavior location identification device 2000 aggregates the results of the behavior location identification performed on the center clip 100 and the results of the behavior location identification performed on the panoramic clip 110, thereby locating the behavior of the behavior class detected for the entire target clip 10.

[0078] In another implementation, the behavior location identification device 2000 aggregates the behavior score vector 130 and the behavior score vector 140 into a single vector called an "aggregated behavior score vector." In this case, behavior classes whose behavior scores are equal to or greater than the threshold Tc are treated as detection targets. Instead of using the behavior score vectors 130 and 140 separately, the behavior location identification device 2000 uses the aggregated behavior score vector to perform behavior location on the central clip 100 and the panoramic clip 110 separately for each detected behavior class, and aggregates the behavior location results.

[0079] Except that the aggregate behavior score vector is used for class activation mapping, the behavior localization performed for center clip 100 in this case is the same as when behavior score vectors 130 and 140 are used separately. Similarly, the behavior localization performed for panoramic clip 110 in this case is the same as when behavior score vectors 130 and 140 are used separately, except that the aggregate behavior score vector is used for class activation mapping.

[0080] <Example of functional configuration> FIG. 13 is a block diagram showing an example of the functional configuration of a behavior location identification device 2000 according to the second embodiment. The behavior location identification device 2000 includes an acquisition unit 2020, a center clip generation unit 2120, a panoramic clip generation unit 2140, a fisheye processing unit 2160, a panoramic processing unit 2180, and a position identification unit 2100. As described in the eleventh embodiment, the acquisition unit 2020 acquires a target clip 10. The center clip generation unit 2120 generates a center clip 100 from the target clip 10. The panoramic clip generation unit 2140 generates a panoramic clip 110 from the target clip 10. The fisheye processing unit 2160 performs fisheye processing on the center clip 100 to calculate a behavior score vector 130. Although not shown in FIG. 13, the person clip generation unit 2040, the feature extraction unit 2060, and the score calculation unit 2080 are included in the fisheye processing unit 2160. The panorama processing unit 2180 performs panorama processing on the panorama clip 110 to calculate the behavior score vector 140. The position identifying unit 2100 uses the behavior score vector 130 and the behavior score vector 140 to identify the behavior position for the target clip 10.

[0081] <Example of hardware configuration> Like the behavior location identification device 2000 of embodiment 11, the behavior location identification device 2000 of embodiment 2 may have the hardware configuration shown in Fig. 3. However, the storage device 1080 of embodiment 2 may further include a program that realizes the functional configuration of the behavior location identification device 2000 of embodiment 2.

[0082] <Processing flow> 14 is a flowchart showing the flow of processing performed by the behavior location identification device 2000 of the second embodiment. The acquisition unit 2020 acquires a target clip 10 (S302). Between steps S302 and S316, there are two processing sequences shown to be performed in parallel. The first processing sequence includes steps S304 to S308 and is performed to locate behavior within the central clip 100. Meanwhile, the second processing sequence includes steps S310 to S314 that are performed to locate behavior within the panoramic clip 110. Note that in other implementations, these sequences may be performed sequentially rather than in parallel.

[0083] The first sequence of processing is performed as follows: The center clip generation unit 2120 generates the center clip 100 (S304). The fisheye processing unit 2160 performs fisheye processing on the center clip 100 to calculate the behavior score vector 130 (S306). The location identification unit 2100 uses the behavior score vector 130 to locate the behavior of the behavior class detected for the center clip 100 (S308).

[0084] The second flow of processing is performed as follows: The panoramic clip generation unit 2140 generates the center clip 100 (S310). The panoramic processing unit 2180 performs panoramic processing on the panoramic clip 110 to calculate the behavior score vector 140 (S312). The position identification unit 2100 uses the behavior score vector 140 to calculate the panoramic clip 100. 1 0, locate the detected behavior of the behavior class (S308).

[0085] The location unit 2100 aggregates the results of the activity location determination for the central clip 100 and the activity location determination for the panoramic clip 110 .

[0086] 14 assumes that the behavior score vectors 130 and 140 are used separately to localize the behavior. However, as described above, the behavior localization may be performed using an aggregate behavior score vector instead of using the behavior score vectors 130 and 140 separately. In this case, after the fisheye processing and the panoramic processing are completed, the aggregate behavior score vector is calculated, and then the behavior localization for the center clip 100 and the panoramic clip 110 is performed using the aggregate behavior score vector.

[0087] <Output from behavior location identification device 2000> The behavior location identification device 2000 of the second embodiment may output output information similar to that output by the behavior location identification device 2000 of the eleventh embodiment. However, the output information of the second embodiment may indicate an aggregated result of behavior location identification for the center clip 100 and the panoramic clip 110. For example, as shown in FIG. 10, when the output information includes a target clip 10 on which a bounding box 220 and an annotation 230 are superimposed, the behavior location identification device 2000 generates a bounding box 220 and an annotation 230 for both a person whose behavior is identified based on the results of fisheye processing and a person whose behavior is identified based on the results of panoramic processing. The same applies when a map 240 and an annotation 230 are superimposed on the target clip 10, as shown in FIG. 11.

[0088] In some cases, there may be a person included in both the central clip 100 and the panoramic clip 110. In this case, the localization unit 2100 may obtain class activation maps for the person from both the central clip 100 and the panoramic clip 110. The localization unit 2100 may combine these class activation maps and localize the person's activity based on the combined class activation map. The class activation maps may be combined by taking their intersections or intermediate points.

[0089] <Trainable parameter optimization> The behavior localization device 2000 of the second embodiment further includes trainable parameters used for panoramic processing in addition to the trainable parameters mentioned in the eleventh embodiment. Similar to those employed in the eleventh embodiment, the trainable parameters of the behavior localization device 2000 of the second embodiment may be optimized in advance using a plurality of training data. However, in this embodiment, the loss may be calculated using an aggregate behavior score vector instead of the behavior score vector 50. Specifically, the loss may be calculated to represent the difference between the ground truth data of the aggregate behavior score indicated by the training data and the aggregate behavior score calculated by the behavior localization device 2000 of the second embodiment in response to the input of the test clip.

[0090] The program can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROMs, CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, programmable ROMs (PROMs), erasable PROMs (EPROMs), flash ROMs, and RAMs). The program may also be provided to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can provide the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.

[0091] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the arrangements and details of the present disclosure within the scope of the present invention.

[0092] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. (Appendix 1) A behavior location identification device, at least one memory configured to store instructions; Execute the instructions, acquiring a target clip, the target clip being a sequence of target images, the target images being fisheye images of one or more people captured from a generally overhead perspective; Detecting one or more people from the clip of interest; generating a person clip from the object clip for each person detected in the object clip, the person clip being a sequence of person images each of which is a region of the object image and includes the detected person corresponding to the person clip; extracting a feature map from each said person clip; calculating, for each predetermined behavior class, a behavior score indicating a reliability of the behavior of the behavior class included in the target clip based on the feature map extracted from the person clip, the behavior class being a type of behavior; performing class activation mapping for each of the person clips to locate each behavior in a behavior class having the behavior score equal to or greater than a threshold; and at least one processor configured to: (Appendix 2) The activity localization includes, for each person clip, generating a class activation map for each behavior class whose behavior score is greater than or equal to the threshold using the behavior score for that behavior class and the feature map extracted from the person clip; identifying the class activation map that exhibits the highest correlation with the behavioral score of the corresponding behavioral class; and identifying the behavior class corresponding to the identified class activation map as the behavior class of the behavior included in the person clip. (Appendix 3) The calculation of the behavioral score includes: calculating, for each person clip, an intermediate vector indicating a confidence level of the behavior of the behavior class included in the person clip for each of the predetermined behavior classes; aggregating the intermediate vectors into a behavior score vector indicating the behavior score of the predetermined behavior class. (Appendix 4) the extraction of the feature map and the calculation of the behavioral score are performed by a pre-trained neural network; 4. The behavior localization device of claim 3, wherein the pre-trained neural network is trained using a training dataset including test clips and vectors that indicate maximum confidence for each of the behavior classes included in the test clips, and the test clips include fisheye images of one or more people captured from a generally overhead perspective. (Appendix 5) 5. The behavioral location determination apparatus of claim 3 or 4, wherein the intermediate vectors are aggregated into the behavior score vector using a log-sum-exponential function. (Appendix 6) The at least one processor further executes the instructions to: generating a sequence of center clips, each of which is generated by cropping a center region from a corresponding said target image; generating a panoramic clip, which is a sequence of panoramic images, each generated by converting a corresponding said target image into a panoramic image; The method is configured to locate the activity included in the target clip by locating the activity included in the central clip, locating the activity included in the panoramic clip, and aggregating results of locating the activity included in the central clip and locating the activity included in the panoramic clip; The location of the action included in the central clip is detecting one or more people from the center clip; generating the people clips from the center clip for each person detected in the center clip; extracting the feature map from each of the people clips; calculating the behavior score for each of the predetermined behavior classes based on the feature map extracted from the person clip; and locating each behavior of a behavior class having the behavior score above a threshold by performing class activation mapping on each of the person clips. (Appendix 7) A control method performed by a computer, comprising: acquiring a target clip, the target clip being a sequence of target images, the target images being fisheye images of one or more people captured from a generally overhead perspective; Detecting one or more people from the clip of interest; generating a person clip from the target clip for each person detected in the target clip, the person clip being a sequence of person images each of which is a region of the target image and includes the detected person corresponding to the person clip; extracting a feature map from each said person clip; calculating, for each predetermined behavior class, a behavior score indicating a reliability of the behavior of the behavior class included in the target clip based on the feature map extracted from the person clip, wherein the behavior class is a type of behavior; A control method comprising the step of performing class activation mapping for each of said person clips to locate each behavior of a behavior class having said behavior score above a threshold. (Appendix 8) The activity localization includes, for each person clip, generating a class activation map for each behavior class whose behavior score is greater than or equal to the threshold using the behavior score for that behavior class and the feature map extracted from the person clip; identifying the class activation map that exhibits the highest correlation with the behavioral score of the corresponding behavioral class; A control method as described in Appendix 7, including identifying the behavior class corresponding to the identified class activation map as the behavior class of the behavior included in the person clip. (Appendix 9) The calculation of the behavioral score includes: calculating, for each person clip, an intermediate vector indicating a confidence level of the behavior of the behavior class included in the person clip for each of the predetermined behavior classes; aggregating the intermediate vectors into a behavior score vector indicating the behavior scores of the predetermined behavior class. (Appendix 10) the extraction of the feature map and the calculation of the behavioral score are performed by a pre-trained neural network; 10. The control method of claim 9, wherein the pre-trained neural network is trained using a training dataset including test clips and vectors that indicate maximum confidence for each of the behavioral classes contained in the test clips, and the test clips include fisheye images of one or more people captured from a generally overhead perspective. (Appendix 11) 11. The control method of claim 9 or 10, wherein the intermediate vectors are aggregated into the behavioral score vector using a log-sum-exponential function. (Appendix 12) generating a sequence of center clips, each of which is generated by cropping a center region from a corresponding said target image; generating a panoramic clip, which is a sequence of panoramic images, each generated by converting a corresponding one of the target images into a panoramic image; locating the activity included in the target clip by locating the activity included in the central clip, locating the activity included in the panoramic clip, and aggregating results of locating the activity included in the central clip and locating the activity included in the panoramic clip; The location of the action included in the central clip is detecting one or more people from the center clip; generating the people clips from the center clip for each person detected in the center clip; extracting the feature map from each of the people clips; calculating the behavior score for each of the predetermined behavior classes based on the feature map extracted from the person clip; 12. A control method according to any one of appendices 7 to 11, comprising: locating each behavior of a behavior class having the behavior score above a threshold by performing class activation mapping on each of the person clips. (Appendix 13) A non-transitory computer-readable storage medium storing a program, the program being configured to: acquiring a target clip, the target clip being a sequence of target images, the target images being fisheye images of one or more people captured from a generally overhead perspective; Detecting one or more people from the clip of interest; for each person detected from the target clip, generating a person clip from the target clip, wherein the person clip is a sequence of person images, each of the person clips being a region of the target image, and including the detected person corresponding to the person clip; extracting a feature map from each said person clip; calculating an action score for each predetermined action class based on the feature map extracted from the person clip, the action score indicating a reliability of the action of the action class included in the target clip, the action class being a type of action; A non-transitory computer-readable storage medium for causing a step of locating each behavior of a behavior class having the behavior score equal to or greater than a threshold by performing class activation mapping on each of the person clips. (Appendix 14) The activity localization includes, for each person clip, generating a class activation map for each behavior class whose behavior score is greater than or equal to the threshold using the behavior score for that behavior class and the feature map extracted from the person clip; identifying the class activation map that exhibits the highest correlation with the behavioral score of the corresponding behavioral class; and identifying the behavior class corresponding to the identified class activation map as the behavior class of the behavior included in the person clip. (Appendix 15) The calculation of the behavioral score includes: calculating, for each person clip, an intermediate vector indicating a confidence level of the behavior of the behavior class included in the person clip for each of the predetermined behavior classes; aggregating the intermediate vectors into a behavior score vector indicative of the behavior scores for the predetermined behavior class. (Appendix 16) the extraction of the feature map and the calculation of the behavioral score are performed by a pre-trained neural network; 16. The storage medium of claim 15, wherein the pre-trained neural network is trained using a training dataset including test clips and vectors that indicate maximum confidence for each of the behavioral classes contained in the test clips, and the test clips include fisheye images of one or more people captured from a generally overhead perspective. (Appendix 17) 17. The storage medium of claim 15 or 16, wherein the intermediate vectors are aggregated into the behavioral score vector using a log-sum-exponential function. (Appendix 18) The program causes the computer to generating a sequence of center clips, each of which is generated by cropping a center region from a corresponding said target image; generating a panoramic clip, which is a sequence of panoramic images, each generated by converting a corresponding one of the target images into a panoramic image; locating the activity in the target clip by locating the activity in the central clip, locating the activity in the panoramic clip, and aggregating results of locating the activity in the central clip and locating the activity in the panoramic clip; The location of the action included in the central clip is detecting one or more people from the center clip; generating the people clips from the center clip for each person detected in the center clip; extracting the feature map from each of the people clips; calculating the behavior score for each of the predetermined behavior classes based on the feature map extracted from the person clip; 18. A storage medium described in any one of Appendixes 13 to 17, comprising: locating each behavior of a behavior class having the behavior score above a threshold by performing class activation mapping on each of the person clips. [Explanation of symbols]

[0093] 10 Target Clips 12 Target image 14 Central area 20 videos 22 video frames 30 Fisheye Camera 40 people 50, 130, 140 Behavioral score vector 50 60 People Clips 62 Portraits 70 Network 80, 90 feature maps 100 Center Clip 102 Center image 110 Panorama Clip 112 panoramic images 200 fully connected layers 210 Intermediate Vector 220 Bounding Box 230 comments 240 maps 1000 computers 1020 Bus 1040 processor 1060 memory 1080 storage device 1100 Input / Output Interface 1120 Network Interface 2000 Behavioral location identification device 2020 Acquisition Department 2040 Person Clip Generation Unit 2060 Feature Extraction Unit 2080 Score Calculation Department 2100 Location identification part 2120 Center Clip Generation Unit 2140 Panorama Clip Generation Unit 2160 Fisheye Processing Unit 2180 Panorama Processing Unit

Claims

1. an acquisition means for acquiring a target clip, the target clip being a sequence of fisheye images of one or more people; a person clip generating means for generating, for each person detected from the target clip, a person clip showing a sequence of an area including the person and being a part of the fisheye image; feature extraction means for extracting a feature map from each said person clip; a score calculation means for calculating, for each behavior class indicating a type of behavior based on the feature map, a behavior score indicating a reliability of the behavior of the behavior class included in the target clip; and a position identifying means for identifying the position of each behavior of the behavior class whose behavior score is equal to or greater than a threshold by performing class activation mapping on each of the person clips.

2. The position specifying means, for each person clip, generating a class activation map for each of the behavior classes whose behavior scores are equal to or greater than the threshold using the behavior scores of the behavior classes and the feature maps extracted from the person clips; identifying the class activation map that exhibits the highest correlation with the behavioral score of the corresponding behavioral class; The behavior location identification device according to claim 1 , wherein the behavior class corresponding to the identified class activation map is identified as the behavior class of the behavior included in the person clip.

3. The score calculation means For each person clip, an intermediate vector is calculated that indicates the reliability of the behavior of the behavior class included in the person clip for each of a plurality of predetermined behavior classes; The behavior location identifying device according to claim 1 , wherein the intermediate vectors are aggregated into a behavior score vector indicating the behavior score of the predetermined behavior class.

4. the extraction of the feature map and the calculation of the behavioral score are performed by a pre-trained neural network; 4. The behavior location identification device of claim 3, wherein the pre-trained neural network is trained using a training dataset including test clips and vectors that indicate maximum confidence for each of the behavior classes included in the test clips, and the test clips include fisheye images of one or more people captured from a generally overhead perspective.

5. The behavior location determining device according to claim 3 or 4, wherein the intermediate vectors are aggregated into the behavior score vector using a log-sum-exponential function.

6. a center clip generating means for generating a center clip, which is a sequence of center images, each center clip being generated by cropping a center region from a corresponding said fisheye image; and a panorama clip generating means for generating a panorama clip, which is a sequence of panorama images each generated by converting a corresponding one of the fisheye images into a panorama image; the location identifying means identifies the activity included in the central clip, identifies the activity included in the panoramic clip, and identifies the activity included in the target clip by aggregating results of the location identifying the activity included in the central clip and the location identifying the activity included in the panoramic clip; The location of the action included in the central clip is detecting one or more people from the center clip; generating the people clips from the center clip for each person detected in the center clip; extracting the feature map from each of the people clips; calculating the behavior score for each of a plurality of predetermined behavior classes based on the feature map extracted from the person clip; and locating each activity of an activity class having the activity score equal to or greater than a threshold by performing class activation mapping on each of the person clips.

7. acquiring a clip of interest, the clip being a sequence of fisheye images of one or more people; generating, for each person detected in the object clip, a person clip showing a sequence of areas of the fisheye image that include the person; extracting a feature map from each said person clip; calculating, for each behavior class indicating a type of behavior based on the feature map, a behavior score indicating a reliability of the behavior of the behavior class included in the target clip; and performing class activation mapping for each of the person clips to identify the location of each behavior in the behavior class for which the behavior score is equal to or greater than a threshold.

8. The activity location is determined by, for each person clip, generating a class activation map for each of the behavior classes whose behavior scores are equal to or greater than the threshold using the behavior scores of the behavior classes and the feature maps extracted from the person clips; identifying the class activation map that exhibits the highest correlation with the behavioral score of the corresponding behavioral class; The control method of claim 7 , further comprising: identifying the behavior class corresponding to the identified class activation map as the behavior class of the behavior included in the person clip.

9. acquiring a clip of interest, the clip being a sequence of fisheye images of one or more people; generating, for each person detected in the object clip, a person clip showing a sequence of areas of the fisheye image that include the person; extracting a feature map from each said person clip; calculating, for each behavior class indicating a type of behavior based on the feature map, a behavior score indicating a reliability of the behavior of the behavior class included in the target clip; and performing class activation mapping for each of the person clips to identify the location of each behavior in the behavior class for which the behavior score is equal to or greater than a threshold.

10. The activity location is determined by, for each person clip, generating a class activation map for each of the behavior classes whose behavior scores are equal to or greater than the threshold using the behavior scores of the behavior classes and the feature maps extracted from the person clips; identifying the class activation map that exhibits the highest correlation with the behavioral score of the corresponding behavioral class; and identifying the behavior class corresponding to the identified class activation map as the behavior class of the behavior included in the person clip.

Citation Information

Patent Citations

  • Object detection method based on human-object interaction weak supervision label

    CN111931703A

  • Weakly-Supervised Action Localization by Sparse Temporal Pooling Network

    US20200272823A1

  • 4d convolutional neural networks for video recognition

    US20210248377A1