Behavior recognition system and method
By estimating and transmitting skeletal information from edge devices, the system addresses real-time behavior recognition challenges, reducing network load and delay while ensuring accurate and private processing.
Patent Information
- Application Number
- PCT/JP2024/015765
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-10-30
AI Technical Summary
Existing video-based behavior recognition technologies face challenges in real-time processing due to high network transmission requirements and processing loads when transmitting large amounts of video data to a central server.
The system recognizes behavior by estimating a person's skeleton from video data on an edge device and transmitting skeletal information to a behavior recognition device, reducing data transmission and processing load.
This approach enables accurate real-time behavior recognition with reduced network load and delay by transmitting less data, maintaining performance and enhancing privacy through encoding.
Smart Images

Figure JP2024015765_30102025_PF_FP_ABST
Abstract
Description
Activity Recognition System and Method
[0001] One aspect of the present invention relates to an activity recognition system and method used to recognize an activity of a person from, for example, the person's movement.
[0002] Various technologies have been proposed that use video or sensor information to acquire information about a target object, such as a person, and recognize the behavior of the object based on the acquired information. Among these, video-based behavior recognition technologies generally capture an image of the object with a camera, transmit the acquired video data to a computer, such as a server computer, and recognize the behavior of the object based on the video data. On the other hand, behavior recognition technologies that use sensor information acquired by, for example, an acceleration sensor, perform some processing on the sensor information on the sensor terminal and then transmit the information to a computer, which recognizes the behavior of the person wearing the sensor (see, for example, Non-Patent Document 1).
[0003] Huang X, Yuan Y, Chang C, Gao Y, Zheng C, Yan L. “Human Activity Recognition Method Based on Edge Computing-Assisted and GRU Deep Learning Network.” Applied Sciences. 2023; 13(16):9059.
[0004] However, when trying to recognize the behavior of an object such as a person in real time, it is necessary to transmit video data captured by, for example, a camera to a computer such as a GPU server without delay. However, in order to achieve this, it is necessary to ensure a network environment that can transmit video data with a large amount of information without delay, and an information processing environment that can perform all calculations related to behavior recognition processing on a single computer.
[0005] This invention has been made in light of the above circumstances, and aims to provide a technology that reduces the amount of data transmitted over a network and the processing load in information processing, thereby enabling smooth behavior recognition in real time.
[0006] In order to solve the above problems, one aspect of the behavior recognition system according to the present invention is to recognize the behavior of an object to be recognized by an edge device, which acquires video data capturing an area including the object, estimates a skeleton of the object from the video data, and transmits skeleton information representing the estimation result to the behavior recognition device, which then receives the skeleton information, recognizes the behavior of the object based on the received skeleton information, and outputs information representing the behavior recognition result.
[0007] According to one aspect of the present invention, skeletal information is transmitted from the edge device to the behavior recognition device, which reduces the amount of transmitted data compared to when video data is transmitted, thereby reducing the network transmission load and transmission delay, and enabling accurate behavior recognition while maintaining real-time performance.
[0008] That is, according to one aspect of the present invention, it is possible to provide a technology that reduces the amount of data transmitted over a network and the processing load in information processing, thereby enabling smooth behavior recognition in real time.
[0009] FIG. 1 is a diagram showing an example of the overall configuration of a behavior recognition system according to a first embodiment of the present invention. FIG. 2 is a block diagram showing an example of the hardware configuration of a network camera device according to the first embodiment of the present invention. FIG. 3 is a block diagram showing an example of the software configuration of a network camera device according to the first embodiment of the present invention. FIG. 4 is a block diagram showing an example of the hardware configuration of a behavior recognition device according to the first embodiment of the present invention. FIG. 5 is a block diagram showing an example of the software configuration of a behavior recognition device according to the first embodiment of the present invention. FIG. 6 is a flowchart showing an example of the processing procedure and processing content of a first example of skeleton estimation processing executed by the network camera device shown in FIG. 3. FIG. 7 is a flowchart showing an example of the processing procedure and processing content of a second example of skeleton estimation processing executed by the network camera device shown in FIG. 3. FIG. 8 is a flowchart showing an example of the processing procedure and processing content of skeleton information counting processing executed by the behavior recognition device shown in FIG. 5. FIG. 9 is a flowchart showing an example of the processing procedure and processing content of skeleton-behavior estimation processing executed by the behavior recognition device shown in FIG. 5. FIG. 10 is a block diagram showing an example of the software configuration of a network camera device according to a second embodiment of the present invention. FIG. 11 is a block diagram showing an example of the software configuration of a behavior recognition device according to the second embodiment of the present invention. FIG. 12 is a flowchart showing an example of the processing procedure and processing content of the skeleton information encoding processing executed by the network camera device shown in FIG. 10. FIG. 13 is a flowchart showing an example of the processing procedure and processing content of the skeleton information encoding processing executed by the behavior recognition device shown in FIG. 11. FIG. 14 is a block diagram showing an example of the software configuration of the behavior recognition learning device provided in the behavior recognition system according to the third embodiment of the present invention. FIG. 15 is a flowchart showing an example of the processing procedure and processing content of the learning processing executed by the skeleton transformation learning unit of the behavior recognition learning device shown in FIG. 14. FIG. 16 is a flowchart showing an example of the processing procedure and processing content of the learning processing executed by the skeleton estimation learning unit of the behavior recognition learning device shown in FIG. 14. FIG. 17 is a flowchart showing an example of the processing procedure and processing content of the learning processing executed by the skeleton behavior recognition learning unit of the behavior recognition learning device shown in FIG. 14.
[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0011] First Embodiment (Configuration Example) (1) System FIG. 1 is a diagram showing an example of the overall configuration of a behavior recognition system according to a first embodiment of the present invention.
[0012] The behavior recognition system according to the first embodiment is composed of a plurality of network camera devices CM11 to CM1n arranged as edge devices, and a behavior recognition device SV1 capable of communicating and transmitting data between each of the plurality of network camera devices CM11 to CM1n via a network NW.
[0013] The network camera devices CM11 to CM1n are installed in a space where a person, for example, a user US1 to USn, who is an object to be recognized, exists. Note that the edge device may be an information processing terminal such as a personal computer used by a user, a router, a switch, a gateway, or the like, in addition to the network camera devices CM11 to CM1n.
[0014] The network NW is composed of a wide area network such as the Internet and an access network for accessing this wide area network. The access network may be a wired or wireless local area network (LAN), an optical transmission network, or a mobile communication network that adopts the 5G standard.
[0015] (2) Network Camera Devices CM11 to CM1n FIGS. 2 and 3 are block diagrams showing examples of the hardware and software configurations of the network camera devices CM11 to CM1n, respectively.
[0016] Each of the network camera devices CM11 to CM1n includes a camera 6 as an imaging device and a control unit 1A that uses a hardware processor such as a graphics processing unit (CPU) and has the information processing functions of a personal computer. Each of the network camera devices CM11 to CM1n has a storage unit including a program storage unit 2A and a data storage unit 3A, a communication interface (hereinafter, the interface will be abbreviated as I / F) unit 4A, and a sensor I / F unit 5A connected to the control unit 1A via a bus.
[0017] The communication I / F unit 4A transmits and receives information data to and from a behavior recognition device SV1, which will be described later, in accordance with a communication protocol defined by the network NW.
[0018] The sensor I / F unit 5A is connected to a camera 6. The sensor I / F unit 5A takes in video data output from the camera CM.
[0019] The program storage unit 2A is, for example, a combination of a non-volatile memory such as a HDD (Hard Disk Drive) or SSD (Solid State Drive) as a storage medium that can be written to and read from at any time, and a non-volatile memory such as a ROM (Read Only Memory), and stores application programs necessary for executing various processes related to the first embodiment of the present invention, in addition to middleware such as an OS (Operating System).
[0020] The data storage unit 3A is, for example, a combination of a nonvolatile memory such as an HDD or SSD as a storage medium that can be written to and read from at any time, and a volatile memory such as a RAM (Random Access Memory), and its storage area includes a video data storage unit 31A, a person information storage unit 32A, a skeleton estimation model storage unit 33A, and a skeleton information storage unit 34A. The storage area also includes a buffer area for temporarily storing data when the control unit 1A executes various processes.
[0021] The control unit 1A has the control functions necessary to realize the first embodiment of the present invention, including a video data acquisition processing unit 11A, a person detection processing unit 12A, a skeleton estimation processing unit 13A, and a skeleton information transmission processing unit 14A.
[0022] Each of the processing units 11A to 14A is realized by causing the processor of the control unit 1A to execute an application program stored in the program storage unit 2A. Note that some or all of the processing units 11A to 14A may be realized using hardware such as an LSI (Large Scale Integration) or an ASIC (Application Specific Integrated Circuit).
[0023] The video data acquisition processing unit 11A captures video data of the space captured by the camera CM via the sensor I / F unit 5A, samples the captured video data at a predetermined frame period, and stores the sampled frame image data in chronological order in the video data storage unit 31A.
[0024] The person detection processing unit 12A reads frame image data from the video data storage unit 31A and detects users US1 to USn to be recognized from the read frame image data.The person detection processing unit 12A then stores coordinate information representing the areas in the frame image data of the detected users US1 to USn as person information in the person information storage unit 32A, in association with user identification information (user ID).An example of the person detection processing will be described in the operation example.
[0025] The skeleton estimation processing unit 13A cuts out image data from the area represented by the coordinate information for each detected user US1 to USn based on the person information stored in the person information storage unit 32A, and estimates the user's skeleton from the cut-out image data.The skeleton estimation processing unit 13A then stores the estimated skeleton information representing the user's skeleton in the skeleton information storage unit 34A in association with the user ID.The skeleton estimation processing is performed using, for example, a trained skeleton estimation model stored in the skeleton estimation model storage unit 33A, an example of which will be described in the operation example.
[0026] The skeletal information transmission processing unit 14A reads out the skeletal information of users US1 to USn from the skeletal information storage unit 34A, and transmits the read-out skeletal information of users US1 to USn together with the identification information of the edge device (also called terminal ID) from the communication I / F unit 4A to the behavior recognition device SV1.
[0027] (3) Behavior Recognition Device SV1 FIGS. 4 and 5 are block diagrams showing an example of the hardware configuration and software configuration of the behavior recognition device SV1, respectively.
[0028] The behavior recognition device SV1 consists of a server computer located, for example, on the web or cloud, and is configured by connecting a processor 1B constituting a control unit (GPU) to a storage unit having a program storage unit 2B and a data storage unit 3B, and a communication I / F unit 4B via a bus.
[0029] The communication I / F unit 4B receives information data transmitted from the network camera devices CM11 to CM1n in accordance with a communication protocol defined by the network NW.
[0030] Like the network camera devices CM11 to CM1n, the program storage unit 2B is a combination of a non-volatile memory such as an HDD or SSD that can be written to and read from at any time, and a non-volatile memory such as a ROM, and stores application programs necessary to execute various processes related to the first embodiment of this invention, in addition to middleware such as an OS.
[0031] The data storage unit 3B is, for example, a combination of a non-volatile memory such as an HDD or SSD as a storage medium that can be written to and read from at any time, and a volatile memory such as a RAM, and its storage area is provided with a skeleton information storage unit 31B, a behavior recognition model storage unit 32B, and a behavior recognition information storage unit 33B. The storage area also includes a buffer area for temporarily storing data when the control unit 1B executes various processes.
[0032] The control unit 1B includes a skeletal information counting processing unit 11B, a skeletal behavior recognition processing unit 12B, and a behavior recognition information transmission processing unit 13B as control functions for realizing the first embodiment of the present invention.
[0033] Each of the processing units 11B to 13B is realized by causing the processor of the control unit 1B to execute an application program stored in the program storage unit 2B. Note that some or all of the processing units 11B to 14B may be realized using hardware such as an LSI (Large Scale Integration) or an ASIC.
[0034] The skeleton information aggregation processing unit 11B receives skeleton information transmitted from each network camera device CM11 to CM1n via the communication I / F unit 4B, and stores each received skeleton information in the skeleton information storage unit 31B in association with the terminal ID of the edge device that transmitted it.
[0035] The skeletal behavior recognition processing unit 12B reads skeletal information for each terminal ID of the edge device from the skeletal information storage unit 31B and executes a process of recognizing the behavior of users US1 to USn based on the read skeletal information.The skeletal behavior recognition processing unit 12B then stores information indicating the recognition results of the behavior of users US1 to USn as behavior recognition information in the behavior recognition information storage unit 33B in association with the terminal ID and user ID.The behavior recognition processing is performed using a trained behavior recognition model stored in the behavior recognition model storage unit 32B, an example of which will be described in the operation example.
[0036] The behavior recognition information transmission processing unit 13B reads out behavior recognition information for each terminal ID of the edge device from the behavior recognition information storage unit 33B, and transmits the read out behavior recognition information from the communication I / F unit 4B to, for example, an administrator terminal (not shown).
[0037] (Example of Operation) Next, an example of operation of the behavior recognition system configured as above will be described.
[0038] (1) Operation of Network Camera Devices CM11 to CM1n Since the operation of each of the network camera devices CM11 to CM1n is the same, the network camera device CM11 will be taken as an example for explanation.
[0039] First Example FIG. 6 is a flowchart showing a first example of the processing procedure and processing content of a series of processes executed by the control unit 1A of the network camera device CM11 to estimate skeletal information.
[0040] (1-1) Acquisition of video data First, in step S11, the control unit 1A of the network camera device CM11, under the control of the video data acquisition processing unit 11A, acquires the video data output from the camera 6 via the sensor I / F unit 5A. The video data acquisition processing unit 11A then samples the acquired video data at a predetermined frame period and stores the sampled frame image data in chronological order in the video data storage unit 31A in association with the terminal ID of the edge device.
[0041] (1-2) Person Detection Next, in step S12, the control unit 1A of the network camera device CM11, under the control of the person detection processing unit 12A, sequentially reads frame image data from the video data storage unit 31A, and detects the user US1 to be recognized from each of the read frame image data.
[0042] This person detection can be performed using, for example, MobileNet v1, which uses depthwise separable convolutions (a general term for convolutions that combine depthwise convolution and pointwise convolution (1x1 convolution)), SSD (Single Shot MultiBox Detector), and NMS (Non-Maximum Suppression). In this example, the person detection result is output as a set of coordinate information (bbox coordinates) that represents the area in the frame image data of user US1 and information (user ID) for identifying the detected user.
[0043] The person detection processing unit 12A then stores coordinate information representing the area of the detected user US1 in the frame image data in association with the user ID in the person information storage unit 32A as person information.
[0044] In addition, if multiple users who are the target of behavior recognition are detected from one frame of image data, personal information consisting of a pair of coordinate information representing the area of each user and a user ID is generated for each user and stored in the personal information storage unit 32A.
[0045] (1-3) Skeleton Estimation The control unit 1A of the network camera device CM11 estimates the user's skeleton as follows under the control of the skeleton estimation processing unit 13A.
[0046] That is, the skeleton estimation processing unit 13A selects one of the detected users from the person information storage unit 32A, reads coordinate information (bbox coordinates) representing an area in the frame image of the selected user, cuts out image data from the area represented by the read coordinate information, and estimates the user's skeleton based on this image.
[0047] For example, the skeleton estimation processing unit 13A inputs an RGB image cut out from the bbox coordinate area into a skeleton estimation model and receives skeleton estimation information from the skeleton estimation model. For example, HRNet is used as the skeleton estimation model, which outputs the x and y coordinates of multiple predefined body parts, such as the nose and shoulders, for example, 17 locations. Note that any method may be used for skeleton estimation as long as it is a top-down type method.
[0048] The skeleton estimation processing unit 13A repeats the above skeleton estimation process for each user, and stores the obtained skeleton information of each user in the skeleton information storage unit 34A with the user ID attached.
[0049] (1-4) Transmission of Skeleton Information When the skeleton estimation process is completed, the control unit 1A of the network camera device CM11 proceeds to step S15. Then, under the control of the skeleton information transmission processing unit 14A, the control unit 1A reads out the skeleton information from the skeleton information storage unit 34A and transmits the read out skeleton information from the communication I / F unit 4A to the behavior recognition device SV1.
[0050] Second Example FIG. 7 is a flowchart showing a second example of the processing procedure and processing content of a series of processes executed by the control unit 1A of the network camera device CM11 to estimate skeletal information.
[0051] (1-1) Acquisition of video data First, in step S21, the control unit 1A of the network camera device CM11, under the control of the video data acquisition processing unit 11A, acquires the video data output from the camera 6 via the sensor I / F unit 5A. The video data acquisition processing unit 11A then samples the acquired video data at a predetermined frame period and stores the sampled frame image data in chronological order in the video data storage unit 31A in association with the terminal ID of the edge device.
[0052] (1-2) Skeleton Estimation Next, in step S22, the control unit 1A of the network camera device CM11, under the control of the skeleton estimation processing unit 13A, reads frame image data from the video data storage unit 31A and estimates the skeleton of the user US1 to be recognized from each of the read frame image data.
[0053] For example, the skeleton estimation processing unit 13A uses Higher HRNet or Open Pose as a skeleton estimation model, inputs the frame image data to this skeleton estimation model, and receives skeleton estimation results output from the skeleton estimation model. This allows obtaining the x and y coordinates of each of multiple predefined body parts, such as the nose and shoulders, of the user appearing in the frame image. Note that any skeleton estimation method can be used as long as it is a bottom-up type method.
[0054] The skeleton estimation processing unit 13A generates skeleton information for each user by associating the user ID with coordinate information representing the estimated skeleton, and stores this skeleton information in the skeleton information storage unit 34A.
[0055] (1-3) Transmission of Skeleton Information When the skeleton estimation process is completed, the control unit 1A of the network camera device CM11 proceeds to step S23. Then, under the control of the skeleton information transmission processing unit 14A, the control unit 1A reads out the skeleton information from the skeleton information storage unit 34A and transmits the read out skeleton information from the communication I / F unit 4A to the behavior recognition device SV1.
[0056] (2) Operation of the Behavior Recognition Device SV1 On the other hand, the control unit 1B of the behavior recognition device SV1 executes the process of recognizing the user's behavior based on the skeleton information transmitted from the network camera devices CM11 to CM1n as follows.
[0057] (2-1) Counting of Skeleton Information FIG. 8 is a flowchart showing an example of the processing procedure and processing content of the skeletal information counting process executed by the control unit 1B of the behavior recognition device SV1.
[0058] First, in step S31, the control unit 1B of the behavior recognition device SV1 receives the skeletal information transmitted from each of the network camera devices CM11 to CM1n via the communication I / F unit 4B under the control of the skeletal information compilation processing unit 11B.
[0059] Next, in step S32, the skeletal information aggregation processing unit 11B selects the skeletal information received for each terminal ID of the edge device, and stores the skeletal coordinate information of each user included in the selected skeletal information in the skeletal information storage unit 31B in step S33.
[0060] The skeletal information compilation processing unit 11B then determines in step S34 whether the accumulated number of pieces of skeletal coordinate information for each user has reached a preset threshold. If the accumulated number has not reached the threshold, the skeletal information compilation processing unit 11B repeatedly executes the skeletal coordinate information accumulation process. On the other hand, if the accumulated number has reached the threshold, the skeletal information compilation processing unit 11B reads out the accumulated set of skeletal coordinate information for each user from the skeletal information storage unit 31B in step S35 and outputs it to the skeletal behavior recognition processing unit 12B. After the output, the skeletal information compilation processing unit 11B discards the read set of skeletal information from the skeletal information storage unit 31B in step S36.
[0061] (2-2) Behavior Recognition Using Skeletons When the control unit 1B of the behavior recognition device SV1 receives a set of skeletal coordinate information of each user from the skeletal information aggregation processing unit 11B, it executes the process of recognizing the behavior of each user under the control of the skeletal behavior recognition processing unit 12B as follows.
[0062] FIG. 9 is a flowchart showing an example of the processing procedure and processing content of the behavior recognition processing using a skeleton, which is executed by the control unit 1B of the behavior recognition device SV1.
[0063] That is, when the skeletal behavior recognition processing unit 12B receives a set of skeletal coordinate information of each user from the skeletal information counting processing unit 11B in step S41, it selects one user from among the multiple users in step S42. Then, in step S43, the skeletal behavior recognition processing unit 12B recognizes the behavior of the selected user based on the skeletal coordinate information.
[0064] This behavior recognition process is performed by inputting the skeletal coordinate information into a trained behavior recognition model stored in the behavior recognition model storage unit 32B and receiving information representing the behavior recognition result output from the behavior recognition model. For example, an Efficient Graph Convolutional Neural Network is used as the behavior recognition model, and the skeletal coordinate information is input into this neural network to obtain a behavior label representing the user's behavior recognition result. Note that any method may be used for the behavior recognition process as long as it is a method for recognizing a person's behavior based on their skeletal information.
[0065] The skeletal behavior recognition processing unit 12B stores the behavior recognition information read from the communication I / F unit 4B, for example, as recognition information in the behavior recognition information storage unit 33B, while associating a behavior label representing the recognition result of the behavior with the user ID.
[0066] The skeleton behavior recognition processing unit 12B repeatedly executes the behavior recognition process based on the skeleton coordinate information in steps S42 and S43 for each user.
[0067] (2-3) Transmission of Behavior Recognition Information When the behavior recognition process for all users is completed, in step S44, the control unit 1B of the behavior recognition device SV1 reads out the behavior recognition information of each user from the behavior recognition information storage unit 33B under the control of the behavior recognition information transmission processing unit 13B. Then, the behavior recognition information transmission processing unit 13B transmits the read out behavior recognition information from the communication I / F unit 4B to, for example, an administrator terminal used by the user's administrator.
[0068] (Effects) As described above, in the first embodiment, the network camera devices CM11 to CM1n serving as edge devices detect users US1 to USn from video data captured by the camera 6, estimate the skeletons of users US1 to USn from an image of an area including users US1 to USn using a skeleton estimation model, and transmit the estimated skeleton information of users US1 to USn to the behavior recognition device SV1 via the network NW. In response to this, the behavior recognition device SV1 recognizes the behavior of the users using the behavior recognition model based on the skeleton information transmitted from the network camera devices CM11 to CM1n, and outputs the obtained behavior recognition information.
[0069] Therefore, the amount of data sent from the network camera devices CM11 to CM1n to the behavior recognition device SV1 is reduced compared to when the video data is transmitted as is, thereby reducing the transmission load and transmission delay on the network NW. Also, the processing load can be reduced compared to when the behavior recognition device SV1 recognizes user behavior from video data, and this, combined with the reduction in the transmission load and transmission delay on the network NW, enables accurate behavior recognition while maintaining real-time performance.
[0070] [Second embodiment] In the second embodiment of the present invention, a network camera device serving as an edge device encodes skeletal information when transmitting the skeletal information, and the behavior recognition device SV1 decodes the skeletal information transmitted from the network camera device and then performs behavior recognition processing.
[0071] (Configuration Example) (1) Network Camera Devices CM21 to CM2n Figure 10 is a block diagram showing an example of the software configuration of the network camera devices CM21 to CM2n according to the second embodiment of the present invention. In Figure 10, the same parts as in Figure 3 are given the same reference numerals, and detailed explanations will be omitted. In addition, the hardware configuration of the network camera devices CM21 to CM2n is the same as in Figure 2, so it will not be shown in the figure.
[0072] The control unit 10A of each of the network camera devices CM21 to CM2n includes a video data acquisition processing unit 11A, a person detection processing unit 12A, a skeleton estimation processing unit 13A, a skeleton encoding processing unit 15A, and an encoded data transmission processing unit 16A.
[0073] The skeleton encoding processing unit 15A encodes the skeleton information read from the skeleton information storage unit 34A.
[0074] The encoded data transmission processing unit 16A transmits the skeleton information including the encoded skeleton coordinate information from the communication I / F unit 4A to the behavior recognition device SV2.
[0075] (2) Behavior recognition device SV2 Fig. 11 is a block diagram showing an example of the software configuration of the behavior recognition device SV2 according to the second embodiment of the present invention. In Fig. 11, the same parts as those in Fig. 5 are designated by the same reference numerals, and detailed explanations thereof will be omitted. In addition, the hardware configuration of the behavior recognition device SV2 is the same as that in Fig. 4, and therefore will not be shown.
[0076] The data storage unit 30B of the behavior recognition device SV2 includes an encoded data storage unit 34B in addition to a skeleton information storage unit 31B, a behavior recognition model storage unit 32B, and a behavior recognition information storage unit 33B.
[0077] The control unit 10B of the behavior recognition device SV2 includes an encoded data receiving processing unit 14B and a skeleton decoding processing unit 15B in addition to a skeleton behavior recognition processing unit 12B and a behavior recognition information transmitting processing unit 13B.
[0078] The coded data receiving processing unit 14B receives the coded skeleton information transmitted from the network camera devices CM21 to CM2n via the communication I / F unit 4B, and stores the received coded skeleton information in the coded data storage unit 34B.
[0079] The skeleton decoding processing unit 15B reads and decodes the encoded skeleton information from the encoded data storage unit 34B, and stores the skeleton information reproduced by this decoding in the skeleton information storage unit 31B.
[0080] (Example of Operation) (1) Operation of Network Camera Devices CM21 to CM2n Since the operation of the network camera devices CM21 to CM2n is the same, the network camera device CM21 will be taken as an example for explanation, as in the first embodiment.
[0081] (1-1) Encoding of Skeleton Information When the control unit 10A of the network camera device CM21 completes the estimation process of the skeleton coordinate information of each user under the control of the skeleton estimation processing unit 13A, it then executes the encoding process of the skeleton information as follows under the control of the skeleton encoding processing unit 15A.
[0082] FIG. 12 is a flowchart showing an example of the processing procedure and processing contents of the skeleton encoding processing executed in the skeleton encoding processing unit 15A.
[0083] That is, first, in step S51, the skeleton encoding processing unit 15A reads the skeleton coordinate information included in the skeleton information for each user from the skeleton information storage unit 34A. Then, in steps S52 and S53, the skeleton encoding processing unit 15A encodes the skeleton coordinate information for each user.
[0084] For the encoding process, for example, the encoder function of an autoencoder that has learned to compress and decompress skeletal coordinate information can be used. Note that the encoding of skeletal coordinate information is not limited to the above autoencoder, and any other encoder may be used.
[0085] The skeleton encoding processing unit 15A repeatedly executes the encoding process of skeleton coordinate information for all detected users in steps S52 and S53. Then, when the encoding process of skeleton information for all users is completed, the skeleton encoding processing unit 15A passes the encoded skeleton information of each user to the encoded data transmission processing unit 16A in step S54.
[0086] (1-2) Transmission of Encoded Skeleton Information When the encoding process of the skeletal information is completed, the control unit 10A of the network camera device CM21, under the control of the encoded data transmission processing unit 16A, transmits the encoded skeletal information passed from the skeletal encoding processing unit 15A from the communication I / F unit 4A to the behavior recognition device SV2.
[0087] (2) Operation of the behavior recognition device SV2 (2-1) Receiving encoded skeleton information The control unit 10B of the behavior recognition device SV2 receives the encoded skeleton information transmitted from the network camera devices CM11 to CM1 n via the communication I / F unit 4B under the control of the encoded data reception processing unit 14B, and temporarily stores the received encoded skeleton information in the encoded data storage unit 34B.
[0088] (2-2) Decoding of Encoded Skeleton Information When the control unit 10B of the behavior recognition device SV2 receives the encoded skeleton information, the control unit 10B decodes the encoded skeleton information as follows.
[0089] FIG. 13 is a flowchart showing an example of the processing procedure and processing content of the decoding process of the encoded skeleton information executed by the control unit 10B of the behavior recognition device SV2.
[0090] That is, the skeleton decoding processing unit 15B first reads the encoded skeleton information from the encoded data storage unit 34B in step S61. Then, in steps S62 and S63, the skeleton decoding processing unit 15B selects one user from each user and performs a decoding process on the encoded skeleton information of the selected user. The decoding process is performed, for example, using the decoder function of an autoencoder that has learned the process of compressing and restoring skeleton coordinate information. Note that the decoding process may be performed using a decoder other than the autoencoder.
[0091] The skeleton decoding processing unit 15B repeatedly executes the above-mentioned decoding process in steps S62 and S63 each time a user is selected.
[0092] The skeleton decoding processing unit 15B stores the skeleton information restored by the decoding process for each user in the skeleton information storage unit 31B in association with the user ID in step S64.
[0093] Thereafter, as in the first embodiment, the control unit 10B of the behavior recognition device SV2 executes a process of recognizing the user's behavior based on the above-mentioned skeletal information for each user under the control of the skeletal behavior recognition processing unit 12B, and transmits behavior recognition information indicating the recognition results to an administrator terminal or the like under the control of the behavior recognition information transmission processing unit 13B.
[0094] (Effect) As described above, in the second embodiment of the present invention, in the network camera devices CM21 to CM2n, the skeleton information of the user estimated by the skeleton estimation process is encoded by the skeleton encoding processing unit 15A and then transmitted to the behavior recognition device SV2. After receiving the encoded skeleton information sent from the network camera devices CM21 to CM2n, the behavior recognition device SV2 decodes it by the skeleton decoding processing unit 15B to restore the original skeleton information, and recognizes the user's behavior based on the restored skeleton information.
[0095] Therefore, since the skeleton information is transmitted in an encoded state, it is possible to increase the confidentiality of the skeleton information, thereby protecting the privacy of the user. In addition, since the skeleton information is compressed by the encoding process before transmission, the amount of data transmitted can be further reduced, thereby further reducing the transmission load on the network NW and reducing the impact of transmission delays.
[0096] [Third embodiment] The third embodiment of the present invention shows an example of a learning device and a learning method for respectively learning the learning model used by the network camera devices CM21 to CM2n for the skeleton estimation process and the skeleton encoding process in the second embodiment, and the learning model used by the behavior recognition device SV2 for the skeleton decoding process and the behavior recognition process.
[0097] FIG. 14 is a block diagram showing an example of the functional configuration of a behavior recognition learning device SVL according to the third embodiment of the present invention.
[0098] The behavior recognition learning device SVL is configured by a server computer similar to the behavior recognition device SV2 described above, and its hardware configuration is the same as that of the behavior recognition device SV2, so it is not shown in the drawings.
[0099] The behavior recognition learning device SVL includes a skeleton transformation learning unit 101 , a skeleton estimation learning unit 102 , and a skeleton behavior recognition learning unit 103 .
[0100] The skeleton transformation learning unit 101 updates the learning parameters of the learning model used in the skeleton encoding processing unit 15A of the network camera devices CM21 to CM2n shown in Figure 10 and the skeleton decoding processing unit 15B of the behavior recognition device SV2 shown in Figure 11.
[0101] The skeleton estimation learning unit 102 updates the learning parameters of the learning model used in the skeleton estimation processing unit 13A and the skeleton encoding processing unit 15A of the network camera devices CM21 to CM2n.
[0102] The skeleton behavior recognition learning unit 103 updates the learning parameters of the learning model used in the skeleton decoding processing unit 15B and the skeleton behavior recognition processing unit 12B of the behavior recognition device SV2.
[0103] (Example of Operation) (1) Operation of the Skeleton Transformation Learning Unit 101 FIG. 15 is a flowchart showing an example of the processing procedure and processing content of the learning parameter update processing executed by the skeleton transformation learning unit 101.
[0104] Learning video data including a user to be recognized as a target of behavior recognition and correct answer information are input to the skeleton transformation learning unit 101. The correct answer information consists of a user ID, bbox coordinates, skeleton coordinates, and behavior labels.
[0105] When the skeleton transformation learning unit 101 acquires the video data and the skeleton coordinates included in the correct answer information in step S71, it first calculates a first loss based on the current learning parameters in step S72.
[0106] For example, when the skeleton encoding processing unit 15A and the skeleton decoding processing unit 15B are configured with the encoder function and decoder function of an autoencoder, respectively, the number of predefined skeleton point coordinates is N, the correct x-coordinate and y-coordinate of a certain skeleton point are (x, y), and the predicted x-coordinate and y-coordinate of a certain skeleton point after encoding and decoding are (x', y'), the first loss can be obtained by calculating the mean square error L1 using the following formula.
[0107]
[0108] The above-described first loss calculation method may be any method that can calculate a loss by learning the learning parameters of the skeleton encoding processing unit 15A and the skeleton decoding processing unit 15B.
[0109] Next, in step S73, the skeleton transformation learning unit 101 updates the learning parameters of the skeleton encoding processing unit 15A and the skeleton decoding processing unit 15B, and in step S74, determines whether the loss due to the updated learning parameters has decreased to below a predetermined threshold, i.e., whether convergence has occurred.
[0110] If the result of this determination is that convergence has not occurred, the skeleton transformation learning unit 101 repeatedly executes the learning parameter update process in steps S72 to S74. If it is determined that convergence has occurred, the skeleton transformation learning unit 101 ends the learning parameter update process and saves the final learning parameters.
[0111] (2) Operation of the Skeleton Estimation Learning Unit 102 FIG. 16 is a flowchart showing an example of the processing procedure and processing content of the learning parameter update processing executed by the skeleton estimation learning unit 102.
[0112] The skeleton estimation learning unit 102 receives as input the video data, the user image cut out in the bbox area prepared as correct answer information, skeleton coordinates, and the encoded skeleton coordinates output from the learned skeleton encoding processing unit 15A.
[0113] In step S81, the skeleton estimation learning unit 102 acquires the user image, skeleton coordinates, and encoded skeleton coordinates, and in step S82, calculates a second loss based on the current learning parameters.
[0114] For example, let N be the number of predefined skeleton point coordinates, M be the number of skeleton point coordinates after encoding, and let (x r , y r ), the correct values of the x and y coordinates after encoding at a skeleton point are (x e , y e ), the x-coordinate and y-coordinate of the predicted coordinates of a certain skeleton point output from the skeleton estimation processing unit 13A are (x r ', y r'), the x and y coordinates of the predicted coordinates at a certain skeleton point output from the skeleton encoding processing unit 15A are (x e ', y e '), the second loss can be obtained by calculating the weighted mean square error L using the following formula: r and γ e is a hyperparameter that adjusts the influence of each term.
[0115]
[0116] The second loss calculation method may be any method that can learn the learning parameters of both or either of the skeleton estimation processing unit 13A and the skeleton encoding processing unit 15A.
[0117] Next, in step S83, the skeleton estimation learning unit 102 updates the learning parameters of the skeleton estimation processing unit 13A and the skeleton encoding processing unit 15A, and in step S84, determines whether the loss due to the updated learning parameters has decreased to below a predetermined threshold, i.e., whether convergence has occurred.
[0118] If the result of this determination is that convergence has not occurred, the skeleton estimation learning unit 102 repeatedly executes the learning parameter update process in steps S82 to S84. If it is determined that convergence has occurred, the skeleton estimation learning unit 102 ends the learning parameter update process and saves the final learning parameters.
[0119] (3) Operation of the Skeleton Behavior Recognition Learning Unit 103 FIG. 17 is a flowchart showing an example of the processing procedure and processing content of the learning parameter update processing executed by the skeleton behavior recognition learning unit 103.
[0120] The skeleton behavior recognition learning unit 103 receives as input the video data, the encoded skeleton coordinates output from the trained skeleton encoding processing unit 15A, which are prepared as correct answer information, the skeleton coordinates, and the correct answer label.
[0121] In step S91, the skeleton behavior recognition learning unit 103 acquires the encoded skeleton coordinates, skeleton coordinates, and correct label, and in step S92, calculates a third loss based on the current learning parameters.
[0122] For example, let N be the number of predefined skeleton point coordinates, T be the number of accumulated skeleton coordinate points, (x, y) be the correct x- and y-coordinates k at a certain skeleton point, (x', y') be the x- and y-coordinates of the predicted coordinates after encoding and decoding at a certain skeleton point, a be the correct action label distribution, and a' be the predicted action label distribution. Then, the third loss can be obtained by calculating the error L3, which is expressed as a weighted sum of the mean square error and cross entropy, using the following formula:
[0123]
[0124] The third loss calculation method may be any method that can learn the learning parameters of both or either of the skeletal decoding processing unit 15B and the skeletal behavior recognition processing unit 12B.
[0125] Next, in step S93, the skeletal behavior recognition learning unit 103 updates the learning parameters of the skeletal decoding processing unit 15B and the skeletal behavior recognition processing unit 12B, and in step S94, determines whether the loss due to the updated learning parameters has decreased to below a preset threshold, that is, whether convergence has occurred.
[0126] If the result of this determination is that convergence has not occurred, the skeletal behavior recognition learning unit 103 repeatedly executes the learning parameter update process in steps S92 to S94. If it is determined that convergence has occurred, the skeletal behavior recognition learning unit 103 ends the learning parameter update process and saves the final learning parameters.
[0127] Finally, the behavior recognition learning device SVL transfers the learning parameters obtained by the skeleton transformation learning unit 101 and the skeleton estimation learning unit 102 to the network camera devices CM21 to CM2n, and sets them in the skeleton estimation model storage unit 33A and the skeleton encoding processing unit 15A.
[0128] At the same time, the behavior recognition learning device SVL transfers the learning parameters obtained by the skeleton behavior recognition learning unit 103 to the behavior recognition device SV2 and sets them in the skeleton decoding processing unit 15B and the behavior recognition model storage unit 32B.
[0129] (Effect) As described above, according to the third embodiment of the present invention, it is possible to generate the learning parameters of each learning model used in the network camera devices CM21 to CM2n and the behavior recognition device SV2 collectively in the behavior recognition learning device SVL.
[0130] [Other Embodiments] (1) In the first embodiment, the case where there are multiple edge devices has been described as an example, but there may be only one. Also, the case where there are multiple objects (users) as behavior recognition targets in each edge device has been described as an example, but there may be only one. Furthermore, the object is not limited to a person, but may also be an animal or a machine such as a robot.
[0131] (2) In the third embodiment, an example has been described in which an independently provided behavior recognition learning device SVL generates learning parameters for a learning model used by the network camera devices CM21 to CM2n and the behavior recognition device SV2. However, the present invention is not limited to this, and the functional units 101 to 103 included in the behavior recognition learning device SVL may be included in the behavior recognition device SV2.
[0132] (3) In addition, the functional configurations of the edge device and the behavior recognition device, the processing procedures and processing contents of each process executed by each device, and the like can be modified and implemented in various ways without departing from the spirit of the present invention.
[0133] Although the embodiments of the present invention have been described in detail above, the above description is merely an example of the present invention in every respect. It goes without saying that various improvements and modifications can be made without departing from the scope of the present invention. In other words, when implementing the present invention, specific configurations according to the embodiments may be appropriately adopted.
[0134] In short, this invention is not limited to the above-described embodiments, and in the implementation stage, the components can be modified and embodied without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.
[0135] SV1, SV2... Behavior recognition device CM11 to NCM1n, CM21 to CM2n... Network camera device NW... Network US1 to USn... User to be recognized 1A, 1B, 10A, 10B... Control unit 2A, 2B... Program storage unit 3A, 3B, 30A, 30B... Data storage unit 4A, 4B... Communication I / F unit 5A... Sensor I / F unit 6A... Camera 11A... Video data acquisition processing unit 12A... Person detection processing unit 13A... Skeleton estimation processing unit 14A... Skeleton information transmission processing unit 15A... Skeleton encoding processing unit 16A... Encoded data transmission processing unit 11B... Skeleton information compilation processing unit 12B... Skeleton behavior recognition processing unit 13B... Behavior recognition information transmission processing unit 14B... Encoded data reception processing unit 15B... Skeleton decoding processing unit 31A... Video data storage unit 32A... Person information storage unit 33A... Skeleton estimation model storage unit 34A... Skeleton information storage unit 31B... Skeleton information storage unit 32B... Behavior recognition model storage unit 33B... Behavior recognition information storage unit SVL... Behavior recognition learning device 101... Skeleton transformation learning unit 102... Skeleton estimation learning unit 103... Skeleton behavior recognition learning unit
Claims
1. A behavior recognition system comprising: at least one edge device placed at a location where an object to be recognized is present; and a behavior recognition device capable of communicating with said edge device via a network, wherein said edge device comprises: a first processing unit that acquires video data capturing an area including said object; a second processing unit that estimates a skeleton of said object from said video data and outputs skeleton information representing the estimation result; and a third processing unit that transmits said skeleton information to said behavior recognition device, wherein said behavior recognition device comprises: a fourth processing unit that receives the skeleton information; a fifth processing unit that recognizes the behavior of said object based on the received skeleton information; and a sixth processing unit that outputs information representing the recognition result of said behavior.
2. The behavior recognition system according to claim 1, wherein the edge device further comprises an encoding processing unit that encodes the skeletal information before the skeletal information is transmitted by the third processing unit, and the behavior recognition device further comprises a decoding processing unit that decodes the skeletal information received by the fourth processing unit and provides the decoded skeletal information to the fifth processing unit.
3. The behavior recognition system of claim 2, further comprising a learning device, the learning device comprising: a skeleton transformation learning unit that learns learning parameters of a learning model used by the encoding processing unit and the decoding processing unit to perform encoding processing and decoding processing of the skeleton information; a skeleton estimation learning unit that learns learning parameters of a learning model used by the first processing unit and the encoding processing unit to perform the skeleton estimation processing and encoding processing of the skeleton information; and a skeleton behavior recognition learning unit that learns learning parameters of a learning model used by the decoding processing unit and the fifth processing unit to perform the decoding processing and processing to recognize the behavior.
4. A behavior recognition method executed by a system having at least one edge device placed at a location where an object to be recognized is present, and a behavior recognition device capable of communicating with the edge device via a network, the behavior recognition method comprising the steps of: the edge device acquiring video data capturing an area including the object; the edge device estimating a skeleton of the object from the video data and outputting skeleton information representing the estimation result; the edge device transmitting the skeleton information to the behavior recognition device; the behavior recognition device receiving the skeleton information; the behavior recognition device recognizing the behavior of the object based on the received skeleton information; and the behavior recognition device outputting information representing the recognition result of the behavior.
Citation Information
Patent Citations
Movement determination method, movement determination device, and movement determination program
JP2015114950A
Method for Processing Key Point Trajectory in Video
JP2018537880A
Posture analysis device, posture analysis method and program
JP2020052867A
Monitoring system, monitoring method, and method for training image recognition device for monitoring system
JP2024004972A
Information processing apparatus and control method of the same
JP2024048120A