Video processing system, video processing device, and video processing method

JPWO2024057469A5Active Publication Date: 2025-05-26NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024546614
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-05-26
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

Existing video processing systems cannot accurately recognize objects in videos when the frame rate changes, as they do not support frame rate adjustments, leading to suboptimal object recognition accuracy.

Method used

A video processing system that acquires input video and time difference information between frames, using a trained recognition model to improve object recognition accuracy by considering frame rate changes, employing a recurrent neural network (RNN) model with a state predictor to handle frame skipping and dynamic bit rate control.

Benefits of technology

Enhances object recognition accuracy in videos by dynamically adjusting to frame rate changes and frame skipping, improving the overall performance of video processing systems in recognizing objects and behaviors.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The objective of the present invention is to provide a video processing system, a video processing device, and a video processing method that can be expected to improve the recognition accuracy of objects in a video. A video processing system (10) comprises a video acquiring means (11), a time difference information acquiring means (12), and a recognizing means (13). The video acquiring means (11) acquires an input video. The time difference information acquiring means (12) acquires first time difference information between frames of the input video. The recognizing means (13) recognizes an object in the input video by inputting the input video and the first time difference information between the frames of the input video into a trained recognition model that has been trained using a training video and second time difference information between frames of the training video.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing system, image processing device, and image processing method

[0001] The present disclosure relates to a video processing system, a video processing device, and a video processing method.

[0002] A technology has been developed in which an edge device transmits captured video to a central server, which then uses an AI engine to recognize objects in the video. For example, the server recognizes the type of work being performed by a worker. The edge device then changes the frame rate of the video by frame filtering or other methods to ensure efficient use of computing resources and network bandwidth.

[0003] As a related technique, Patent Document 1 discloses a technique for performing video scene recognition from time-series frames extracted from a video using a deep learning algorithm such as a recurrent neural network (RNN).

[0004] Japanese Patent Application Laid-Open No. 2018-005638

[0005] In technologies such as those disclosed in Patent Document 1, the server at the center does not support changes in the frame rate of the video, so it is not possible to recognize objects that correspond to changes in the frame rate of the video, and there is room for improvement in the accuracy of recognizing objects in the video.

[0006] In consideration of such problems, the present disclosure aims to provide an image processing system, an image processing device, and an image processing method that are expected to improve the accuracy of recognizing objects in an image.

[0007] The video processing system of the present disclosure comprises: a video acquisition means for acquiring an input video; a time difference information acquisition means for acquiring first time difference information between frames of the input video; and a recognition means for inputting the input video and the first time difference information between frames of the input video into a trained recognition model trained using a training video and second time difference information between frames of the training video, and recognizing an object in the input video.

[0008] The video processing device of the present disclosure comprises: a video acquisition means for acquiring an input video; a time difference information acquisition means for acquiring first time difference information between frames of the input video; and a recognition means for inputting the input video and the first time difference information between frames of the input video into a trained recognition model trained using a training video and second time difference information between frames of the training video, and recognizing an object in the input video.

[0009] The video processing method disclosed herein includes a computer acquiring an input video, acquiring first time difference information between frames of the input video, inputting the input video and the first time difference information between frames of the input video into a trained recognition model trained using a training video and second time difference information between frames of the training video, and recognizing an object in the input video.

[0010] The present disclosure can provide an image processing system, an image processing device, and an image processing method that are expected to improve the accuracy of recognizing objects in an image.

[0011] 1 is a block diagram showing the configuration of a video processing system according to an overview of an embodiment. FIG. 2 is a block diagram showing the configuration of a video processing device according to an overview of an embodiment. FIG. 3 is a flowchart showing a video processing method according to an overview of an embodiment. FIG. 4 is a block diagram showing the configuration of a video processing system according to a first embodiment. FIG. 5 is a block diagram showing the configuration of a terminal according to the first embodiment. FIG. 6 is a block diagram showing the configuration of a center server according to the first embodiment. FIG. 7 is a flowchart showing the operation of the video processing system according to the first embodiment. FIG. 8 is a diagram showing an example of input information of a trained recognition model according to the first embodiment. FIG. 9 is a diagram showing an example of the configuration of a trained recognition model and recognition operation according to the first embodiment. FIG. 10 is a block diagram showing the configuration of a center server according to a second embodiment. FIG. 11 is a flowchart showing an example of the operation of the video processing system according to the second embodiment. FIG. 12 is a diagram showing an example of the configuration of a trained recognition model and recognition operation according to the second embodiment. FIG. 13 is a diagram showing an example of a first learning operation of a recognition model according to the second embodiment. FIG. 14 is a diagram showing another example of the configuration of a trained recognition model and recognition operation according to the second embodiment.

[0012] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, the same or corresponding elements are designated by the same reference numerals, and for clarity of explanation, duplicate explanations will be omitted as necessary.

[0013] (Outline of the embodiment) First, a video processing system 10 according to an outline of the embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the video processing system 10 according to the outline of the embodiment. The video processing system 10 can be applied to, for example, a remote monitoring system that collects video via a network and recognizes the video.

[0014] As shown in Fig. 1, video processing system 10 includes video acquisition unit 11, time difference information acquisition unit 12, and recognition unit 13. Video acquisition unit 11 acquires input video. Time difference information acquisition unit 12 acquires first time difference information between frames of the input video. Recognition unit 13 inputs the input video and the first time difference information between frames of the input video into a trained recognition model trained using the training video and second time difference information between frames of the training video, and recognizes objects in the input video.

[0015] Next, the configuration of the video processing device 20 according to the outline of the embodiment will be described with reference to FIG. 2. FIG. 2 is a block diagram showing the configuration of the video processing device 20 according to the outline of the embodiment. As shown in FIG. 2, the video processing device 20 includes the video acquisition unit 11, the time difference information acquisition unit 12, and the recognition unit 13 shown in FIG. 1. Furthermore, when the video processing device 20 is realized using edge computing, part or all of the video processing device 20 may be located on the edge or in the cloud. For example, the video acquisition unit 11 and the time difference information acquisition unit 12 may be located on an edge terminal, and the recognition unit 13 may be located on a cloud server. Furthermore, each function may be distributed in the cloud. Furthermore, the video processing device 20 may be realized using virtualization technology such as a virtualization server. Furthermore, part or all of the video processing device 20 may be located on the site side or on the server side. A device located at the site where a terminal is installed, a device located close to the site, or a device close to the terminal in terms of a network hierarchy is referred to as a device located on the site side. A device located away from the site is referred to as a device located on the center side. A device located on the center side may also be located on the cloud, and the center side may also be referred to as the cloud side.

[0016] Next, a video processing method according to an embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the video processing method according to the embodiment. For example, the video processing method according to the embodiment is executed by the video processing system 10 of Fig. 1 or the video processing device 20 of Fig. 2.

[0017] As shown in Fig. 3, an input video is acquired (step S11). Next, first time difference information between frames of the input video is acquired (step S12). Next, the input video and the first time difference information between frames of the input video are input to a trained recognition model trained using the training video and the second time difference information between frames of the training video, and objects in the input video are recognized (step S13).

[0018] As described above, the image processing system 10 can recognize objects in response to changes in the frame rate of the image by taking into account information about the time difference between frames of the input image. The image processing system 10 is expected to improve the accuracy of recognizing objects in the image.

[0019] (Basic Configuration of Video Processing System) Next, a video processing system 1, which is an example of a system to which an embodiment is applied, will be described with reference to FIG. 4 . FIG. 4 is a block diagram showing the configuration of the video processing system 1 according to the first embodiment. As shown in FIG. 4 , the video processing system 1 is a system that monitors an area captured by a camera using the captured video. In this embodiment, the system will be described as a system that remotely monitors the work of workers at a site. For example, the site may be an area where people and machines are operating, such as a work site such as a construction site, a public square where people gather, or a school. In this embodiment, the work will be described as construction work, civil engineering work, etc., but is not limited to this. Note that video includes time-series frames, which are images in a time series, so the terms video and image are interchangeable. In other words, the video processing system can be described as a video processing system that processes video and an image processing system that processes images.

[0020] The video processing system 1 includes a plurality of terminals 100, a center server 200, a base station 300, and an MEC 400. The terminals 100, the base station 300, and the MEC 400 are located on the site side, and the center server 200 is located on the center side. For example, the center server 200 is located in a data center or the like that is located away from the site. The site side is the edge side of the system, and the center side is also the cloud side.

[0021] The terminal 100 and the base station 300 are communicatively connected via a network NW1. The network NW1 is, for example, a wireless network such as 4G, local 5G / 5G, LTE (Long Term Evolution), or wireless LAN. The base station 300 and the center server 200 are communicatively connected via a network NW2. The network NW2 includes, for example, a core network such as 5GC (5th Generation Core network) or EPC (Evolved Packet Core), or the Internet. It can also be said that the terminal 100 and the center server 200 are communicatively connected via the base station 300. The base station 300 and the MEC 400 are communicatively connected via any communication method, but the base station 300 and the MEC 400 may be a single device.

[0022] The terminal 100 is a terminal device connected to the network NW1 and also serves as an image generating device that generates images of the site. The terminal 100 acquires images captured by a camera 101 installed at the site and transmits the acquired images to the center server 200 via the base station 300. The camera 101 may be located outside the terminal 100 or inside the terminal 100.

[0023] The terminal 100 compresses video from the camera 101 to a predetermined bit rate and transmits the compressed video. The terminal 100 has a compression efficiency optimization function 102 that optimizes compression efficiency, and a video distribution function 103. The compression efficiency optimization function 102 performs ROI control that controls the image quality of an ROI (Region of Interest; also called a fixation region). The compression efficiency optimization function 102 reduces the bit rate by maintaining the image quality of an ROI that includes a person or object while lowering the image quality of the surrounding area. The video distribution function 103 distributes the video with controlled image quality to the center server 200.

[0024] The base station 300 is a base station device of the network NW1, and also a relay device that relays communication between the terminal 100 and the center server 200. For example, the base station 300 is a local 5G base station, a 5G gNB (next generation Node B), an LTE eNB (evolved Node B), a wireless LAN access point, or the like, but may also be another relay device.

[0025] The MEC (Multi-access Edge Computing) 400 is an edge processing device located on the edge side of the system. The MEC 400 is an edge server that controls the terminal 100 and has a compression bit rate control function 401 and a terminal control function 402 that control the bit rate of the terminal. The compression bit rate control function 401 controls the bit rate of the terminal 100 through adaptive video distribution control and QoE (quality of experience) control. For example, the compression bit rate control function 401 predicts the recognition accuracy to be obtained while suppressing the bit rate according to the communication environment of the networks NW1 and NW2, and allocates a bit rate to the camera 101 of each terminal 100 so as to improve the recognition accuracy. The terminal control function 402 controls the terminal 100 to distribute video at the allocated bit rate. The terminal 100 encodes the video so as to obtain the allocated bit rate and distributes the encoded video.

[0026] The center server 200 is a server installed on the center side of the system. The center server 200 may be one or more physical servers, or may be a cloud server built on the cloud or other virtualized server. The center server 200 is a monitoring device that monitors on-site work by recognizing the work of people from on-site camera images. The center server 200 is also an image recognition device that recognizes the behavior of people in images transmitted from the terminal 100.

[0027] The center server 200 has an image recognition function 201, an alert generation function 202, a GUI drawing function 203, and a screen display function 204. The image recognition function 201 inputs the image transmitted from the terminal 100 into an AI engine (for example, a trained recognition model) to recognize the work performed by the worker, i.e., the type of human behavior. The alert generation function 202 generates an alert according to the recognized work. The GUI drawing function 203 displays a GUI (Graphical User Interface) on the screen of the display device. The screen display function 204 displays the image of the terminal 100, the recognition results, alerts, etc. on the GUI.

[0028] First Embodiment First, the configuration of a video processing system 1 according to a first embodiment will be described with reference to Fig. 4. As shown in Fig. 4, the video processing system 1 includes a plurality of terminals 100, a center server 200, a base station 300, and an MEC 400. Note that the configuration of each device is an example, and other configurations may be used as long as the operation according to this embodiment, which will be described later, is possible. For example, some of the functions of the terminal 100 may be arranged in the center server 200 or another device, or some of the functions of the center server 200 may be arranged in the terminal 100 or another device.

[0029] The video processing system 1 is a specific implementation of the video processing system 10 according to the outline of the embodiment. The center server 200 is a specific implementation of the video processing device 20 according to the outline of the embodiment.

[0030] Next, the configuration of the terminal 100 of the video processing system 1 according to the first embodiment will be described with reference to Fig. 5. Fig. 5 is a block diagram showing the configuration of the terminal 100 of the video processing system 1 according to the first embodiment. As shown in Fig. 5, the terminal 100 includes a video acquisition unit 110, a frame filter unit 120, an encoding unit 130, and a terminal communication unit 140.

[0031] The video acquisition unit 110 acquires video (also referred to as input video) captured by the camera 101. The input video is, for example, data capturing a person who is a worker performing work on-site, a work object used by the person, etc. The input video includes frames in a time series.

[0032] The frame filter unit 120 filters (selects) time-series frames included in the input video. The frame filter unit 120 performs filtering, for example, to adjust the bit rate of the video to be transmitted to the center server 200. Here, frames included in the input video that are not filtered are skipped.

[0033] The encoding unit 130 encodes the filtered input video. The encoding unit 130 changes the frame rate of the input video by filtering the frames. The encoding unit 130 may encode the input video so that the attention area of ​​the frame has higher image quality than other areas. Specifically, the encoding unit 130 detects objects in the input video using a trained neural network model (e.g., a model such as a convolutional neural network) and surrounds the detected objects with a box. The encoding unit 130 may surround the detected objects with shapes other than boxes, such as circles, ellipses, or irregular shapes that match a silhouette. The encoding unit 130 then recognizes the objects within the boxes. The encoding unit 130 extracts objects classified as people or work objects from the recognized objects and determines the area within the box of the extracted object as the attention area. The encoding unit 130 encodes the input video so that the attention area has higher image quality than other areas.

[0034] The terminal communication unit 140 transmits the encoded data to the center server 200 .

[0035] Next, the configuration of the center server 200 of the video processing system 1 according to the first embodiment will be described with reference to Fig. 6. Fig. 6 is a block diagram showing an example of the configuration of the center server 200 of the video processing system 1 according to the first embodiment. As shown in Fig. 6, the center server 200 includes a center communication unit 210, a decoding unit 220, a time difference information acquisition unit 230, a recognition unit 240, a storage unit 250, and a learning unit 260.

[0036] The center communication unit 210 receives the encoded data transmitted from the terminal 100 via the base station 300. The center communication unit 210 is an interface capable of communicating with the Internet or a core network, and is, for example, a wired interface for IP communication, but may also be a wired or wireless interface for any other communication method.

[0037] The decoding unit 220 decodes (decodes) encoded data received from the terminal 100. The decoding unit 220 decodes using a video encoding method compatible with the encoding method of the terminal 100, such as H.264 or H.265. The decoding unit 220 performs decoding according to the compression rate of each region in the frame, and generates decoded input video.

[0038] The time difference information acquisition unit 230 acquires time difference information ΔT (ΔT is a natural number) based on time stamp information acquired from a video compression codec or the like. The time difference information ΔT corresponds to a specific frame included in the input video and is information representing the time difference between the specific frame and the previous frame. In other words, the time difference information ΔT is 1 if no frames are skipped between the specific frame and the previous frame. On the other hand, the time difference information ΔT is 1+n if n frames are skipped between the specific frame and the previous frame. The time stamp information is information indicating the timing at which each frame included in the input video was captured by the camera 101. Note that the time stamp information may also be information indicating the timing at which each frame was encoded by the encoding unit 130 of the terminal 100.

[0039] The recognition unit 240 inputs the time-series frames included in the input video and the time difference information ΔT between the frames of the input video as input information into the trained recognition model M1, and recognizes objects in the input video. The recognition unit 240, for example, recognizes the work performed by the worker in the input video, i.e., the type of human behavior. Specifically, the trained recognition model M1 is a recurrent neural network (RNN) model that inputs the time-series frames included in the input video, and includes multiple RNN cells. The multiple RNN cells input parameters corresponding to the time difference information between the frames of the input video. More specifically, the multiple RNN cells input parameters obtained by decoding the time difference information between the frames of the input video by a decoder.

[0040] The storage unit 250 stores the trained recognition model M1. The training unit 260 generates the trained recognition model M1 by training using time difference information ΔT between frames of training videos and correct answer data.

[0041] Next, the recognition operation of the video processing system 1 according to the first embodiment will be described with reference to FIGS.

[0042] 7 is a flowchart showing the operation of the video processing system 1 according to the first embodiment. As shown in Fig. 7, first, the video acquisition unit 110 of the terminal 100 of the video processing system 1 acquires an input video of a site from the camera 101 (step S101). The input video includes frames in time series.

[0043] The frame filter unit 120 filters the time-series frames included in the input video (step S102), where frames that have not been filtered out are skipped.

[0044] The encoding unit 130 encodes the filtered input video (step S103). Next, the terminal communication unit 140 transmits the encoded data to the center server 200 via the base station 300 (step S104).

[0045] Next, the center communication unit 210 of the center server 200 receives the encoded data from the terminal 100 (step S105). Next, the decoding unit 220 obtains the input video by decoding the encoded data (step S106).

[0046] Next, the time difference information acquisition unit 230 acquires time difference information ΔT between frames corresponding to the frames of the input video (step S107). Specifically, the time difference information acquisition unit 230 acquires the time difference information ΔT based on time stamp information acquired from a video compression codec or the like. The time stamp information is, for example, information on the timing at which each frame included in the input video was captured by the camera 101.

[0047] Next, the recognition unit 240 inputs the time series frames included in the input video and the time difference information ΔT corresponding to the frames of the input video as input information to the trained recognition model M1 (step S108).

[0048] 8 is a diagram showing an example of input information input to the trained recognition model M1. As shown in FIG. 8, the input information includes time-series frames included in the input video and time difference information ΔT corresponding to the frames. The time difference information ΔT represents the time difference between a corresponding specific frame and the previous frame. For example, the time difference information ΔT is 1 if no frames are skipped between the corresponding specific frame and the previous frame. Furthermore, the time difference information ΔT is 1+n if n frames are skipped between the corresponding specific frame and the previous frame.

[0049] 7 , the recognition unit 240 then recognizes objects in the input video using the trained recognition model M1 (step S109). The recognition unit 240 recognizes, for example, the work performed by the worker in the input video, i.e., the type of person's behavior.

[0050] 9 is a diagram showing an example of the configuration and recognition operation of a trained recognition model M1. As shown in FIG. 9, the trained recognition model M1 is a recurrent neural network (RNN) model and includes multiple RNN time-series cells M11. If the RNN structure is classified into an input layer, an intermediate layer, and an output layer, the cells M11 correspond to the intermediate layers of the RNN. Furthermore, the trained recognition model M1 includes decoders M12 corresponding to each cell M11.

[0051] At a predetermined time, decoder M12 receives input of time difference information ΔT and outputs parameters obtained by decoding the received time difference information ΔT to cell M11. Next, cell M11 receives input of a frame, a state vector output by cell M11 at a previous time, and parameter information output by decoder M12, and outputs a state vector to cell M11 at a later time. Note that the initial state of the state vector input to cell M11 may be, for example, all elements of which are 0.

[0052] For example, at time t, decoder M12 accepts input of time difference information ΔT of 1, and outputs parameters obtained by decoding the time difference information ΔT of 1 to cell M11. Here, since no frame skip occurs between the frame input to cell M11 at time t-1 and the frame input to cell M11 at time t, the time difference information ΔT input to decoder M12 at time t is 1. Then, at time t, cell M11 accepts input of the frame, the state vector output by cell M11 at time t-1, and parameters, and outputs the state vector to cell M11 at time t+1.

[0053] Meanwhile, at time t+1, decoder M12 accepts input of time difference information ΔT of 2 and outputs parameters obtained by decoding time difference information ΔT of 2 to cell M11. Here, because one frame skip occurs between the frame input to cell M11 at time t and the frame input to cell M11 at time t+1, the time difference information ΔT input to decoder M12 at time t+1 is 2. Then, at time t+1, cell M11 accepts input of the frame, the state vector output by cell M11 at time t, and parameters, and outputs the state vector to cell M11 at time t+2.

[0054] Next, the learning operation of the recognition model M1 in the video processing system 1 according to the first embodiment will be described.

[0055] The learning unit 260 inputs a time series of frames included in a training video and time difference information ΔT corresponding to the frames into the recognition model M1. The training video includes, for example, a time series of frames in which frame skipping occurs according to a predetermined pattern. The configuration of the recognition model M1 has been described above (see FIG. 9 ). The learning unit 260 learns the recognition model M1 by comparing the output result of the recognition model M1 with the correct answer data, and generates a trained recognition model M1.

[0056] As described above, the trained recognition model M1 of the video processing system 1 decodes the time difference information ΔT and dynamically determines the parameters to be input to cell M11. In other words, the trained recognition model M1 can improve the object recognition accuracy by reflecting the frame time difference information in object recognition, taking into account cases where frames are skipped due to a change in the video frame rate, etc.

[0057] Second Embodiment Next, the configuration of a video processing system 2 according to a second embodiment will be described. Similar to the video processing system 1 according to the first embodiment, the video processing system 2 includes a plurality of terminals 100, a center server 200, a base station 300, and an MEC 400. The center server 200 of the video processing system 2 differs from the center server 200 of the video processing system 1 in the following configuration.

[0058] Fig. 10 is a diagram showing the configuration of the center server 200 of the video processing system 2. As shown in Fig. 10, the center server 200 of the video processing system 2 includes a center communication unit 210, a decoding unit 220, a time difference information acquisition unit 230, a recognition unit 270, a storage unit 280, and a learning unit 290.

[0059] The recognition unit 270 inputs the time series frames included in the input video and the time difference information ΔT between the frames of the input video as input information to the trained recognition model M2, and recognizes objects in the input video. The trained recognition model M2 inputs the time series frames included in the input video and includes multiple cells of a recurrent neural network (RNN) that inputs and outputs state vectors in a time series. The trained recognition model M2 inserts a state predictor that predicts a state vector based on the time difference information ΔT between frames between predetermined cells, such as between cells where a frame skip has occurred. The storage unit 280 stores the recognition model M2.

[0060] The learning unit 290 uses time difference information between frames in a time series that have undergone frame skipping in a predetermined pattern contained in the training video and frames in the training video, as well as correct answer data, to train a trained recognition model M2 with a state predictor inserted.

[0061] The learning unit 290 also uses time-series frames in the training video that do not have a frame skip and the correct answer data to train multiple cells of the trained recognition model M2. The learning unit 290 then inputs the time-series frames in the training video that do not have a frame skip into multiple cells of the trained recognition model M2. The learning unit 290 then trains a state predictor using state vectors output by the multiple cells at time t (t is a natural number) and state vectors output at time t + N (N is a natural number).

[0062] On the other hand, the recognition unit 270 inputs the time difference information between the frames of the input video and the movement between the frames of the input video into a trained recognition model that has been trained using the time difference information between the frames of the training video and the movement between the frames of the input video, and recognizes objects in the input video.

[0063] The learning unit 290 uses time difference information between frames in a time series that have undergone frame skipping in a predetermined pattern contained in the training video and frames in the training video, as well as the motion between frames in the input video and correct answer data to train a trained recognition model M2 with a state predictor inserted.

[0064] Next, the recognition operation of the video processing system 2 according to the second embodiment will be described with reference to FIG.

[0065] 11 is a flowchart showing an example of the operation of the video processing system 2 according to the second embodiment. As shown in FIG. 11, the video processing system 2 first executes the above-described processes of steps S101 to S107 (see FIG. 7). A description of the processes of steps S101 to S107 will be omitted.

[0066] Next, the recognition unit 270 of the center server 200 inputs the time-series frames included in the input video and the time difference information ΔT corresponding to the frames as input information to the trained recognition model M2 (step S201). An example of the input information has been described above (see FIG. 8). However, in this embodiment, the recognition unit 240 inputs the time difference information ΔT where ΔT≠1 as input information to the trained recognition model M2.

[0067] Next, the recognition unit 270 recognizes objects in the input video using the trained recognition model M2 (step S202). The recognition unit 240 recognizes, for example, the work performed by the worker in the input video, i.e., the type of person's behavior.

[0068] 12 is a diagram showing an example of the configuration and recognition operation of a trained recognition model M2 according to the second embodiment. As shown in FIG. 12, the trained recognition model M2 is a recurrent neural network (RNN) that includes multiple time-series RNN cells M21. At a given time, cell M21 receives input of a frame and a state vector output by cell M21 at a previous time, and outputs the state vector to cell M21 at a later time. Note that the initial state of the state vector may be, for example, all elements of which are 0.

[0069] Furthermore, in the trained recognition model M2, if a frame skip occurs between a frame input to cell M21 at a predetermined time and a frame input to cell M21 at a time preceding the predetermined time, a state predictor M22 is inserted between cell M21 at the predetermined time and cell M21 at the previous time. The occurrence of a frame skip can be determined from time difference information ΔT corresponding to the frame input to cell M21 at the predetermined time. The inserted state predictor M22 receives as input the state vector output by cell M21 at the previous time and time difference information ΔT corresponding to the frame input to cell M21 at the predetermined time. The state predictor M22 then predicts the state vector and outputs the predicted state vector to cell M21 at the predetermined time.

[0070] For example, a frame skip occurs between the frame input to cell M21 at time t+1 and the frame input to cell M21 at time t. In this case, the trained recognition model M2 inserts a state predictor M22 between cell M21 at time t+1 and cell M21 at time t. The state predictor M22 receives as input the state vector output by cell M21 at time t and time difference information ΔT of 2 corresponding to the frame input to cell M21 at time t+1. Here, the input time difference information ΔT is 2 because one frame skip occurs between the frame input to cell M21 at time t+1 and the frame input to cell M21 at time t. The state predictor M22 predicts the state vector and outputs the predicted state vector to cell M21 at time t+1.

[0071] Next, the learning operation of the recognition model M2 of the video processing system 2 according to the second embodiment will be described with reference to FIGS.

[0072] FIG. 13 is a diagram illustrating an example of a first learning operation of the recognition model M2. As illustrated in FIG. 13, the learning unit 290 inputs time-series frames included in a training video and time difference information ΔT (ΔT≠1) corresponding to the frames to the recognition model M2. Specifically, the learning unit 290 inputs the time-series frames included in the training video to multiple cells M21 of the recognition model M2. The input time-series frames have undergone frame skipping according to a predetermined pattern. Furthermore, when a frame skip occurs between a frame input to the cell M21 of the recognition model M2 at a predetermined time and a frame input to the cell M21 at a time preceding the predetermined time, the learning unit 290 inserts a state predictor M22 between the cell M21 at the predetermined time and the cell M21 at the previous time. The learning unit 290 inputs time difference information ΔT corresponding to the frame input to the cell M21 at the predetermined time to the state predictor M22. For example, one frame skip occurs between a frame input to a predetermined cell M21 at time t+1 and a frame input to the cell M21 at time t. In this case, the learning unit 290 inserts a state predictor M22 between the cell M21 at time t+1 and the cell M21 at time t, and inputs time difference information ΔT of 2 corresponding to the frame input to the predetermined cell M21 at time t+1.

[0073] The learning unit 290 then learns the recognition model M2 into which the state predictor M22 has been inserted by comparing the output result of the recognition model M2 into which the state predictor M22 has been inserted with the correct answer data. Note that the learning unit 290 may learn the recognition model M2 into which the state predictor M22 has been inserted separately from the learning of the recognition model M2.

[0074] 14 is a diagram showing an example of a second learning operation of the recognition model M2. As shown in FIG. 14, the learning unit 290 inputs time-series frames included in the learning video into multiple cells M21 of the recognition model M2. The input time-series frames have not undergone frame skipping. The learning unit 290 then learns the recognition model M2 by comparing the output results of the recognition model M2 with the correct answer data.

[0075] Next, the learning unit 290 inputs time-series frames included in the training video into a plurality of cells M21 of the trained recognition model M2. The input time-series frames are free of frame skipping.

[0076] Next, the learning unit 290 acquires a data set consisting of the state vector output from the cell M21 at time t and the state vector output from the cell M21 at time t+N (N is a natural number), and trains the state predictor M22 using the acquired data set as training data. In detail, the learning unit 290 trains the state predictor M22 by, for example, performing a regression analysis on the output result when the state vector at time t and N are input to the state predictor M22 so as to bring the output result closer to the state vector at time t+N.

[0077] FIG. 15 is a diagram showing another example of the recognition operation of the trained recognition model M2 according to the second embodiment.

[0078] 15, the trained recognition model M2 is a recurrent neural network (RNN) that includes multiple time-series RNN cells M21. At a given time, cell M21 receives input of a frame and a state vector output by cell M21 at a previous time, and outputs the state vector to cell M21 at a later time. Note that the initial state of the state vector may be, for example, all elements of which are 0.

[0079] Furthermore, if a frame skip occurs between a frame input to cell M21 at a predetermined time and a frame input to cell M21 at a time preceding the predetermined time, the trained recognition model M2 inserts a state predictor M23 between the predetermined cell M21 and the cell M21 at the previous time. The state predictor M23 receives as input a state vector output by cell M21 at the previous time, time difference information ΔT corresponding to the frame input to cell M21 at the predetermined time, and a motion vector. The motion vector is the difference between the frame at the predetermined time and the frame at the previous time, i.e., information obtained by vectorizing the motion. The state predictor M23 predicts the state vector and outputs the predicted state vector to cell M21 at the predetermined time.

[0080] For example, a frame skip occurs between the frame input to cell M21 at time t+1 and the frame input to cell M21 at time t. In this case, the trained recognition model M2 inserts a state predictor M23 between cell M21 at time t+1 and cell M21 at time t. The state predictor M23 receives input of a motion vector along with time difference information ΔT corresponding to the state vector output by cell M21 at time t and the frame input to cell M21 at time t+1. The input motion vector represents the difference, i.e., the movement, between the frame input to cell M21 at time t and the frame input to cell M21 at time t+1. The state predictor M23 predicts the state vector and outputs the predicted state vector to cell M21 at time t+1.

[0081] Next, the learning operation of the recognition model M2 in the video processing system 2 will be described. The learning unit 290 inputs time difference information ΔT (ΔT≠1) and motion vectors corresponding to time-series frames included in the learning video into the recognition model M2. The learning unit 290 learns the recognition model M2 into which the state predictor M22 has been inserted by comparing the output result of the recognition model M2 into which the state predictor M22 has been inserted with correct answer data. Note that the learning unit 290 may perform the learning of the state predictor M22 separately from the learning of the recognition model M2.

[0082] As described above, when a frame skip occurs, the trained recognition model M2 of the video processing system 2 according to the second embodiment inserts the state predictor M22 or the state predictor M23 between the cells M21 to predict a state vector. The trained recognition model M1 can improve the object recognition accuracy by reflecting frame time difference information in object recognition, taking into account cases where frames are skipped due to a change in the video frame rate, etc.

[0083] Each component in the above-described embodiments may be configured with hardware or software, or both, and may be configured with a single piece of hardware or software, or may be configured with multiple pieces of hardware or software. Each device and each function (processing) may be realized by a computer 1000 having a processor 1001 such as a CPU (Central Processing Unit) and a memory 1002 serving as a storage device, as shown in FIG. 19 . For example, a program for performing the method (video processing method) in the embodiment may be stored in the memory 1002, and each function may be realized by the processor 1001 executing the program stored in the memory 1002.

[0084] These programs include instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The programs may be stored on a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The programs may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals.

[0085] Some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes. (Supplementary Note 1) A video processing system comprising: video acquisition means for acquiring input video; time difference information acquisition means for acquiring first time difference information between frames of the input video; and recognition means for inputting the input video and the first time difference information between frames of the input video into a trained recognition model trained using a training video and second time difference information between frames of the training video, and recognizing an object in the input video. (Supplementary Note 2) The trained recognition model is a model including multiple cells of a recurrent neural network (RNN) that inputs frames in time series included in the input video, and the multiple cells input parameters corresponding to the first time difference information between frames of the input video. (Supplementary Note 3) The video processing system according to Supplementary Note 2, wherein the multiple cells input parameters obtained by decoded the first time difference information between frames of the input video. (Supplementary Note 4) The video processing system according to Supplementary Note 1, wherein the trained recognition model is a model that includes a plurality of cells of a recurrent neural network that inputs time-series frames included in the input video and inputs and outputs state vectors in time series, and has a state predictor inserted between predetermined cells that predicts a state vector based on first time difference information between frames of the input video. (Supplementary Note 5) The video processing system according to Supplementary Note 4, wherein the trained recognition model with the state predictor inserted is trained using time-series frames that have undergone frame skipping in a predetermined pattern included in the training video, second time difference information between frames of the training video, and correct answer data.(Supplementary Note 6) The video processing system according to Supplementary Note 4, wherein the plurality of cells of the trained recognition model are trained using time-series frames included in the training video that do not have a frame skip and correct answer data, and the state predictor inserted into the trained recognition model is trained using a state vector output by the plurality of cells at time t (t is a natural number) and a state vector output at time t+N (N is a natural number) when time-series frames included in the training video that do not have a frame skip are input to the trained plurality of cells. (Supplementary Note 7) The video processing system according to Supplementary Note 1, wherein the recognition means inputs first time difference information between the input video and frames of the input video and the movement between frames of the input video to a trained recognition model trained using second time difference information between the training video and frames of the training video and the movement between frames of the training video, and recognizes an object in the input video. (Supplementary Note 8) A video processing device comprising: video acquisition means for acquiring an input video; time difference information acquisition means for acquiring first time difference information between frames of the input video; and recognition means for inputting the input video and the first time difference information between frames of the input video into a trained recognition model trained using a training video and second time difference information between frames of the training video, and recognizing an object in the input video. (Supplementary Note 9) The trained recognition model is a model including a plurality of cells of a recurrent neural network (RNN) that inputs frames in time series included in the input video, and the plurality of cells input parameters corresponding to the first time difference information between frames of the input video. (Supplementary Note 10) The video processing device according to Supplementary Note 9, wherein the plurality of cells input parameters obtained by decoded the first time difference information between frames of the input video.(Supplementary Note 11) The video processing device according to Supplementary Note 8, wherein the trained recognition model is a model that inputs time-series frames included in the input video and includes a plurality of cells of a recurrent neural network that inputs and outputs state vectors in time series, and has a state predictor inserted between predetermined cells that predicts a state vector based on first time difference information between frames of the input video. (Supplementary Note 12) The video processing device according to Supplementary Note 11, wherein the trained recognition model with the state predictor inserted is trained using time-series frames that have undergone frame skipping in a predetermined pattern included in the training video, second time difference information between frames of the training video, and correct answer data. (Supplementary Note 13) The video processing device according to Supplementary Note 11, wherein a plurality of cells of the trained recognition model are trained using time-series frames included in the training video that do not cause frame skipping and correct answer data, and the state predictor inserted into the trained recognition model is trained using state vectors output by the plurality of cells at time t (t is a natural number) and state vectors output at time t+N (N is a natural number) when time-series frames included in the training video that do not cause frame skipping are input to the trained plurality of cells. (Supplementary Note 14) The video processing device according to Supplementary Note 8, wherein the recognition means inputs first time difference information between the input video and frames of the input video and movement between frames of the input video to a trained recognition model trained using second time difference information between the training video and frames of the training video and movement between frames of the training video, and recognizes an object in the input video. (Supplementary Note 15) A video processing method in which a computer acquires an input video, acquires first time difference information between frames of the input video, inputs the input video and the first time difference information between frames of the input video to a trained recognition model trained using a training video and second time difference information between frames of the training video, and recognizes an object in the input video.(Supplementary Note 16) The video processing method of Supplementary Note 15, wherein the trained recognition model is a model including a plurality of cells of a recurrent neural network (RNN) that inputs time-series frames included in the input video, and the plurality of cells input parameters corresponding to first time difference information between frames of the input video. (Supplementary Note 17) The video processing method of Supplementary Note 16, wherein the plurality of cells input parameters obtained by decoded the first time difference information between frames of the input video. (Supplementary Note 18) The video processing method of Supplementary Note 15, wherein the trained recognition model is a model that inputs time-series frames included in the input video and includes a plurality of cells of a recurrent neural network that inputs and outputs state vectors in a time series, and has a state predictor inserted between predetermined cells that predicts a state vector based on the first time difference information between frames of the input video. (Supplementary Note 19) The video processing method of Supplementary Note 18, wherein the trained recognition model into which the state predictor has been inserted is trained using second time difference information between time-series frames in which frame skipping has occurred in a predetermined pattern included in the training video and frames of the training video, and correct answer data. (Supplementary Note 20) The video processing method of Supplementary Note 18, wherein a plurality of cells of the trained recognition model are trained using time-series frames in which frame skipping has occurred included in the training video and correct answer data, and the state predictor inserted into the trained recognition model is trained using state vectors output by the plurality of cells at time t (t is a natural number) and state vectors output at time t+N (N is a natural number) when time-series frames in the training video in which frame skipping has occurred are input to the plurality of trained cells. (Supplementary Note 21) The video processing method described in Supplementary Note 15, wherein a computer inputs first time difference information between the input video and frames of the input video and motion between frames of the input video into a trained recognition model trained using second time difference information between the training video and frames of the training video and motion between frames of the training video, and recognizes objects in the input video.

[0086] 1, 2, 10 Video processing system 11 Video acquisition unit 12 Time difference information acquisition unit 13 Recognition unit 20 Video processing device 100 Terminal 101 Camera 102 Compression efficiency optimization function 110 Video acquisition unit 120 Frame filter unit 130 Encoding unit 140 Terminal communication unit 200 Center server 201 Video recognition function 202 Alert generation function 203 GUI drawing function 204 Screen display function 210 Center communication unit 220 Decoding unit 230 Time difference information acquisition unit 240, 270 Recognition unit 250, 280 Storage unit 260, 290 Learning unit 300 Base station 401 Compression bit rate control function 1000 Computer 1001 Processor 1002 Memory M1, M2 Recognition model M11, M21 Cell M12 Decoder M22, M23 State predictor

Claims

1. An image acquisition means for acquiring an input image; a time difference information acquiring means for acquiring first time difference information between frames of the input video; a recognition means for inputting the input video and the first time difference information between frames of the input video into a trained recognition model trained using a training video and second time difference information between frames of the training video, and recognizing an object in the input video. Video processing system.

2. The trained recognition model is a model including a plurality of cells of a recurrent neural network (RNN) to which a time series of frames included in the input video is input, The plurality of cells input a parameter corresponding to first time difference information between frames of the input video. The video processing system according to claim 1 .

3. The trained recognition model is a plurality of cells of a recurrent neural network that receives frames in a time series included in the input video and inputs and outputs state vectors in a time series; A model in which a state predictor is inserted between predetermined cells to predict a state vector based on first time difference information between frames of the input video. The video processing system according to claim 1 .

4. The trained recognition model into which the state predictor is inserted is Learning is performed using second time difference information between a time-series frame in which a frame skip occurs in a predetermined pattern included in the learning video and a frame of the learning video, and correct answer data. The video processing system according to claim 3 .

5. The plurality of cells of the trained recognition model include Learning is performed using time-series frames that are included in the learning video and correct answer data that do not cause frame skips, The state predictor to be inserted into the trained recognition model is When a time-series frame without frame skip included in the learning video is input to the learned cells, learning is performed using a state vector output by the learned cells at time t (t is a natural number) and a state vector output by the learned cells at time t+N (N is a natural number). The video processing system according to claim 3 .

6. The recognition means includes: The first time difference information between the input video and the frames of the input video and the motion between the frames of the input video are input to a trained recognition model trained using the training video and second time difference information between the frames of the training video and the motion between the frames of the training video, thereby recognizing an object in the input video. The video processing system according to claim 1 .

7. An image acquisition means for acquiring an input image; a time difference information acquiring means for acquiring first time difference information between frames of the input video; a recognition means for inputting the input video and the first time difference information between frames of the input video into a trained recognition model trained using a training video and second time difference information between frames of the training video, and recognizing an object in the input video. Image processing device.

8. The trained recognition model is a model including a plurality of cells of a recurrent neural network (RNN) to which a time series of frames included in the input video is input, The plurality of cells input a parameter corresponding to first time difference information between frames of the input video. The image processing device according to claim 7.

9. The computer Get the input video, Obtaining first time difference information between frames of the input video; The input video and the first time difference information between frames of the input video are input to a trained recognition model trained using a training video and second time difference information between frames of the training video, and an object in the input video is recognized. Image processing method.

10. The trained recognition model is a model including a plurality of cells of a recurrent neural network (RNN) to which a time series of frames included in the input video is input, The plurality of cells input a parameter corresponding to first time difference information between frames of the input video. The video processing method according to claim 9.