Moving image processing device, data processing device, image processing system, vehicle, moving image processing method, and program
By generating object feature data and sampling frame images, the video processing device reduces communication load and maintains object movement reproduction, addressing the inefficiency of transferring large moving image data.
Patent Information
- Application Number
- JP2024031440
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-01
- Publication Date
- 2025-09-11
AI Technical Summary
The large amount of data in moving images imposes a significant communication load when transferring video images, making it inefficient and burdensome.
A video processing device that performs skeletal detection on frame images, generates object feature data, and creates an output video by sampling some frame images, reducing the data volume by generating image set data that includes both image data and object feature data.
The reduced data volume allows for effective communication load reduction while maintaining the ability to reproduce object movement, even if some frame images are missing.
Smart Images

Figure 2025133467000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a moving image processing device, a data processing device, an image processing system, a vehicle, a moving image processing method, and a program. [Background technology]
[0002] Technologies have been developed that detect the position and posture of an object (mainly a person) from an image captured by a camera, and understand or predict the behavior of the object based on the detection results (see, for example, Patent Document 1 below). There have also been attempts to transfer video images captured by a camera to a data processing device (such as a server device on the cloud) via wireless communication, and have the data processing device perform status analysis of the object in the video. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-98484 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the amount of image data for a moving image is often enormous, and transmitting the image data itself imposes a heavy communication load.
[0005] An object of the present invention is to provide a technique that contributes to reducing the communication load involved in transferring moving images. [Means for solving the problem]
[0006] The video processing device of the present invention has a controller that performs skeletal detection on objects in each frame image based on a plurality of frame images that make up an input video, generates object feature data for each frame image that includes the results of the skeletal detection, generates an output video by sampling some of the plurality of frame images, and generates image set data that includes image data of the output video and the object feature data for each of the plurality of frame images. [Effects of the Invention]
[0007] The amount of data for each object feature data is much smaller than the amount of data for each frame image. Therefore, the amount of data for the image set data can be made smaller than the amount of data for the input video by roughly the amount of data for the frame images that were not included in the output video from the multiple frame images that make up the input video. On the other hand, even if some frame images are missing from the input video in the output video, the movement of the object can be reproduced based on the image set data including the results of skeletal detection. In other words, by using the image processing device according to the present invention, the movement of the object can be reproduced on the receiving device of the image set data, and in this case, the communication load can be reduced compared to when all image data for the input video is transmitted and received. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 2 is a diagram illustrating the relationship between a user and other components according to an embodiment of the present invention. [Figure 2] 1 is a schematic block diagram of an in-vehicle system according to an embodiment of the present invention. [Figure 3] FIG. 2 is a diagram showing a shooting area of a camera according to an embodiment of the present invention. [Figure 4] 1 is a diagram illustrating an internal configuration of an in-vehicle device according to an embodiment of the present invention. [Figure 5] 1 is a diagram illustrating an internal configuration of a data processing device according to an embodiment of the present invention. [Figure 6]1A and 1B are diagrams illustrating the relationship between an input video sequence and thinned-out video sequences, and the structure of the input video sequence, according to an embodiment of the present invention. [Figure 7] FIG. 2 is a functional block diagram of an image processing unit provided in the in-vehicle device according to the embodiment of the present invention. [Figure 8] FIG. 10 is a diagram showing how a bounding box is set for one input image according to the embodiment of the present invention. [Figure 9] FIG. 2 is a functional block diagram relating to an embodiment of the present invention and relating to behavior recognition processing. [Figure 10] FIG. 2 is a diagram illustrating the relationship between input video sequences, object feature data, thinned video sequences, and image set data according to an embodiment of the present invention. [Figure 11] 2 is a functional block diagram of an image processing unit provided in the data processing device according to the embodiment of the present invention. FIG. [Figure 12] FIG. 1 relates to a first example belonging to an embodiment of the present invention, and shows the relationship between an input video sequence, object feature data, thinned video sequences, and restored video sequences. [Figure 13] 4 is a flowchart showing the operation of a controller in an in-vehicle device according to a first example of an embodiment of the present invention. [Figure 14] 4 is a flowchart illustrating the operation of a controller in a data processing device according to a first example of an embodiment of the present invention. [Figure 15] FIG. 10 is a diagram showing the internal configuration of a machine learning device according to a second example of an embodiment of the present invention. [Figure 16] FIG. 10 is a functional block diagram of a controller provided in a machine learning device according to a second example of an embodiment of the present invention. [Figure 17] FIG. 10 relates to a third example belonging to an embodiment of the present invention and shows the relationship between input video sequences, object feature data, and thinned video sequences. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, examples of embodiments of the present invention will be described in detail with reference to the drawings. In each of the drawings, the same parts are designated by the same reference numerals, and duplicate descriptions of the same parts will be omitted as a general rule. In this specification, for the sake of simplicity, symbols or signs referring to information, signals, physical quantities, functional units, circuits, elements, or components may be used, and the names of the information, signals, physical quantities, functional units, circuits, elements, or components corresponding to the symbols or signs may be omitted or abbreviated.
[0010] FIG. 1 shows the relationship between user U1 and other components assumed in an embodiment of the present invention. User U1 is an occupant of vehicle V1. User U1 is the driver of vehicle V1. However, user U1 may also be an occupant other than the driver (i.e., a passenger in vehicle V1). Vehicle V1 is any type of vehicle. Here, vehicle V1 is assumed to be an automobile or the like that runs on a road. An in-vehicle system 1 is mounted on vehicle V1, and each component of the in-vehicle system 1 is installed in an appropriate location in vehicle V1.
[0011] A seat ST1 is installed in the cabin of the vehicle V1. A user U1 sits in the seat ST1. Since it is assumed that the user U1 is the driver, the seat ST1 is the driver's seat. Hereinafter, when simply referring to the cabin, unless otherwise specified, this refers to the cabin of the vehicle V1. Also, below, unless otherwise specified, the inside of the vehicle refers to the internal area of the vehicle V1, and the outside of the vehicle refers to the external area of the vehicle V1. Figure 1 shows an information terminal TM carried by the user U1 located inside the cabin. The information terminal TM is a portable information terminal, such as a smartphone or tablet.
[0012] The direction from the driver's seat of vehicle V1 toward the steering wheel is defined as "forward," and the direction from the steering wheel of vehicle V1 toward the driver's seat is defined as "rearward." The direction perpendicular to the fore-and-aft direction and parallel to the road surface on which vehicle V1 is traveling is defined as the left-right direction. User U1 sits in seat ST1 facing forward. The fore-and-aft direction and left-and-right direction correspond to the fore-and-aft direction and left-and-right direction as seen from user U1's perspective. Unless otherwise specified below, vehicle V1 is assumed to be located on a horizontal road surface, and the traveling direction of vehicle V1 is assumed to be forward.
[0013] 2 shows a schematic block diagram of the in-vehicle system 1. The in-vehicle system 1 includes an in-vehicle device 10, a cruise control device 20, an actuator unit 30, a vehicle sensor unit 40, a camera unit 50, and an HMI 60. The components of the in-vehicle system 1 can transmit and receive any signals and information to and from each other through an in-vehicle network formed in the vehicle V1. The in-vehicle network includes, for example, a CAN (Controller Area Network) and an AVCLAN (Audio Visual Communication Local Area Network).
[0014] Some of the components of the in-vehicle system 1 are wirelessly connected to a communication network NET (see FIG. 1). At least the in-vehicle device 10 is wirelessly connected to the communication network NET. The communication network NET includes the Internet, an intranet, etc. A data processing device 200 and a database DB are connected to the communication network NET by wire or wirelessly. The data processing device 200 is an example of an external device provided outside the in-vehicle device 10.
[0015] The in-vehicle device 10 is capable of two-way communication with the data processing device 200 via the communication network NET. The in-vehicle device 10 can also access the database DB via the data processing device 200. The in-vehicle device 10 may also be able to directly access the database DB without going through the data processing device 200. The data processing device 200 can also access the database DB via the communication network NET. However, the database DB may be built into the data processing device 200. The database DB is a large-capacity recording medium formed of a magnetic disk, semiconductor memory, or the like. Access to the database DB includes a write operation for recording any data (information) in the database DB and a read operation for reading any data (information) recorded in the database DB. The data processing device 200 is composed of one or more computers connected to the communication network NET. The data processing device 200 may also be configured using cloud computing. In this specification, the terms "recording information or data" and "storing information or data" are synonymous. In this specification, the terms "information" and "data" are interchangeable.
[0016] Each component shown in Figure 2 will be explained. The in-vehicle device 10 cooperates with the camera unit 50 to record the situation outside or inside the vehicle (details will be described later). The driving control device 20 controls the driving of the vehicle V1 using the actuator unit 30. The actuator unit 30 has various driving parts such as a motor that realizes the driving of the vehicle V1. Specifically, the actuator unit 30 includes an engine and a motor that generate the driving force of the vehicle V1, a steering actuator that drives the steering of the vehicle V1, and a brake actuator that drives the brakes of the vehicle V1.
[0017] The vehicle sensor unit 40 has sensors that detect the details of the driving operation of the vehicle V1 by the driver of the vehicle V1 and sensors that detect various states of the vehicle V1. The vehicle sensor unit 40 outputs vehicle sensor information containing the detection results. The driving control device 20 realizes driving control of the vehicle V1 by driving and controlling the actuator unit 30 in accordance with the vehicle sensor information.
[0018] The camera unit 50 consists of one or more unit cameras that capture images of the outside or inside of the vehicle V1. Each unit camera captures images at a predetermined frame rate. Some of the unit cameras provided in the camera unit 50 are exterior cameras. The exterior cameras have a capture area set outside the vehicle V1, and generate exterior camera images by capturing images of the situation within the capture area. The exterior camera images are images captured of the capture area by the exterior camera. Some of the unit cameras provided in the camera unit 50 are interior cameras. The interior cameras have a capture area set inside the vehicle V1 (i.e., the interior of the vehicle V1), and generate interior camera images by capturing images of the situation in the capture area. The interior camera images are images captured of the capture area by the interior camera. Data representing the content of any image is called image data.
[0019] The following description focuses primarily on camera 51, one of the unit cameras provided in camera unit 50. While camera 51 may be an in-vehicle camera, it is assumed below that it is an exterior camera. For the sake of clarity, camera 51 is assumed to capture images of the front side of vehicle V1. Therefore, as shown in FIG. 3, the capture area of camera 51 includes the front area of vehicle V1. In FIG. 3, a shaded area SR1 represents a portion of the capture area of camera 51. However, the capture area of camera 51 may also be the rear area, right side area, or left side area of vehicle V1. The capture area of camera 51 may also include all or part of the front area, rear area, right side area, and left side area of vehicle V1. The front area, rear area, right side area, and left side area of vehicle V1 are areas located in the exterior area of vehicle V1, in front of vehicle V1, rear of vehicle V1, right side area, and left side area of vehicle V1, respectively. Image data of the image captured by camera 51 is sent to in-vehicle device 10.
[0020] The HMI 60 is a human machine interface and is provided with a display device 61, a speaker 62, and an operation input unit 63.
[0021] The display device 61 has a display screen such as a liquid crystal display panel, and displays any video (image) under the control of the in-vehicle device 10, the driving control device 20, or a display control device not shown. The display device 61 is installed in an appropriate location in the cabin of the vehicle V1 so that each occupant of the vehicle V1 can see the display content of the display device 61. Multiple display devices 61 may be installed in the cabin of the vehicle V1. The display device 61 may be a component of a car navigation system installed in the vehicle V1. The car navigation system may be included in the in-vehicle system 1. The display device 61 may be a display device provided in the information terminal TM.
[0022] The speaker 62 outputs any sound (message, warning sound, music, etc.) under the control of the in-vehicle device 10, the driving control device 20, or an audio device (not shown). The speaker 62 is installed at an appropriate location in the cabin of the vehicle V1 so that each occupant of the vehicle V1 can hear the sound output from the speaker 62. Multiple speakers 62 may be installed in the cabin of the vehicle V1. The speaker 62 may be a speaker provided in the information terminal TM.
[0023] The operation input unit 63 receives arbitrary operations from each occupant of the vehicle V1. The operation input unit 63 can be configured with operation buttons, a touch panel, or the like. A microphone may be provided in the HMI 60, and voice operations using the microphone may be input to the operation input unit 63. The operation input unit 63 may be an operation input unit provided in the information terminal TM. In addition, a vibration device that applies vibrations to the occupants (particularly the driver) of the vehicle V1 may be provided in the HMI 60.
[0024] 4 shows the internal configuration of the in-vehicle device 10. The in-vehicle device 10 includes a controller 11, a memory 12, a communication unit 13, and a recording medium 14.
[0025] The controller 11 includes a processing unit including a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit) as hardware resources. The controller 11 may execute a program recorded in the memory 12 or any other recording medium to realize any function, operation, and process that should be realized by the controller 11. The controller 11 is provided with an image processing unit 11a as part of the processing unit within the controller 11.
[0026] The memory 12 is configured to include a non-volatile memory such as a ROM (Read Only Memory) or a flash memory, and a volatile memory such as a RAM (Random Access Memory). The memory 12 stores various data referenced by the controller 11 as well as various programs to be executed by the controller 11.
[0027] The communication unit 13 is a communication circuit (communication module) that transmits and receives any signal between the in-vehicle device 10 and a counterpart device different from the in-vehicle device 10. The counterpart device for the communication unit 13 includes components other than the in-vehicle device 10 among the components of the in-vehicle system 1 shown in FIG. 2. The communication unit 13 can communicate with the counterpart device via an in-vehicle network formed in the vehicle V1. The counterpart device for the communication unit 13 includes an external device connected to the communication network NET, and therefore includes the data processing device 200. Note that the controller 11 can transmit and receive any information to and from the counterpart device using the communication unit 13, but the description of the communication unit 13 may be omitted below.
[0028] The recording medium 14 is a non-volatile recording medium made of a magnetic disk, a flash memory, or the like, and stores (records) any information in a non-volatile manner. The controller 11 is capable of recording any information on the recording medium 14 and is also capable of reading any information recorded on the recording medium 14. The recording medium 14 may be detachable from the in-vehicle device 10. The recording medium 14 may also be external to the in-vehicle device 10 and installed within the in-vehicle system 1.
[0029] FIG. 5 shows the internal configuration of the data processing device 200. The data processing device 200 includes a controller 210, a memory 220, and a communication unit 230. The controller 210 includes a processing unit including a CPU, a GPU, and the like as hardware resources. The controller 210 may execute a program recorded in the memory 220 or any other recording medium to realize any function, operation, or process to be realized by the controller 210. The controller 210 includes an image processing unit 210a as part of the processing unit within the controller 210. The memory 220 includes a non-volatile memory such as a ROM or flash memory, and a volatile memory such as a RAM. The memory 220 stores various data referenced by the controller 210 as well as various programs to be executed by the controller 210. The communication unit 230 transmits and receives any signal between the data processing device 200 and a different counterpart device. The counterpart device for the communication unit 230 includes the in-vehicle device 10. The controller 210 can use the communication unit 230 to send and receive any information to and from a partner device, but in the following, the description of the communication unit 230 may be omitted.
[0030] The in-vehicle device 10 and the camera unit 50 form a drive recorder. In the drive recorder, the controller 11 executes a recording process in which image data of images captured by each unit camera in the camera unit 50 and additional data including vehicle sensor information are associated with each other and recorded on the recording medium 14. After the controller 11 is started up, the controller 11 may execute the recording process all the time, or may execute the recording process only when it detects that a predetermined event (such as a sudden braking event) has occurred.
[0031] Referring to FIG. 6(a), the image processing unit 11a has the function of generating a thinned video MVb from the input video MVa. FIG. 6(b) shows the structures of the input video MVa and the thinned video MVb. A single still image forming a given video is called a frame image. A captured image obtained by one capture by the camera 51 functions as a frame image in the video. The camera 51 captures images sequentially at a predetermined frame rate, generating multiple frame images arranged in chronological order. Of the multiple frame images generated by the camera 51, each frame image generated during the capture period of the input video MVa is specifically referred to as a frame image Fa. Therefore, the input video MVa is composed of multiple frame images Fa. In the input video MVa, the multiple frame images Fa are arranged in chronological order at intervals equal to the reciprocal of the frame rate RTa. The frame rate RTa represents the number of frame images Fa per second in the input video MVa. The frame rate RTa corresponds to the frame rate of the camera 51. However, the frame rate RTa may be at least twice as high as the shooting frame rate of the camera 51 and may be an integer multiple. Each frame image Fa contains an image of a detection target, and the detection target is detected by object detection for each frame image Fa.
[0032] The image processing unit 11a generates a thinned video MVb from the input video MVa through sampling. In the sampling process, the image processing unit 11a generates the thinned video MVb by sampling some of the frame images Fa that make up the input video MVa. In other words, in the sampling process, the image processing unit 11a generates the thinned video MVb by thinning out some of the frame images Fa that make up the input video MVa. For example, in the sampling process, the image processing unit 11a samples the frame images Fa at a fixed sampling interval and generates a collection of the sampled frame images Fa as the thinned video MVb. The example of FIG. 6(b) shows how the frame images Fa that make up the input video MVa are sampled at intervals of (2 / RTa). That is, the sampling interval in the example of FIG. 6(b) is (2 / RTa). However, the sampling interval may be any interval, such as (4 / RTa), (6 / RTa), or (8 / RTa).
[0033] In the following, for clarity and specificity of explanation, it is assumed that the input video sequence MVa is composed of n frame images Fa. As shown in FIG. 6(c), the n frame images Fa that make up the input video sequence MVa are referred to as frame images Fa[1] to Fa[n]. Frame image Fa[i] is an image captured by camera 51 at time t[i]. Time t[i+1] is a time later than time t[i]. Therefore, in the input video sequence MVa, frame images Fa[1], Fa[2], Fa[3], ..., Fa[n] are arranged in this order along the direction of time progression. n represents an integer greater than or equal to 2, and i represents any integer.
[0034] Referring to FIG. 7, the characteristic operations performed by the image processing unit 11a will be described. FIG. 7 is a functional block diagram of the image processing unit 11a. The image processing unit 11a includes functional blocks F11 to F13. The controller 11 including the image processing unit 11a is a program execution device (computer) capable of executing any program. All or part of the functions of the functional blocks F11 to F13 may be realized by the controller 11 executing a program recorded in the memory 12 or any other recording medium. Each frame image (more specifically, image data of each frame image) generated by the camera 51 is sequentially input to the image processing unit 11a. The frame images generated by the camera 51 and input to the image processing unit 11a are referred to as input images IN. An input video MVa is composed of a plurality of input images IN arranged in time series, and at this time, the plurality of input images IN that compose the input video MVa are frame images Fa[1] to Fa[n]. It should be noted that, with respect to any image, the input, output, recording, saving, generating, and obtaining of the image are synonymous with the input, output, recording, saving, generating, and obtaining of image data of the image. Similar expressions are also interpreted in the same manner. Furthermore, a process or operation based on any image is, in detail, a process or operation based on the image data of the image.
[0035] ---Detection part F11--- The functional block F11 is a detection unit. The detection unit F11 performs object detection processing and skeleton detection processing on each input image IN. Hereinafter, the object detection processing may be referred to as object detection, and the skeleton detection processing may be referred to as skeleton detection.
[0036] In object detection, the detection unit F11 detects whether a detection target exists within a detection target area in the input image IN based on the image data of the input image IN. The presence of a detection target in a target image or a target area within a target image specifically refers to the presence of an image of the detection target (in other words, image data of the detection target) within the target image or target area. The detection target area may be the entire image area of the input image IN, or a partial area of the entire image area of the input image IN. When a detection target is detected within a detection target area during object detection, the position and shape of the detection target in the input image IN are detected, as well as the type of the detection target. The detection target in the detection unit F11 includes at least a person (human). The detection target may be only a person. The detection target may also include an animal. In this embodiment, an animal refers to a vertebrate other than a person (human), such as a dog, cat, cow, or pig. Hereinafter, unless otherwise specified, only a person is considered as the detection target.
[0037] The detection unit F11 generates and outputs BBOX information and class information as a result of object detection for the input image IN. The BBOX information includes position information of the detected object in the input image IN (information specifying the position of the detected object). In detail, the BBOX information indicates the position and shape of the detected object in the input image IN. In object detection, a rectangular area in the input image IN where the image of the detected object exists is specified as a bounding box (hereinafter referred to as BBOX). The BBOX information indicates the position and shape of the BBOX. The class information indicates the type of the detected object in the input image IN.
[0038] FIG. 8 shows an input image 600 as an example of the input image IN. When the input image 600 is obtained by capturing an image with the camera 51, an object 610 is located within the capturing area of the camera 51, and as a result, the input image 600 includes an image of the object 610. The object 610 is a person. Therefore, by object detection on the input image 600, the position and shape of the object 610 in the input image 600 are detected, and the type of the object 610 is detected as a person. In FIG. 8, an area 620 within the dashed rectangular frame is a BBOX set for the object 610. The BBOX information for the object 610 specifies the coordinates of one of the four corners of the BBOX 620 (for example, the coordinates of the upper left corner), as well as the width and height of the BBOX 620. The width of the BBOX 620 is the horizontal length of the BBOX 620, and the height of the BBOX 620 is the vertical length of the BBOX 620. In object detection for the input image 600, class information indicating that the type of the object 610 is a person is generated in association with the object 610 and the BBOX 620. Hereinafter, the object 610 may be referred to as a person 610.
[0039] The object detection method itself is publicly known, and object detectors that perform object detection have been put to practical use. The object detector may be an object detection AI configured with a neural network that has undergone machine learning. In this specification, AI is an abbreviation for artificial intelligence. The detection unit F11 may include a publicly known object detector (not shown), and the detection unit F11 may realize object detection using the publicly known object detector.
[0040] Skeleton detection is only meaningful when an image of the detection target object exists within the detection target area in the input image IN and the detection target object is detected by object detection. In skeleton detection, the detection unit F11 detects the positions of specific parts of the detection target object in the input image IN. The detection unit F11 generates and outputs the results of skeleton detection as skeleton information. There are multiple specific parts. The multiple specific parts whose positions should be detected include, for example, the head, nose, neck, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, center of the spine, right hip, left hip, right knee, left knee, right ankle, and left ankle. However, the specific parts are not limited to the examples given here, and the total number of specific parts is arbitrary. The positions of the specific parts are called key points. Therefore, in skeleton detection, each key point of the detection target object in the input image IN is detected. However, depending on the posture of the detection target object in the input image IN, it may be impossible to detect some key points. The skeleton information indicates each key point detected for the detection object (that is, the detected position of each specific part of the detection object in the input image IN).
[0041] In skeleton detection for the input image 600 in Fig. 8, each key point of a person 610, which is the detection target, is detected. Therefore, skeleton information for the input image 600 indicates each key point detected for the person 610 (i.e., the detected position of each specific part of the person 610 in the input image 600). The connection relationship between key points in the detection target is called a skeleton. Skeleton information may also be included in the skeleton information.
[0042] The method of skeletal detection itself is publicly known, and skeletal detectors that perform skeletal detection have been put to practical use. The skeletal detector may be a skeletal detection AI configured with a neural network that has undergone machine learning. The detection unit F11 may include a publicly known skeletal detector (not shown), and the detection unit F11 may realize skeletal detection using the publicly known skeletal detector.
[0043] For the sake of convenience, object detection and skeleton detection have been described separately here, but the detection unit F11 can adopt either the top-down or bottom-up method as desired. A detection unit F11 that adopts the top-down method performs object detection on each input image IN, and then performs skeleton detection on the detection target object detected by object detection. In contrast, a detection unit F11 that adopts the top-down method performs object detection and skeleton detection simultaneously in parallel without dividing them.
[0044] If there are multiple detection targets within the detection target area in the input image IN, BBOX information, class information, and skeleton information are generated for each detection target. Also, if the detection target in the detection unit F11 is only a person, the generation and output of class information by the detection unit F11 can be omitted.
[0045] ---Action recognition section F12--- The functional block F12 is a behavior recognition unit. The behavior recognition unit F12 performs behavior recognition processing on the input image IN for each input image IN. The behavior recognition processing is performed on the detection target object detected by object detection. A person as the detection target object to which the behavior recognition processing is applied is called a target person. The output information of the detection unit F11 includes BBOX information, class information, and skeleton information, and is input to the behavior recognition unit F12. The behavior recognition unit F12 performs behavior recognition processing based on the output information of the detection unit F11 (BBOX information, class information, and skeleton information). In the behavior recognition processing, the behavior recognition unit F12 recognizes the behavior of the target person by detecting the posture, etc. of the target person, and generates and outputs behavior recognition information indicating the recognition result.
[0046] The method of behavior recognition using skeleton detection itself is publicly known, and the behavior recognition unit F12 may be configured using a publicly known method related to behavior recognition. The behavior recognition unit F12 may also be configured using an existing behavior recognition AI. When the detection target is a person, what the person is holding in their hand may also be detected in the behavior recognition process and included in the behavior recognition information.
[0047] In the behavior recognition process, the behavior recognition unit F12 may recognize which of multiple candidate behaviors the target person's behavior is. The multiple candidate behaviors include, for example, the following first to twelfth candidate behaviors. The first, second, third, and fourth candidate behaviors are walking while facing right, walking while facing left, walking toward vehicle V1, and walking away from vehicle V1, respectively. The fifth, sixth, seventh, and eighth candidate behaviors are running while facing right, running while facing left, running toward vehicle V1, and running away from vehicle V1, respectively. The ninth, tenth, eleventh, and twelfth candidate behaviors are stopping while facing vehicle V1, stopping while facing the opposite side of vehicle V1, stopping while facing right, and stopping while facing left, respectively. Based on the behavior recognition result for the input image 600 in FIG. 8, the behavior recognition information indicates that the behavior of person 610 as the target person is "walking while facing right." More specifically, for example, in the behavior recognition process, the behavior recognition unit F12 may identify one or more candidate behaviors from among a plurality of candidate behaviors, and derive, for each identified candidate behavior, a probability that the target person's behavior corresponds to the identified candidate behavior. In this case, the identified candidate behavior and the derived probability may be included in the behavior recognition information.
[0048] The behavior recognition unit F12 can refer to image data of the input image IN in the behavior recognition processing. The behavior recognition unit F12 may perform the behavior recognition processing for each input image IN based on the image data of the input image IN and output information (BBOX information, class information, and skeletal information) of the detection unit F11 for the input image IN. That is, for example, in the behavior recognition processing for the input image 600 (see FIG. 8 ), the behavior recognition unit F12 may recognize the behavior of the person 610 (target person) based on the output information (BBOX information, class information, and skeletal information) of the detection unit F11 for the input image 600 and the image data of the input image 600. In the behavior recognition processing for a certain input image IN that has been focused on, image data of multiple input images IN including the focused input image IN may be referenced. That is, for example, in the behavior recognition processing for the input image 600, the behavior recognition unit F12 may recognize the behavior of the person 610 (target person) based on the output information of the detection unit F11 for the input image 600 and the image data of the input image sequence that includes the input image 600. The input image sequence including the input image 600 refers to a plurality of input images arranged in time series obtained by capturing images multiple times with the camera 51, and including the input image 600.
[0049] Fig. 9 shows an example of the internal configuration of the action recognition unit F12. The action recognition unit F12 in Fig. 9 has three functional blocks: an action understanding unit F12a, an action prediction unit F12b, and an action evaluation unit F12c. An input image IN at a certain time t is particularly referred to as an input image IN_t. Output information of the detection unit F11 is referred to as output information DET, and output information of the detection unit F11 for the input image IN_t is particularly referred to as output information DET_t. The output information DET_t includes BBOX information, class information, and skeleton information generated by the detection unit F11 by object detection and skeleton detection for the input image IN_t.
[0050] The behavior understanding unit F12a understands the current behavior of the target person based on the output information DET, and generates and outputs current behavior information A indicating the results of this understanding. Understanding in the behavior understanding unit F12a corresponds to understanding the current behavior of the target person. The current behavior information A generated based on the output information D_t is particularly referred to as current behavior information A_t. The present with respect to the current behavior information A_t refers to the time t. In other words, the current behavior information A_t represents the results of understanding the behavior of the target person at time t. The behavior understanding unit F12a may generate the current behavior information A_t not only based on the output information DET_t, but also based on image data of the input image IN_t or image data of an input image sequence including the input image IN_t.
[0051] The behavior prediction unit F12b predicts the future behavior of the target person based on the output information DET, and generates and outputs predicted behavior information FA indicating the result of the prediction. Here, the future is defined as a unit time T UNIT It refers to a point in time after time t. UNIT The time after that is called time (t+1). UNIT is a time period predetermined by the behavior prediction unit F12b, for example, 1 second. The predicted behavior information FA generated based on the output information DET_t is particularly referred to as predicted behavior information FA_t+1. The predicted behavior information FA_t+1 represents the prediction result of the behavior prediction process performed at time t, and is a prediction of the behavior of the target person at time (t+1) at time t. That is, at time t, the behavior prediction unit F12b calculates the predicted behavior information FA_t+1 from the time t by calculating the unit time T UNIT The behavior prediction unit F12b generates predicted behavior information FA_t+1 by predicting the behavior of the target person in the future. The behavior prediction unit F12b may generate the predicted behavior information FA_t+1 based not only on the output information DET_t but also on the image data of the input image IN_t or on the image data of an input image sequence including the input image IN_t.
[0052] The behavior evaluation unit F12c generates the above-mentioned behavior recognition information by evaluating the behavior of the target person based on the current behavior information A and the predicted behavior information FA. The behavior recognition information based on the current behavior information A_t and the predicted behavior information FA_t+1 represents the recognition result of the behavior of the target person at time t.
[0053] ---Sampling section F13--- The functional block F13 is a sampling unit (see FIG. 7). The sampling unit F13 generates a thinned video MVb from the input video MVa by performing the sampling process described above (see FIGS. 6(a) to 6(c)). The thinned video MVb corresponds to the video output by the image processing unit 11a. The n input images IN arranged in time series are frame images Fa[1] to Fa[n]. The sampling unit F13 also determines which of the input images IN generated sequentially in time series will be set as frame image Fa[1] and which other input image IN will be set as frame image Fa[n] (details will be described later). The input video MVa is composed of n frame images Fa, while the thinned video MVb is composed of fewer than n frame images Fa (for example, n / 2, n / 4, or n / 8 frame images Fa).
[0054] ---Object feature information, image data--- Information including the BBOX information, class information, and skeleton information generated by the detection unit F11 and the behavior recognition information generated by the behavior recognition unit F12 is referred to as object feature data. The image processing unit 11a generates object feature data for each input image IN. Therefore, the image processing unit 11a generates object feature data for each of the frame images Fa[1] to Fa[n]. With reference to FIG. 10, the object feature data for the frame images Fa[1] to Fa[n] are particularly referred to as object feature data D[1] to D[n], respectively. The image processing unit 11a generates the object feature data D[1] to D[n] based on the frame images Fa[1] to Fa[n].
[0055] The image processing unit 11a (detection unit F11) performs skeleton detection on the detection object in each frame image Fa based on the image data of the frame images Fa[1] to Fa[n]. The image processing unit 11a (detection unit F11 and behavior recognition unit F12) then generates object feature data including the results of skeleton detection for each frame image Fa. That is, the image processing unit 11a (detection unit F11 and behavior recognition unit F12) generates object feature data D[i] including the results of skeleton detection on the detection object in the frame image Fa[i] based on the image data of the frame image Fa[i].
[0056] In skeleton detection for each frame image Fa, the image processing unit 11a detects the positions of multiple specific parts of the detection object that are also the positions in that frame image Fa, and includes the detection results of those positions in the skeleton information (and therefore in the object feature data). Therefore, in skeleton detection for frame image Fa[i], the positions of multiple specific parts of the detection object that are also the positions in frame image Fa[i] are detected, and the detection results of those positions are included in the skeleton information for frame image Fa[i] (and therefore in the object feature data D[i]).
[0057] Furthermore, the image processing unit 11a (behavior recognition unit F12) recognizes the behavior of the detection object for each frame image Fa in the behavior recognition process for the input moving image MVa, and derives the recognition result for each frame image Fa as behavior recognition information. The behavior recognition information is included in the corresponding object feature data. That is, the image processing unit 11a (behavior recognition unit F12) recognizes the behavior of the detection object in frame image Fa[i] in the behavior recognition process for frame image Fa[i]. The behavior of the detection object in frame image Fa[i] is, in other words, the behavior of the detection object at time t[i]. The image processing unit 11a (behavior recognition unit F12) then derives the recognition result of the behavior of the detection object for frame image Fa[i] as behavior recognition information for frame image Fa[i], and includes it in the object feature data D[i].
[0058] The image processing unit 11a can perform behavior recognition processing on frame image Fa[i] based on the image data of frame image Fa[i] and the output information (BBOX information, class information, and skeleton information) of the detection unit F11 for frame image Fa[i]. In the behavior recognition processing on frame image Fa[i], the image processing unit 11a may refer to image data of a frame image sequence that includes frame image Fa[i]. A frame image sequence that includes frame image Fa[i] refers to a plurality of frame images that are arranged in time series obtained by multiple captures by the camera 51 and that include frame image Fa[i], such as frame images Fa[ij] to Fa[i] (j is an integer greater than or equal to 1).
[0059] Furthermore, as can be understood from the above, the object feature data for each frame image Fa includes position information (BBOX information) of the detection target object in that frame image Fa. That is, the object feature data D[i] of frame image Fa[i] includes position information (BBOX information) of the detection target object in frame image Fa[i]. However, the object feature data for each frame image Fa does not have to include position information (BBOX information) of the detection target object. This is because the position of the detection target object in the frame image Fa is identified by the skeletal information of the detection target object.
[0060] The image processing unit 11a then generates image set data including image data of the thinned video MVb (image data of the output video) and object feature data D[1] to D[n]. In the image set data, the image data of the thinned video MVb and the object feature data D[1] to D[n] are associated with each other. In the image set data, the object feature data D[1] to D[n] are associated with times t[1] to t[n], respectively.
[0061] In the image set data, each frame image constituting the thinned video MVb is associated with the capture time of the frame image. That is, for example, if the thinned video MVb includes frame image Fa[1], frame image Fa[1] in the thinned video MVb is associated with time t[1]. Similarly, if the thinned video MVb includes frame image Fa[5], frame image Fa[5] in the thinned video MVb is associated with time t[5]. The same applies when the thinned video MVb includes other frame images Fa. When frame image Fa[i] is included among the multiple frame images constituting the thinned video MVb, timestamp information indicating that frame image Fa[i] was captured at time t[i] may be added to the image set data.
[0062] ---Image restoration processing--- The controller 11 transmits the image set data generated by the image processing unit 11a to the data processing device 200 (see FIG. 1) via the communication network NET using the communication unit 13. In the data processing device 200, the image set data is received by the communication unit 230, and the received image set data is input to the image processing unit 210a (see FIG. 5). When transmitting and receiving the image set data, well-known data compression and data decompression for communication may be performed. The image processing unit 210a performs image restoration processing on the image set data to restore the input video MVa. The restored input video MVa is referred to as restored video MVc. In many cases, the restored video MVc is similar to the original input video MVa generated by the camera 51, but does not completely match the original input video MVa.
[0063] In the image restoration process, the image processing unit 210a estimates the movement of the detection object between two adjacent frame images in the thinned video MVb based on the object feature data D[1] to D[n]. Then, of the frame images Fa[1] to Fa[n], a frame image Fa that is not included in the thinned video MVb is generated based on the result of the estimation, thereby generating a restored video MVc.
[0064] 11 is a functional block diagram of the image processing unit 210a. The image processing unit 210a includes functional blocks F21 and F22. The controller 210 including the image processing unit 210a is a program execution device (computer) capable of executing any program. All or part of the functions of the functional blocks F21 and F22 may be realized by the controller 210 executing a program recorded in the memory 220 or any other recording medium.
[0065] The functional block F21 is an encoder, and the functional block F22 is a decoder. Object feature data D[1] to D[n] is input to the encoder F21 as an input sequence. The encoder F21 extracts features of the input sequence and inputs the extracted information to the decoder F22. Image data of the thinned video MVb is also input to the decoder F22. The decoder F22 generates a new sequence as a restored video MVc (more specifically, image data of the restored video MVc) based on the information extracted by the encoder F21 and the thinned video MVb (more specifically, image data of the thinned video MVb). The controller 210 can store the image data of the generated restored video MVc in a database DB.
[0066] The encoder F21 and the decoder F22 are a behavioral encoder AI and a generative AI formed by machine learning a deep neural network (DNN). The encoder F21 and the decoder F22 can be formed using a model commonly known as a Transformer (Vision Transformer).
[0067] Below, several specific operational examples, application techniques, modified techniques, etc. relating to each component shown in FIG. 1 and other figures will be described in multiple embodiments. The matters described above in this embodiment are applied to each of the following embodiments unless otherwise specified and unless there is a contradiction. If there are any matters in each embodiment that contradict the matters described above, the description in that embodiment may take precedence. Furthermore, unless there is a contradiction, matters described in any of the multiple embodiments shown below can also be applied to any of the other embodiments (i.e., any two or more of the multiple embodiments can be combined).
[0068] <<First Example>> A first embodiment will be described. FIG. 12 shows a specific example of how a restored video MVc is generated from a thinned video MVb based on an input video MVa. In the first embodiment, it is assumed that each of the frame images Fa[1] to Fa[n] includes an image of a person 710, and that the person 710 is detected as a detection target in object detection for each of the frame images Fa[1] to Fa[n]. Therefore, the image processing unit 11a generates object feature data D[1] to D[n] related to the person 710. The object feature data D[i] related to the person 710 includes BBOX information, class information, skeletal information, and behavior recognition information of the person 710 in the frame image Fa[i]. For ease of illustration, FIG. 12 shows only the portion of each frame image Fa in which the image of the person 710 appears. Also, to avoid cluttering the illustration, the reference numeral "710" is assigned to only the person 710 in some of the frame images Fa in FIG. 12, but all the people shown in FIG. 12 represent the same person (710).
[0069] In the example of FIG. 12, the sampling interval when generating the thinned video MVb from the input video MVa is (4 / RTa). (4 / RTa) is four times the reciprocal of the frame rate RTa of the input video MVa (see FIG. 6(b)). The sampling unit F13 samples n frame images Fa (Fa[1] to Fa[n]) at a fixed sampling interval (4 / RTa) starting from frame image Fa[1]. The sampling unit F13 then generates a collection of the sampled frame images Fa as the thinned video MVb. Since sampling is performed starting from frame image Fa[1], the first frame image of the thinned video MVb is frame image Fa[1]. Because the sampling interval is (4 / RTa), the thinned-out video MVb includes frame images Fa[1], Fa[5], and Fa[9], but does not include frame images Fa[2] to Fa[4] and Fa[6] to Fa[8] (assuming here that n≧9). In the example of FIG. 12, n is a value obtained by adding 1 to a multiple of 4 (e.g., 129). Therefore, the total number of frame images Fa that make up the thinned-out video MVb is (((n−1) / 4)+1).
[0070] Image set data consisting of image data of the thinned video MVb and object feature data D[1] to D[n] is transmitted from the in-vehicle device 10 and received by the data processing device 200. In the data processing device 200, the controller 210 inputs the received image set data to the image processing unit 210a. The image processing unit 210a generates image data of the restored video MVc based on the image set data using encoders F21 and F22 (see FIG. 11). The restored video MVc is made up of n still images arranged in chronological order. Each still image constituting the restored video MVc is called a frame image Fc.
[0071] The restored video sequence MVc has n frame images Fc[1] to Fc[n]. Similar to the input video sequence MVa (see FIG. 6(b)), the frame images Fc[1] to Fc[n] in the restored video sequence MVc are arranged in chronological order at intervals equal to the reciprocal of the frame rate RTa. In the restored video sequence MVc, the frame images Fc[1] to Fc[n] correspond to times t[1] to t[n], respectively.
[0072] For any integer i that satisfies "1≦i≦n", if the thinned-out video MVb includes a frame image Fa[i], the image processing unit 210a generates a still image having the same content (image data) as the content (image data) of frame image Fa[i] as frame image Fc[i]. Therefore, for any integer i that satisfies "1≦i≦n", if the thinned-out video MVb includes a frame image Fa[i], frame image Fc[i] is the same image as frame image Fa[i]. 12, frame image Fc[1] is the same as frame image Fa[1], frame image Fc[5] is the same as frame image Fa[5], and so on.
[0073] For any integer i satisfying "1≦i≦n," if the thinned-out video MVb does not include a frame image Fa[i], the image processing unit 210a generates a frame image Fc[i] through image restoration processing. In this case, the content (image data) of frame image Fc[i] is expected to match or generally match the content (image data) of frame image Fa[i], but the former and latter contents may differ in detail. In other words, in the example of FIG. 12, the content (image data) of frame image Fc[2] is expected to match or generally match the content (image data) of frame image Fa[2], but the former and latter contents may differ in detail. Here, the match between the content of frame image Fc[2] and the content of frame image Fa[2] refers to the match between the image of person 710 in frame image Fc[2] and the image of person 710 in frame image Fa[2]. The same applies to frame image Fc[3], etc.
[0074] As described above, the image processing unit 11a of the in-vehicle device 10 generates object feature data D[1] to D[n] including skeletal information through skeletal detection for each frame image Fa, and generates a thinned video MVb from the input video MVa by thinning out the frame images Fa. Then, image set data including the image data of the thinned video MVb and the object feature data D[1] to D[n] is generated. The data volume of each object feature data is much smaller than the data volume per frame image Fa. Therefore, the data volume of the image set data can be made smaller than the data volume of the input video MVa by roughly the amount of data for the frame images Fa[1] to Fa[n] that were not included in the thinned video MVb. Although some frame images are missing from the thinned video MVa, the movement of the detected object can be reproduced based on the image set data including the results of skeletal detection. In other words, by using the in-vehicle device 10, it becomes possible to reproduce the movement of the detected object on the receiving device (data processing device 200) of the image set data, thereby reducing the communication load compared to when all image data of the input moving image MVa is sent and received.
[0075] Since each key point (position of a specific part) is detected by skeleton detection, it becomes possible to reproduce the movement of the detected object on the receiving side device of the image set data (data processing device 200).
[0076] Furthermore, since the object feature data includes the recognition results of the behavior of the detected object in each frame image Fa, it is possible to accurately reproduce the movement of the detected object by using the recognition results on the receiving device (data processing device 200) of the image set data.
[0077] The frame images Fa[1] to Fa[n] are sampled at a fixed interval to generate the thinned video MVb, which reduces the amount of data in the thinned video MVb depending on the sampling interval. Furthermore, the adoption of a fixed sampling interval allows for uniform reproduction accuracy of the motion of the detected object across the entire restored video MVc.
[0078] The data processing device 200, which is the device on the receiving side of the image set data, can generate a restored video sequence MVc based on the image set data. That is, the movement of the detection object can be reproduced in the data processing device 200 while enjoying the effect of reducing the communication load as described above.
[0079] FIG. 13 shows an operation flowchart of the controller 11 in the in-vehicle device 10. Each process of steps S11 to S18 shown in FIG. 13 is executed by the controller 11 (mainly the image processing unit 11a). When the vehicle V1 starts, the controller 11 and the camera 51 start, and the operation of the controller 11 starts from the process of step S11. When the camera 51 starts, sequential acquisition of frame images starts in step S11. Sequential acquisition of frame images means that image data of each frame image generated by the camera 51 through sequential shooting by the camera 51 is sequentially acquired by the controller 11 via the in-vehicle network. Although not particularly shown in FIG. 13, from step S11 onwards, the above-mentioned recording process as a drive recorder is executed by the controller 11.
[0080] In step S12 following step S11, the image processing unit 11a starts executing transmission-side image processing on each frame image captured by the camera 51. The transmission-side image processing includes the object detection, skeleton detection, and behavior recognition processing described above. The image processing unit 11a may perform transmission-side image processing on all frame images captured by the camera 51, or may perform transmission-side image processing on only some of all frame images captured by the camera 51. The transmission-side image processing is performed on at least each frame image Fa forming the input moving image MVa.
[0081] In step S13 following step S12, the controller 11 (or the image processing unit 11a) monitors whether a predetermined start condition is met. If the start condition is met (Y in step S13), the controller 11 (or the image processing unit 11a) causes the process to proceed to step S14, whereas if the start condition is not met (N in step S13), the controller 11 continues monitoring step S13.
[0082] The start condition may be set in any manner. For example, the start condition may be established when a detection target (particularly a person) is detected by object detection. In more detail, the start condition may be established when, for example, a transition from a state in which a detection target (particularly a person) is not detected by object detection to a state in which a detection target (particularly a person) is detected by object detection. Alternatively, for example, the start condition may be established each time a certain time period has elapsed after the controller 11 is started (provided, however, that a detection target is detected in each image captured by the camera 51 by object detection).
[0083] In step S14, the controller 11 (or the image processing unit 11a) sets the start time of the input video MVa based on the time when the start condition is met. The start time of the input video MVa may be set to the time when the start condition is met, or may be set to a time a predetermined time after the time when the start condition is met. When one input video MVa is focused on, the start time of the input video MVa is time t[1] (see FIG. 10). The start time of the input video MVa may also be set to a time a predetermined time before the time when the start condition is met. The controller 11 can store image data of images captured by the camera 51 in the memory 12 for a certain temporary time, and can read image data of past captured images for the temporary time from the memory 12.
[0084] In step S15 following step S14, the controller 11 (or the image processing unit 11a) monitors whether a predetermined termination condition is met. If the termination condition is met (Y in step S15), the controller 11 (or the image processing unit 11a) causes the process to proceed to step S16, whereas if the termination condition is not met (N in step S15), the controller 11 continues monitoring step S15.
[0085] The method for setting the termination condition is arbitrary. Typically, for example, the termination condition may be established when a predetermined shooting time has elapsed since the start time of the input video MVa. When adopting a method in which the start condition is established when a detection target (particularly a person) is detected by object detection, the termination condition may be established when the detection target (particularly a person) is no longer detected by object detection. In detail, for example, the establishment of the termination condition may correspond to a transition from a state in which a detection target (particularly a person) is detected by object detection to a state in which the detection target (particularly a person) is not detected by object detection. In this case, the image processing unit 11a can determine the transition from the former state to the latter state by tracking and identifying the detection target within the input video MVa based on the image data of the input video MVa. Note that an upper limit on the length of the input video MVa is set, and therefore the termination condition is always established when a specified maximum time has elapsed since the start time of the input video MVa.
[0086] The controller 11 (or image processing unit 11a) sets the most recent captured image generated by the camera 51 at the time the termination condition is met as frame image Fa[n]. Alternatively, the controller 11 (or image processing unit 11a) may set the captured image generated by the camera 51 immediately after the termination condition is met as frame image Fa[n]. After step S15, the process proceeds to step S16, and the processes of steps S16 to S18 are executed. At the stage of proceeding to step S16, the image data of the input moving image MVa made up of frame images Fa[1] to Fa[n] has already been acquired by the controller 11.
[0087] In step S16, the image processing unit 11a (sampling unit F13) performs the above-mentioned sampling process to generate a thinned video MVb, which is an output video of the image processing unit 11a, from the input video MVa. Then, the process proceeds to step S17. From step S14 to step S17, the image processing unit 11a (detection unit F11 and behavior recognition unit F12) performs transmission-side image processing on each of the frame images Fa[1] to Fa[n] to generate object feature data D[1] to D[n]. In step S17, the image processing unit 11a generates image set data, which is a combination of the object feature data D[1] to D[n] and the image data of the thinned video MVb. Then, in step S18, the controller 11 wirelessly transmits the image set data to the data processing device 200 via the communication network NET. After step S18, the process returns to step S13, and the first processing group consisting of steps S13 to S18 is repeated. When the first processing group is executed for the jth time, the jth thinned-out video MVb based on the jth input video MVa is generated, and the jth image set data is transmitted to the data processing device 200 (where j is any natural number).
[0088] 14 shows an operation flowchart of the controller 210 in the data processing device 200. When the data processing device 200 is started, the operation of the controller 210 starts from the processing of step S21. In step S21, the controller 210 waits for reception of image set data from the in-vehicle device 10. The processing of step S21 is repeated until the image set data is received, and when the image set data is received (Y in step S21), the process moves from step S21 to step S22.
[0089] In step S22, the image processing unit 210a generates a restored video MVc by the above-mentioned image restoration process based on the received image set data, i.e., generates image data of the restored video MVc. After step S22, in step S23, the controller 210 stores the image data of the generated restored video MVc in the database DB. Then, the process returns to step S21, and the second process group consisting of steps S21 to S23 is repeated. When the second process group is executed for the jth time, the jth restored video MVc is generated based on the jth image set data corresponding to the jth thinned-out video MVb (where j is an arbitrary natural number).
[0090] The restored video MVc can be used in any way. For example, the restored video MVc can be used to verify how the vehicle V1 and the pedestrian behaved when there was a pedestrian in front of the vehicle V1. The restored video MVc can also be used to provide feedback to drivers (fleet drivers, etc.) for safe driving guidance.
[0091] <<Second Example>> A second embodiment will be described. In the second embodiment, a method for generating an encoder F21 and a decoder F22 that implement image restoration processing by machine learning will be described.
[0092] FIG. 15 shows the internal configuration of a machine learning device 300 that performs the machine learning. The machine learning device 300 includes a controller 310, a memory 320, and a communication unit 330. The controller 310 includes a processing unit including a CPU, a GPU, and the like as hardware resources. The controller 310 is a program execution device (computer) capable of executing any program. The controller 310 may implement any function, operation, or process to be realized by the controller 310 by executing a program recorded in the memory 320 or any other recording medium. The memory 320 includes a nonvolatile memory such as a ROM or flash memory, and a volatile memory such as a RAM. The memory 320 stores various data referenced by the controller 310 as well as various programs to be executed by the controller 310. The communication unit 330 transmits and receives any signal between the machine learning device 300 and a different device.
[0093] The machine learning device 300 performs machine learning based on training data. The training data includes a large amount of image data of training videos. For example, hundreds of thousands of training videos are prepared. Videos captured by a camera 51 installed on the vehicle V1 or a camera installed on another vehicle may be used as training videos. The machine learning device 300 is connected to a communication network NET via a wired or wireless connection, and the machine learning device 300 may acquire image data of several training videos from other devices connected to the communication network NET. Each frame image constituting the training video is assumed to contain an image of the detection target.
[0094] The controller 310 is provided with an image processing unit 311 equivalent to the image processing unit 11a. The image processing unit 311 has a detection unit, a behavior recognition unit, and a sampling unit equivalent to the detection unit F11, the behavior recognition unit F12, and the sampling unit F13 in FIG. 7, and generates image set data for each training video based on the image data of the training video. The method for generating image set data based on the image data of the training video by the image processing unit 311 is the same as the method for generating image set data based on the image data of the input video MVa by the image processing unit 11a. The image set data generated by the image processing unit 311 includes image data of a thinned video generated by applying a sampling process to the training video, and object feature data corresponding to the training video. The object feature data corresponding to the training video consists of a collection of object feature data derived for each frame image in the training video.
[0095] Hereinafter, thinned videos generated by applying sampling processing to training videos will be particularly referred to as training thinned videos (see FIG. 16). Object feature data corresponding to training videos will be referred to as training object feature data. Image set data consisting of training thinned videos and training object feature data will be referred to as training image set data. Additionally, training videos will be referred to as training input videos.
[0096] The controller 310 is also provided with an image processing unit 312. As shown in Fig. 16, the image processing unit 312 includes an encoder 312a and a decoder 312b that are the basis for the encoder F21 and the decoder F22 in Fig. 11. The encoder 312a and the decoder 312b are each configured using a DNN (deep neural network).
[0097] In machine learning, the controller 310 causes the image processing unit 312 to execute the following learning arithmetic processing for each set of learning image data. As a result, the image processing unit 312 generates restored learning videos corresponding to videos restored from the input learning videos. In the learning arithmetic processing, object feature data for learning included in the set of learning image data is input as an input sequence to the encoder 312a. In the learning arithmetic processing, the encoder 312a extracts features of the input sequence and inputs the extracted information to the decoder 312b. Image data of the thinned learning videos is also input to the decoder 312b. In the learning arithmetic processing, the decoder 312b generates a new sequence as restored learning videos (more specifically, image data of the thinned learning videos) based on the information extracted by the encoder 312a and the thinned learning videos (more specifically, image data of the thinned learning videos).
[0098] In machine learning, the controller 310 updates the parameters (weights and biases) of the encoder 312a and the decoder 312b based on the error between the generated restored learning video sequence and the learning input video sequence. In machine learning, the controller 310 repeats this update as many times as necessary. The encoder 312a and the decoder 312b after machine learning are incorporated into the image processing unit 210a of the data processing device 200 as the encoder F21 and the decoder F22 (see FIG. 11).
[0099] <<Third Example>> A third embodiment will be described. After the start time of the input video MVa (i.e., time t[1]), the number of detection objects detected from the input video MVa may increase during the period leading up to the end time of the input video MVa (i.e., time t[n]). In the third embodiment, a method for generating a thinned video MVb that can be applied in such a case will be described.
[0100] Please refer to Figure 17. Figure 17 shows the relationship between the input video sequence MVa and the thinned video sequence MVb assumed in the third embodiment. For convenience of illustration, the contents of each frame image Fa are omitted from Figure 17. Also, to avoid cluttering the illustration, Figure 17 omits some of the time symbols "t[1] to t[n]", some of the frame image symbols "Fa[1] to Fa[n]", and some of the object feature data symbols "D[1] to D[n]". As described above, the input video sequence MVa is made up of frame images Fa[1] to Fa[n].
[0101] In the example of Fig. 17, it is assumed that an image of a first person is included in each of the frame images Fa[1] to Fa[n], and that the first person is detected as a detection target in object detection for each of the frame images Fa[1] to Fa[n]. Furthermore, in the example of Fig. 17, it is assumed that an image of a second person is included in each of the frame images Fa[k+1] to Fa[n], and that the second person is detected as another detection target in object detection for each of the frame images Fa[k+1] to Fa[n]. The second person is a different person from the first person. In the example of Fig. 17, an image of the second person is not included in each of the frame images Fa[1] to Fa[k], and therefore the second person is not detected in object detection for the frame images Fa[1] to Fa[k].
[0102] For each integer i that satisfies "1≦i≦k," the image processing unit 11a generates object feature data D[i] by performing object detection, skeletal detection, and behavior recognition processing based on the image data of frame image Fa[i]. For frame image Fa[i] that satisfies "1≦i≦k," a first person is detected by object detection, and skeletal detection and behavior recognition processing are performed on the first person, thereby deriving skeletal information and behavior recognition information for the first person. Therefore, the object feature data D[i] for "1≦i≦k" includes BBOX information, class information, skeletal information, and behavior recognition information related to the first person in frame image Fa[i]. Since a second person is not detected in object detection for frame images Fa[1] to Fa[k], each of the object feature data D[1] to D[k] does not include BBOX information, class information, skeletal information, and behavior recognition information related to the second person.
[0103] For each integer i that satisfies "k+1≦i≦n," the image processing unit 11a generates object feature data D[i] by performing object detection, skeletal detection, and behavior recognition processing based on the image data of frame image Fa[i]. For frame image Fa[i] that satisfies "k+1≦i≦n," a first person is detected by object detection, and skeletal detection and behavior recognition processing are performed on the first person, thereby deriving skeletal information and behavior recognition information for the first person. Therefore, the object feature data D[i] for "k+1≦i≦n" includes BBOX information, class information, skeletal information, and behavior recognition information for the first person in frame image Fa[i]. In addition, for frame image Fa[i] that satisfies "k+1≦i≦n," a second person is also detected by object detection, and skeletal detection and behavior recognition processing are performed on the second person, thereby deriving skeletal information and behavior recognition information for the second person. Therefore, the object feature data D[i] in "k+1≦i≦n" also includes BBOX information, class information, skeleton information, and behavior recognition information related to the second person in the frame image Fa[i].
[0104] In the example of FIG. 17, the sampling interval when generating the thinned video MVb from the input video MVa is (4 / RTa). (4 / RTa) is four times the reciprocal of the frame rate RTa of the input video MVa (see FIG. 6(b)). The sampling unit F13 samples n frame images Fa (Fa[1] to Fa[n]) at a fixed sampling interval (4 / RTa) starting from frame image Fa[1]. The sampling unit F13 then generates a collection of the sampled frame images Fa as the thinned video MVb. Since sampling is performed starting from frame image Fa[1], the first frame image of the thinned video MVb is frame image Fa[1].
[0105] However, in the example of FIG. 17, after the start time t[1] of the input video MVa, the number of detection objects detected from the input video MVa increases from 1 to 2 during the period up to the end time t[n] of the input video MVa. Accordingly, the following sampling is performed. Specifically, the sampling unit F13 in the example of FIG. 17 samples frame images Fa[1] to Fa[k] starting from frame image Fa[1] at a fixed sampling interval (4 / RTa). Thereafter, the sampling unit F13 in the example of FIG. 17 samples frame images Fa[k+1] to Fa[n] starting from frame image Fa[k+1] at a fixed sampling interval (4 / RTa). The sampling unit F13 then generates a collection of the sampled frame images Fa as a thinned-out video MVb. For frame images Fa[k+1] to Fa[n] in which the first and second persons are detected by object detection, sampling is performed starting from frame image Fa[k+1], and so frame image Fa[k+1] is included in the thinned-out video MVb.
[0106] In this case, the thinned-out video MVb includes frame images Fa[1], Fa[5]..., Fa[k-5], and Fa[k-1], but does not include frame images Fa[2] to Fa[4], Fa[k-4] to Fa[k-2], or Fa[k]. The thinned-out video MVb also includes frame images Fa[k+1], Fa[k+5]..., Fa[n-4], and Fa[n], but does not include frame images Fa[k+2] to Fa[k+4] or Fa[n-3] to Fa[n-1]. Note that in the example of FIG. 17, it is assumed that (k-1) has an integer value that is one greater than a multiple of 4, and that n has an integer value that is one less than a multiple of 4. For example, (k-1, k+1, n) = (21, 23, 43).
[0107] Through object detection, only the first person is detected from frame image Fa[k], while the first and second people are detected from frame image Fa[k+1]. That is, the total number of detected objects detected by object detection from one frame image Fa increases from frame image Fa[k] to frame image Fa[k+1]. Based on the object detection results of the detection unit F11, the sampling unit F13 identifies the frame image Fa[k+1] immediately after this increase as the change origin image. Then, after sampling from frame image Fa[1] at a fixed sampling interval (4 / RTa), if the change origin image is identified, the sampling unit F13 samples the group of frame images Fa after the change origin image at a fixed sampling interval (4 / RTa) from the change origin image.
[0108] The conditions for the values of k and n to which the method described in the third embodiment is applied are as follows. It is generally assumed that the sampling unit F13 will sample two or more frame images Fa from among the frame images Fa[1] to Fa[k]. The minimum sampling interval is (2 / RTa). For this reason, k should preferably represent an integer of 3 or greater. In addition, it is generally assumed that the sampling unit F13 will sample two or more frame images Fa from among the frame images Fa[k+1] to Fa[n]. For this reason, it is generally assumed that n should represent an integer of (k+3) or greater.
[0109] The above-mentioned first and second persons are examples of the first and second detection objects. At least one of the first and second detection objects may be an animal (such as a dog or a cat).
[0110] According to the method of the third embodiment, if a second object is detected in the input video sequence MVa after sampling has started, the sampling timing is reset at the point in time when the second object is detected. Therefore, image information from the initial detection of the second object is included in the thinned video sequence MVb. This is believed to contribute to improving the accuracy of reproducing the movement of the second object in the image restoration process.
[0111] <<Fourth Example>> A fourth embodiment will be described. From the viewpoint of privacy protection, the facial image portion of each frame image Fa may be processed in the process of generating a thinned video sequence MVb from an input video sequence MVa. This will be described in detail.
[0112] In the fourth embodiment, a person is a detection target detected by object detection. In the fourth embodiment, in order to clearly distinguish between frame images before and after processing, each frame image constituting the input moving image MVa is referred to as frame image Fa or Fa[i], and each frame image constituting the thinned moving image MVb is referred to as frame image Fb or Fb[i] (i is an arbitrary integer). Processing is performed by processing. Each frame image constituting the input moving image MVa has not been processed. The image processing unit 11a performs processing on frame image Fa[i] in the input moving image MVa, and sets the processed frame image Fa[i] as frame image Fb[i].
[0113] In the processing of frame image Fa[i], the facial image of the person in frame image Fa[i] is deleted. In the processing, for example, a predetermined illustration image or black image is filled in the deleted portion. Alternatively, the processing of frame image Fa[i] may be a mosaic process or a blur process on the facial image of the person in frame image Fa[i].
[0114] An explanation will be given using the input video MVa and thinned video MVb in FIG. 12 as examples. For simplicity of explanation, let us assume that "n=9." In this case, the input video MVa in the fourth embodiment is made up of frame images Fa[1] to Fa[9], and the thinned video MVb in the fourth embodiment is made up of frame images Fb[1], Fb[5], and Fb[9]. The image processing unit 11a applies processing to frame image Fa[1] in the input video MVa, and sets the processed frame image Fa[1] as frame image Fb[1]. Similarly, the image processing unit 11a applies processing to frame image Fa[5] in the input video MVa, and sets the processed frame image Fa[5] as frame image Fb[5]. The same applies to the set of frame images Fa[9] and Fb[9].
[0115] By including the processed thinned-out video MVb in the image set data and transmitting it to the data processing device 200, the privacy of people in the input video MVa can be protected.
[0116] <<Fifth Example>> A fifth embodiment will be described. The controller 11 may issue various notifications based on the object feature data, in addition to generating the thinned video MVb. The optional notification may be a notification by displaying a video on the display device 61 (a notification that affects the user U1's visual sense), a notification by outputting a sound from the speaker 62 (a notification that affects the user U1's auditory sense), or a combination thereof. If the HMI 60 includes the vibration device, the optional notification may include a notification by generating a vibration from the vibration device (a notification that affects the user U1's tactile sense).
[0117] For example, the controller 11 may determine whether there is a possibility that a pedestrian or the like may jump out in the traveling direction of the vehicle V1 based on the object feature data, and if it determines that there is such a possibility, may issue a jumping-out warning notice to the user U1. The jumping-out warning notice is a notice that conveys (warns) to the user U1 that there is a possibility that a pedestrian or the like may jump out in the traveling direction of the vehicle V1.
[0118] Furthermore, when the controller 11 determines that there is a possibility that a pedestrian or the like may jump out in the traveling direction of the vehicle V1, the controller 11 may transmit the determination result to the driving control device 20. The driving control device 20 may perform driving control of the vehicle V1 based on the determination result (driving control to avoid contact between the vehicle V1 and the pedestrian or the like).
[0119] <<Sixth Example>> A sixth embodiment will be described. An in-vehicle device 10 incorporates a video processing device according to the present invention. The video processing device includes a controller 11. The in-vehicle device 10 itself can be considered to be a video processing device, or a part of the in-vehicle device 10 can be considered to correspond to the video processing device. In the above explanations, a vehicle V1 and an in-vehicle system 1 are assumed as application examples of the video processing device. However, application of the video processing device according to the present invention is not limited to the vehicle V1 and the in-vehicle system 1.
[0120] For example, the moving image processing device according to the present invention can be applied to verifying and analyzing the form of a player in sports. In this case, the player corresponds to user U1. More specifically, for example, if user U1 is a baseball batter, the input moving image MVa is generated by capturing the user U1 swinging the bat with camera 51. In this case, for example, the time when the pitcher throwing the ball to user U1 releases the ball may be set as time t[1]. The controller 210 on the data processing device 200 side may evaluate the appropriateness of the form of user U1 based on the restored moving image MVc.
[0121] Alternatively, for example, the video processing device according to the present invention can be applied to verifying and analyzing the working status of a worker in a factory. In this case, the worker corresponds to user U1, and a camera that captures the worker corresponds to camera 51. The input video MVa is generated by capturing the worker's work with camera 51. The controller 210 on the data processing device 200 side may evaluate the safety of the worker's work or the worker's work efficiency, etc., based on the restored video MVc.
[0122] <<Seventh Example>> A seventh embodiment will now be described.
[0123] A program that causes a computer device to execute any of the methods described in the embodiments of the present invention, and a non-volatile recording medium on which the program is recorded, are included within the scope of the embodiments of the present invention. The program that causes a computer device to execute any of the methods described in the embodiments of the present invention may be a subprogram incorporated into any main program or called by any main program. Any processing in the embodiments of the present invention may be realized by hardware such as a semiconductor integrated circuit, software equivalent to the program, or a combination of hardware and software.
[0124] The in-vehicle device 10 is a type of computer device. A method related to image processing executed by the in-vehicle device 10 or a method performed by the moving image processing device according to the present invention can be referred to as a moving image processing method. A program (image processing program) for causing a computer device to execute the moving image processing method can be formed. A system including the in-vehicle device 10, the camera 51, and the data processing device 200 can be referred to as an image processing system. [Explanation of symbols]
[0125] 1. In-vehicle systems V1 vehicle U1 User TM Information Terminal ST1 seat NET communication network 10 Onboard equipment 11 Controller 11a Image processing unit 12 Memory 13 Communications Department 14 Recording media 20 Driving control device 30 Actuator section 40 Vehicle sensor unit 50 Camera Department 51 Camera 60 HMI 61 Display device 62 Speaker 63 Operation input section 200 Data processing device 210 Controller 210a Image processing unit 220 memory 230 Communications Department DB Database 300 Machine Learning Device 310 Controller 311, 312 Image processing unit 320 memory 330 Communications Department MVa Input video MVb thinned video (output video) MVc restored video Fa[1]~Fa[n], Fa[1]~Fc[n] frame images D[1]~D[n] Object feature data F11 Detector F12 Behavior recognition unit F13 Sampling section F21 Encoder F22 decoder
Claims
1. A moving image processing device having a controller, the controller comprising: Based on a plurality of frame images constituting an input video sequence, a skeleton of an object in each frame image is detected, and object feature data including the result of the skeleton detection is generated for each frame image; generating an output moving image by sampling some of the frame images from among the plurality of frame images; generating image set data including image data of the output moving image and the object feature data for each of the plurality of frame images; ,Video Image Processing Device.
2. The controller detects positions of a plurality of specific parts of the object in the frame image in the skeleton detection for each of the frame images. The moving image processing device according to claim 1 .
3. The object is a person or an animal, The controller recognizes the behavior of the object for each frame image, and includes a result of the recognition in the object feature data for each frame image. The moving image processing device according to claim 1 .
4. The controller generates the output video sequence by sampling the plurality of frame images at regular intervals.
4. The moving image processing device according to claim 1.
5. the plurality of frame images are first to nth frame images arranged in chronological order, Among the first to n-th frame images, the (k+1)-th to n-th frame images include an image of a first object as the object and an image of a second object different from the first object, while among the images of the first object and the second object, the first to k-th frame images include only the image of the first object, The controller performing the skeleton detection for the first object in each of the first to k-th frame images based on the first to k-th frame images, and including the result of the skeleton detection for the first object in each of the first to k-th frame images in the corresponding object feature data; based on the (k+1)-th to n-th frame images, performing the skeleton detection for the first object and the second object in each of the (k+1)-th to n-th frame images, and including the results of the skeleton detection for the first object and the second object in each of the (k+1)-th to n-th frame images in the corresponding object feature data; generating the output video sequence by sampling the first to k-th frame images at regular intervals from the first frame image and sampling the (k+1)-th to n-th frame images at the regular intervals from the (k+1) frame image; k represents an integer of 3 or more, and n represents an integer of (k+3) or more.
4. The moving image processing device according to claim 1.
6. 4. A video processing device according to claim 1, wherein the video processing device receives the image set data and generates a restored video by restoring the input video based on the received image set data. , data processing device.
7. An in-vehicle device installed in a vehicle, the in-vehicle device having the moving image processing device according to any one of claims 1 to 3; a camera installed in a vehicle and configured to capture and generate the input moving images; a data processing device that receives the image set data from the in-vehicle device and generates a restored video sequence by restoring the input video sequence based on the received image set data. ,image processing system.
8. an in-vehicle device having the moving image processing device according to any one of claims 1 to 3; a camera for capturing and generating the input moving images; ,vehicle.
9. A moving image processing method executed by a moving image processing device, Based on a plurality of frame images constituting an input video sequence, a skeleton of an object in each frame image is detected, and object feature data including the result of the skeleton detection is generated for each frame image; generating an output moving image by sampling some of the frame images from among the plurality of frame images; generating image set data including image data of the output moving image and the object feature data for each of the plurality of frame images; ,Video Image Processing Method.
10. A program that causes a computer device to execute the moving image processing method according to claim 9.
Citation Information
Patent Citations
Information processing program, method for processing information, and information processor
JP2023098484A