System, program, learned model, learning model generation method and generation device, and the like
The system addresses the challenge of efficiently collecting specific image scenes by using deep learning and machine learning to analyze data from a vehicle's camera, achieving accurate and reliable detection and recording of desired scenes.
Patent Information
- Application Number
- JP2025049193
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-12
AI Technical Summary
Existing technologies lack an efficient method for collecting images of specific scenes, such as those desired by users, from data obtained by a vehicle's camera, especially in abnormal situations.
A system that detects images representing specific scenes by analyzing data from a vehicle's camera at multiple time points, using deep learning and machine learning models to identify scenes with a high degree of coincidence with predefined specific images.
Enables the efficient detection and recording of specific image scenes, improving the accuracy and reliability of image collection, especially in abnormal or desired situations.
Smart Images

Figure 2025089453000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to, for example, a system, a program, a learned model, a method for generating a learning model, a generating apparatus, and the like.
Background Art
[0002] In the conventional art, event recording information is created by using, as a trigger, the occurrence of a predetermined event in a photographing apparatus such as a drive recorder. In such a case, there has been a consideration of efficiently collecting photographing data obtained by photographing in an abnormal situation (Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In addition to photographing data obtained by photographing in an abnormal situation, it is desirable to have a technique for efficiently collecting images of a specific scene such as a scene desired by a user.
[0005] In view of the above problems, one of the objects of the present invention is to provide a technique for detecting a specific image scene that can be photographed by a vehicle.
[0006] The object of the present invention is not limited to this, and the applicant also intends to obtain rights through divisional applications, amendments, etc. for configurations that aim to obtain the effects achieved by parts of the configurations disclosed in this specification, drawings, etc. For example, in this specification, problems obtained by reading parts described as "can" or "is possible" as "is a problem" are disclosed in this specification. The problems are described as independent ones, and the applicant also intends to obtain rights through divisional applications, amendments, etc. alone for the configurations for solving each problem. Even if a problem is implicitly understood from the description of the specification, the applicant intends to make part of the configuration described in this specification the scope of claims through amendment or divisional application. In addition, configurations for solving problems obtained by combining these independent problems are also disclosed, and the applicant intends to obtain rights.
Means for Solving the Problem
[0007] (1) The system according to this invention has a function of detecting an image representing the above scene from images taken by a vehicle at each of a plurality of time points based on data defining a scene of a specific image that can be taken by the vehicle.
[0008] In this way, a technology for detecting a scene of a specific image that can be taken by a vehicle can be provided.
[0009] (2) It has a function of taking pictures with a camera installed in the vehicle and a function of recording images taken at each of a plurality of time points by the above camera on a recording medium, and the above detecting function may detect an image representing a scene whose degree of coincidence with the scene of the above specific image is equal to or greater than a threshold value among the images at each of the above time points recorded on the above recording medium.
[0010] In this way, an image representing a scene whose degree of coincidence with the scene of a specific image is equal to or greater than a threshold value can be detected from the images taken by the camera installed in the vehicle.
[0011] (3) The above data may be an image at one time point.
[0012] If there is a specific image at a single point in time, the time until the scene is detected can be shortened.
[0013] (4) The above data may be images at two or more points in time.
[0014] If the specific image is an image at two or more points in time, the accuracy of the detected scene can be improved. Images at two or more points in time are preferably temporally continuous images like a video (images in which the movement of the main subject within the image is continuous).
[0015] (5) The above data may be generated by deep learning by inputting the same image as the above specific image or a plurality of images approximated to the above specific image.
[0016] In this way, using deep learning technology, it is possible to detect the scene of a specific image that can be captured by a vehicle.
[0017] (6) The above data is a trained first learning model that has performed machine learning using the above specific image as teacher data, and the above detection function may detect an image representing the above scene by inputting the images captured at the above respective points in time into the above trained model.
[0018] In this way, a specific frame can be detected relatively quickly.
[0019] (7) The above specific scene may include at least one of a scene that may lead to an accident or disaster, or a scene where an accident or disaster has occurred.
[0020] In this way, it is possible to detect from a video a scene that may lead to an accident or disaster, or a scene close to a scene where an accident or disaster has occurred.
[0021] Warning control means for controlling a warning device to give a warning in response to detecting an image representing the above scene may be further provided.
[0022] For example, when the above system is installed in a vehicle and a video is recorded while shooting, and a scene close to a specific image is found in the video, the driver of the vehicle can be alerted to the fact that the vehicle is in a situation close to the scene of the specific image.
[0023] (9) The warning control means may control the warning device to change the content of the warning according to the type or degree of danger of the above scene.
[0024] In this way, it is easy for the driver of the vehicle to grasp the degree of attention required.
[0025] (10) First recording control means for controlling a recording device to record the detected image of the scene on a recording medium may be further provided.
[0026] In this way, an image of a frame of a specific scene can be recorded, and for example, it can be made easier to refer to the image of that frame later.
[0027] (11) The above detection function may detect an image representing the above scene based on the above data and a physical quantity indicating the running state of the vehicle.
[0028] In this way, by taking into account the driving situation of the vehicle, a desired scene can be detected more accurately.
[0029] (12) The recording medium stores images taken by the vehicle at each of the above times, and reproduction control means for controlling a reproduction device to reproduce the images taken by the vehicle at each of the above times recorded on the recording medium, and notification control means for controlling a notification device to notify an image representing the detected above scene in association with the image reproduced by the reproduction device under the control of the reproduction control means may be further provided.
[0030] By doing so, it becomes easier to detect a scene approximated to a specific image when reproducing the image recorded on the recording medium.
[0031] (13) In the recording medium, taking the period from when a recording start command is given to when a recording stop command is given as one period, images captured by the vehicle at each of the above time points are stored, and there is provided a generation means for generating an image corresponding to a missing portion of the images occurring within the one period, and a second recording control means for controlling the recording device to record the image generated by the generation means in the missing portion.
[0032] By doing so, the video of the missing portion can be restored.
[0033] (14) The generation means inputs at least one second image before or after a first portion corresponding to the missing portion in the images captured at each of a plurality of time points in another vehicle, and performs machine learning with the first portion as an output to input at least one image before or after the missing data into a learned second learning model that generates an image corresponding to the missing portion.
[0034] By doing so, the video of the missing portion can be generated relatively accurately.
[0035] (15) A program for realizing the functions of the above system may be provided to a computer.
[0036] By doing so, by installing the program in the device, in that device, it is possible to detect a scene of a specific image that can be captured by the vehicle.
[0037] (16) The learned model according to the present invention uses a specific image that can be captured by the vehicle as teacher data, takes the input as images captured by the vehicle at each of a plurality of time points, and takes the output as the detection of an image representing the scene of the specific image from the input images.
[0038] In this way, it is possible to provide a learned model for detecting the scene of a specific image that can be captured by a vehicle.
[0039] (17) The method for generating a learning model according to this invention uses a specific image that can be captured by a vehicle as teacher data, sets the input as the images captured by the vehicle at each of a plurality of time points, and generates a learning model in which the output is an image representing the scene of the specific image from the input images.
[0040] In this way, it is possible to generate a learning model that can detect images of a relatively large number of specific scenes or detect images of specific scenes with high accuracy.
[0041] (18) The specific image may be an image captured by a plurality of imaging devices.
[0042] In this way, it is possible to generate a learning model that can detect images of a relatively large number of specific scenes or detect images of specific scenes with high accuracy.
[0043] (19) The imaging device is a drive recorder installed in the vehicle, and the specific image may represent the scene when a recording command is given to the drive recorder or the scene when an impact is applied to the vehicle on which the drive recorder is mounted.
[0044] In this way, it is possible to generate a learning model for detecting an image of a scene when a recording command is given to the drive recorder or when an impact is applied to the vehicle on which the drive recorder is mounted.
[0045] (20) It is preferable to detect an image representing a scene in which an object is moving at a speed or acceleration equal to or higher than a certain level.
[0046] In this way, it is possible to generate a learning model for detecting an image including an object that has moved at a speed or acceleration equal to or higher than a certain level.
[0047] It is preferable to generate a learning model for detecting an image representing the scene based on the above data and a physical quantity indicating the driving state of the vehicle.
[0048] In this way, by taking into account the driving situation of the vehicle, it is possible to generate a learning model that can detect a desired scene with higher accuracy.
[0049] Using, as teacher data, a first image that is an image of a predetermined period among the images taken at each of a plurality of time points, and a second image that is at least one of the images before or after the first image, performing learning with the first image as the input and the second image as the output, and using, as the input, a third image taken at each of a plurality of time points including a missing part in a predetermined period, and estimating a fourth image that is an image of the missing part from the input third image is preferable.
[0050] In this way, it is possible to generate a learning model for generating an image of the missing part.
[0051] The learning model generation device according to the present invention generates a learning model by the learning model generation method according to any one of (17) to (22).
[0052] In this way, it is possible to generate a learning model that can provide a learned model for detecting a scene of a specific image that can be captured by a vehicle.
[0053] The inventions described in (1) to (22) above can be arbitrarily combined. For example, it may be configured to add at least a part of the configuration of at least one of the inventions from (2) to (23) to all or a part of the configuration of the invention described in (1). In particular, it is preferable to make an invention in which at least a part of the configuration of at least one of the inventions from (2) to (23) is added to the invention described in (1). Also, any configuration can be extracted from the inventions described in (1) to (23), and the extracted configurations can be combined.
[0054] The applicant of the present application intends to obtain rights for inventions including these configurations. Also, even if there are descriptions such as "in the case of ~" or "when ~", it is not described as a configuration limited to that case or that time. These show examples of better configurations, and the applicant also intends to obtain rights for configurations that are not these cases or times. Also, the descriptions with an order are not limited to this order. The applicant also discloses configurations in which some parts are deleted or the order is changed, and intends to obtain rights for them.
Effect of the Invention
[0055] According to the present invention, it is possible to provide a technique for detecting a scene of a specific image that can be photographed by a vehicle.
[0056] Note that the effects of the invention of the present application are not limited to this, and the effects resulting from the disclosed parts of the configurations in this specification, drawings, etc. are also disclosed, and the applicant also intends to obtain rights for the configurations that exhibit such effects by means of divisional applications, amendments, etc. For example, the parts described as "can ~" or "is possible ~" in this specification are descriptions that explicitly indicate the resulting effects, and even if there is no description such as "can ~" or "is possible ~", there are parts that show the effects. Also, even if there is no such description, there are effects grasped by the said configuration.
Brief Description of the Drawings
[0057]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Embodiments for Carrying Out the Invention
[0058] Hereinafter, embodiments for carrying out the present invention will be described with reference to the drawings. Note that the embodiments shown below are one of the embodiments for providing the present invention, and the content of the present invention of the present application is not limitedly interpreted based on the following description.
[0059] [Background of the Invention of the Present Application] In current drive recorders, event recording is automatically (passively) executed according to the detection results of sensors such as acceleration sensors. However, the only way for the user to actively perform event recording is by using a switch. There is a situation where when the sensors do not react and the user cannot press the switch (for example, when the driver is not in the vehicle during driving or parking), video recording cannot be performed. Therefore, the inventor considered a solution method in which, by using a method such as deep learning, even when the sensors do not react, accidents and dangerous situations are learned from video data, and when the same or similar situations occur, event recording is automatically performed (such as in the first and second embodiments). In this way, this solution method complements event recording with machine learning and learns videos in situations where the sensors do not react. Also, the inventor considered that instead of event recording for shooting situations such as accidents, the scenery during driving or the points that the user wants to capture can be learned in advance as teacher images, and similar situations can be automatically determined and recorded by the drive recorder. The inventor thought that if it could be used as a shooting device instead of videos of accidents and the like, the uses of the drive recorder would increase (such as in the third embodiment).
[0060] [Overview of Each Embodiment] Each of the embodiments described below relates to a technique for detecting an image representing the above-described scene from images captured by a vehicle at each of a plurality of time points based on data defining a scene of a specific image that can be captured by the vehicle. In each embodiment, the specific image is an image (also referred to as a desired image) that defines a desired scene. "Detecting an image representing the above-described scene" includes generating a trigger for executing a predetermined process. The predetermined process is event recording in the first and second embodiments, and in the third embodiment, displaying information regarding the image representing the scene, and when the displayed information is selected, displaying an event-recorded video including the image representing the scene based on video data representing the corresponding video, fast-forwarding to the point in time when the image representing the scene is displayed in response to this selection during video playback, and jumping (heading to) to this point in time during image playback. Not limited to these, "detecting an image representing the above-described scene" may include a process such as associating identification information indicating that the image has been detected. In each embodiment, a method is described for finding a frame representing a scene having a degree of coincidence with the scene of the desired image equal to or greater than a threshold value among a plurality of frames constituting a video represented by video data. A video is configured by arranging images (here, still images, also referred to as frames in the embodiment) at each of a plurality of time points on a time axis. When a video is played back based on video data, images are sequentially played back along the time axis. The specific image may be an image at one point in time (i.e., a still image), or any of two or three or more images at different points in time, or two or three or more images captured continuously in time representing frames constituting a video. The desired image may be any image such as an image of an accident such as a car accident (not limited to car accidents), a scene that may lead to a disaster such as a natural disaster (not limited to natural disasters), a scene where an accident or disaster has occurred, or an image of a desired landscape. Also, as described later, it may be an image of a scene in which a recording command is given to a drive recorder by an operation button or the like, or an image of a scene when an impact is applied to a vehicle equipped with a drive recorder.
[0061] [First Embodiment] Figures 1 to 11 show a first embodiment. The first embodiment uses deep learning with a plurality of images that are the same as or approximate to a desired image, or a plurality of images that are the same as and approximate to a desired image, to generate a target image, and uses the generated target image as the desired image. Also, the generation of a learned model will be described.
[0062] Figure 1 is a perspective view of the drive recorder 1 (an example of a system) seen from the obliquely rear side. In the following description, the drive recorder 1 is a device installed on the vehicle from the rear in this embodiment. The vehicle is, for example, a private car or a commercial car (passenger car). However, the vehicle may be a special vehicle such as a bus, a truck, a forklift, or a vehicle in public transportation such as a train, a monorail, or a linear motor car.
[0063] An SD card insertion port 10 is formed on one side surface of the housing of the drive recorder 1. A display 11 is formed on almost the entire back surface of the housing of the drive recorder 1, and a plurality of operation buttons 12 are provided on both the left and right sides with the display 11 in between. A joint rail 13 is provided on the upper surface of the drive recorder 1. Although not shown in FIG. 1, a camera lens is provided on the front surface of the drive recorder. A DC jack of the drive recorder 1 is formed on the side surface not shown in FIG. 1, and a speaker and an HD (high definition) output terminal are formed on the bottom surface of the drive recorder 1.
[0064] A camera including a lens captures an image in front of, for example, a vehicle. Preferably, it may also capture images inside and outside the vehicle cabin (such as the lateral and rear directions of the vehicle). A DC jack is a jack for connecting to a DC power supply via a power cable. The SD card slot 10 is a slot for inserting an SD card [an example of a memory card, which may also be other memory cards such as xD picture cards, or memories such as a USB (Universal Serial Bus) memory other than a memory card]. A speaker outputs sound and voice. An HD output terminal is a terminal for connecting to other information devices via a cable. The joint rail 13 is for attaching a joint for mounting the drive recorder 1 to a vehicle. The display 11 displays various images such as the image captured by the camera of the drive recorder 1. The operation button 12 is for a user to input various commands to the drive recorder 1 by operating it.
[0065] Figure 2 shows a view looking forward from inside the vehicle cabin of the vehicle equipped with the drive recorder 1.
[0066] A rearview mirror 4 is provided near the center of the upper part of the front windshield 3 of the vehicle. The drive recorder 1 is fixed to the vehicle by a joint at a position on the passenger seat side (the left side when looking forward in FIG. 2) adjacent to the rearview mirror 4.
[0067] The DC jack of the drive recorder 1 is connected to the cigarette socket 5 via the power cable 6. When the accessory power of the vehicle is turned on, power is supplied from the cigarette socket 5 to the drive recorder 1. Note that the drive recorder 1 is not limited to the forms shown in FIGS. 1 and 2, and for example, it may be equipped with an omnidirectional camera or a hemispherical camera.
[0068] Figure 3 is a block diagram showing the electrical configuration of the drive recorder 1.
[0069] The drive recorder 1 includes a controller 20. This controller 20 includes a CPU (Central Processing Unit) 20a that oversees the overall operation of the drive recorder 1, a ROM (Read Only Memory) 20b that stores the operation program of the drive recorder 1 and the like, a RAM (Random Access Memory) 20c that temporarily stores data and the like, and a timer 20d. The controller 20 has functions such as a GPS information processing program, a video processing program, and a communication processing program.
[0070] Also, the drive recorder 1 includes a GPS (Global Positioning System) receiver 16, a camera 17, an SD card reader / writer 14, and an acceleration sensor 18. The GPS receiver 16 detects the position of the drive recorder 1, and thus the position of the vehicle on which the drive recorder 1 is mounted can be known. The camera 17 photographs the front etc. of the vehicle on which the drive recorder 1 is mounted as described above. The SD card reader / writer 14 reads the data recorded on the inserted SD card 23 and writes data to the SD card 23 when the SD card 23 is inserted into the SD card insertion slot 10 as described above. The acceleration sensor 18 detects the acceleration in the vertical, horizontal, and front-back directions applied to the drive recorder 1, and thus the acceleration in the vertical, horizontal, and front-back directions of the vehicle on which the drive recorder 1 is mounted can be known.
[0071] The GPS data representing the position of the drive recorder 1 output from the GPS receiver 16, the video data representing the video photographed by the camera 17, and the signal representing the acceleration detected by the acceleration sensor 18 are respectively input to the controller 20. The controller 20 can record the GPS data output from the GPS receiver 16 onto the SD card 23 inserted into the SD card insertion slot 10 by executing the GPS information processing program. Also, the controller 20 can record the video data obtained from the camera 17 on the SD card 23 in association with the time by executing the video processing program.
[0072] Furthermore, the drive recorder 1 includes a speaker 15 for outputting the above-described voice and the like, a display 11 for displaying video and the like, operation buttons 12, a communication circuit 19, and an LTE module 21.
[0073] By outputting voice data and the like from the controller 20 to the speaker 15, the speaker 15 outputs voice and the like, and by outputting video data and the like from the controller 20 to the display 11, the display 11 displays video and the like. Various commands from the operation buttons 12 are input to the controller 20.
[0074] The communication circuit 19 functions as a communication means for performing wireless communication with external devices, such as a server, a personal computer, a smartphone, a tablet terminal, etc. The controller 20 has a function of transmitting data such as video to an external device via the communication circuit 19 by executing a communication processing program. As the communication circuit 19, for example, one compliant with a short-range wireless communication standard such as the WiFi standard or Bluetooth (registered trademark) may be used. The short-range wireless communication standard can be applied, for example, to the communication between a work vehicle (such as a forklift) operating within a facility and an external device. The mobile communication system standard can be applied, for example, to the communication between a vehicle moving within a wider area and an external device.
[0075] The LTE (Long Term Evolution) module 21 is a communication circuit compliant with the communication standard of a mobile phone. Instead of the LTE module 21, a communication circuit compliant with the standard of another mobile communication system such as 4G may be used. Video data and the like captured by the camera 17 can be transmitted to an external server or the like by the LTE module 21.
[0076] The operation program of the drive recorder 1 may be pre-recorded in the SD card 23, or may be received by the LTE module 21 via the Internet or the like and recorded in the SD card 23, or the SD card 23 may be inserted into a personal computer or the like, connected to the Internet via the personal computer, downloaded from the personal computer, and recorded in the SD card 23.
[0077] FIG. 4 is a block diagram showing the electrical configuration of 30 servers that communicate with the drive recorder 1.
[0078] The server 30 includes a control device 31 that oversees the overall operation.
[0079] Connected to the control device 31 of the server 30 are a communication device that communicates with the drive recorder 1 and other client computers, etc., a memory 33 that stores data, etc., and a hard disk drive 34. The hard disk drive 34 accesses a hard disk 35 that stores video data, etc., reads the video data recorded on the hard disk 35, and records video data, etc., on the hard disk 35.
[0080] FIG. 5 is an example of a learning model 40.
[0081] The learning model 40 shown in FIG. 5 generates a target image by inputting a plurality of images by deep learning. The learning model 40 includes an input layer 41, an intermediate layer (hidden layer) 42, and an output layer 43.
[0082] Each of the pixels P1 to PN that make up any one image (including not only the desired image and images approximated to the desired image but also other images) is input from the input layer 41 of the learning model 40, and learning is performed by deep learning through the intermediate layer 42 and the output layer 43, and this is done only for a huge number of arbitrary images. Then, when an arbitrary image containing many desired images and images approximated to the desired image is input from the input layer 41 of the learning model 40, it is possible to find neurons 43a that respond to the desired image and images approximated to the desired image. When an unknown image is input from the input layer 41 of the learning model 40 and the neurons 43a react, it can be determined that the input unknown image is a desired image or an image approximated to the desired image. Thus, image recognition can be performed by the learning model 40.
[0083] Furthermore, by tracing back from the reacted neurons 43a to the input layer 41 side, it is possible to generate a target image that most resembles the desired image. The target image generated in this way becomes the desired image generated by deep learning (unsupervised learning).
[0084] For example, when images that are likely to lead to an accident, images that are likely to lead to a disaster, images that have led to an accident, images that have led to a disaster, and a huge number of other images that have nothing to do with accidents or disasters are input to the learning model 40 and learned by deep learning, it becomes known which neurons (let's call them the first neurons) react when an image that is likely to lead to an accident is input to the learning model 40, which neurons (let's call them the second neurons) react when an image that is likely to lead to a disaster is input to the learning model, which neurons (let's call them the third neurons) react when an image that has led to an accident is input to the learning model, and which neurons (let's call them the fourth neurons) react when an image that has led to a disaster is input to the learning model, and so on.
[0085] By tracing back from the reacted first neuron to the input layer 41 side, a target image that seems to be an image likely to lead to an accident can be generated. By tracing back from the reacted second neuron to the input layer 421, a target image that seems to be an image likely to lead to a disaster can be generated. By tracing back from the reacted third neuron to the input layer 41 side, a target image that seems to be an image that has led to an accident can be generated. By tracing back from the reacted fourth neuron to the input layer 41 side, a target image that seems to be an image that has led to a disaster can be generated. The generated target image can be used as a desired image.
[0086] Figure 6 is an example of the target image 50 generated as described above and is an example of an image likely to lead to an accident.
[0087] The target image 50 is an image generated based on deep learning and is different from an actually photographed image. However, it may accidentally be an image almost identical to an actually photographed image.
[0088] As an example of this embodiment, among the frames constituting the video, frames of scenes whose degree of coincidence with the scene of this target image 50 is equal to or higher than a threshold value can be found.
[0089] Figure 7 is a flowchart showing the target image generation processing procedure.
[0090] The target image generation processing procedure shown in FIG. 7 may be performed at any timing as long as a large number of images are obtained. Also, if image data representing a large number of images is recorded in the SD card 23 or the database 22, it may be performed in the drive recorder 1, or if image data representing a large number of images is recorded in the hard disk 35 of the server 30, it may be performed in the server 30. When collecting images from drive recorders (a plurality of drive recorders) mounted on a plurality of vehicles, it is desirable that the target image generation processing procedure be executed by the server 30. Further, if access to the Internet is possible, even if a large number of images are not recorded in the SD card 23, the database 22, or the hard disk 35, the target image generation processing may be performed either in the drive recorder 1 or the server 30 by reading a large number of images from the Internet. Furthermore, the target image generation processing may be performed in a client computer (personal computer).
[0091] A large number of arbitrary images are input to the learning model 40 (step 61), and deep learning is performed in the learning model 40 (step 62). As described above, the target image is obtained (step 63).
[0092] Needless to say, the program of the learning model 40 is installed in the controller 20 of the drive recorder 1 or the control device 31 of the server 30.
[0093] FIG. 8 shows a method for generating a plurality of target images.
[0094] The first learning model 40A, the second learning model 40B, and the third learning model 40C include an input layer, an intermediate layer, and an output layer, similar to the learning model shown in FIG. 5.
[0095] Let the image input to the first learning model 40A be the first image, the image input to the second learning model 40B be the second image, and the image input to the third learning model 40C be the third image. As described with reference to FIG. 5, when a huge number of images are input to the first learning model 40A and learned by deep learning, neurons that respond to specific images (for example, images with a high possibility of leading to an accident) are found, and by tracing those neurons to the input layer, an image representing the features of the specific image can be obtained. For example, a first target image representing features with a high possibility of leading to an accident is obtained. Similarly, when a huge number of images are input to the second learning model 40B and learned by deep learning, neurons that respond to specific images (for example, images with a high possibility of leading to a disaster) are found, and by tracing those neurons to the input layer, an image representing the features of the specific image can be obtained. For example, a second target image representing features with a high possibility of leading to a disaster is obtained. When a huge number of images are input to the third learning model 40C and learned by deep learning, neurons that respond to specific images (for example, images of an accident that has occurred) are found, and by tracing those neurons to the input layer, an image representing the features of the specific image can be obtained. For example, a third target image representing the features of an image of an accident that has occurred is obtained. In this way, multiple types of target images are obtained.
[0096] In FIG. 8, multiple types of target images are obtained using multiple learning models 40A - 40C. However, even in one learning model 40 as shown in FIG. 5, by finding neurons that respond to the first target image, neurons that respond to the second target image, neurons that respond to the third target image, etc., and tracing them to the input layer 41, the first target image, the second target image, the third target image, etc. can be found.
[0097] In the example shown in FIG. 8, the first image, the second image, and the third image are preferably images that are taken continuously in time series. As a result, the obtained first target image, second target image, and third target image will also represent those that are continuous in order in time series. For example, if the first target image, second target image, and third target image represent the characteristics of an image highly likely to lead to an accident, when frames close to the first target image, second target image, and third target image appear in sequence in the video, it can be determined that the possibility of an accident is higher in such a situation, and the accuracy is higher than when determining that there is a high possibility of an accident based on only one target image.
[0098] FIGS. 9(A), (B), and (C) are examples of the first target image, second target image, and third target image generated as described above.
[0099] As the first target image 51, second target image 52, and third target image 53, it can be seen that the child appearing as a subject approaches the car, and the possibility of an accident is increasing. Considering these first target image 51, second target image 52, and third target image 53 as desired images, if a scene similar to these first target image 51, second target image 52, and third target image 53 is found from the video, it can be determined that there is a high probability of an accident occurring at that scene.
[0100] In the above-described embodiment, the three first target image 51, second target image 52, and third target image 53 are used as desired images, but two or four or more target images may be generated and used as desired images.
[0101] FIG. 10 shows another method of generating a target image, and is a block diagram showing the electrical configuration of an adversarial generation network.
[0102] Feature data representing the features of the desired image is input to the generator 71. In the generator 71, pseudo-image data representing a pseudo-image of the desired image is generated from the input feature data. The generated pseudo-image data is input from the generator 71 to the discriminator 72. Image data representing the desired image is also input to the discriminator 72. In the discriminator 72, it is discriminated whether the desired image represented by the image data representing the desired image and the pseudo-image represented by the pseudo-image data are the same or similar images. Data representing the discrimination result is input to each of the generator 71 and the weighting changer 73. Based on the data representing the discrimination result, the generator 71 is adjusted, and the weighting of the feature amounts of the discriminator 72 is changed by the weighting changer 73.
[0103] When the generation of pseudo-image data in the generator 71, the discrimination in the discriminator 72, the adjustment based on the discrimination result in the generator 71, and the change of the weighting of the discriminator 72 in the weighting changer 73 are repeated, it is determined in the discriminator 72 that the pseudo-image represented by the pseudo-image data generated by the generator 71 and the desired image are identical or approximate. The pseudo-image thus determined is taken as the target image. A plurality of images of the same type may be used as the desired image, and the pseudo-image obtained by repeating the above-described processing may be used as the target image. The target image thus obtained becomes the desired image found from the frames constituting the video.
[0104] In the above embodiments, a target image is generated using an arbitrary image and the generated target image is used as a desired image, or a pseudo-image is generated using a desired image and the generated pseudo-image is used as a desired image. However, when an image with a high possibility of leading to an accident or disaster, an image that has led to an accident or disaster, an image of a scene that the user wants to search for, etc. are already known, the target image, pseudo-image, etc. may be used as the desired image without generating them using deep learning or the like. For example, in the case of a video of an event record (for example, a video recorded when an impact or the like is given to a vehicle when a drive recorder mounted on the vehicle records a video), one or a plurality of frames constituting the video of the event record may be used as the desired image.
[0105] FIG. 11 shows a method of generating a learned model by training a learning model.
[0106] This is another example of generating a target image and shows an example of a learning model.
[0107] The learning model 80 includes an input layer 81, an intermediate layer (hidden layer) 82, and an output layer 83.
[0108] Using a desired image as teacher data, data representing pixels P1 to PN that make up the desired image is input to the input layer 81 in the order in which the image is composed. The features, weightings, etc. of the input layer 81, intermediate layer 82, and output layer 83 are adjusted as necessary so that the same data as pixels P1 to PN input to the input layer 81 via the intermediate layer 82 is output from the output layer 83 in the same order as the order in which the desired image is composed. That is, machine learning is performed so that the output of the learning model 80 becomes almost the same as the input, and the features, weightings, etc. of the learning model 80 are adjusted as necessary.
[0109] The number of data representing pixels input to the input layer 81 or the number of data representing pixels output from the output layer 83 may be reduced. The feature amounts, weightings, etc. of the learning model 80 may be adjusted as necessary so that the data representing pixels at the same position in the image are substantially the same between the input and the output.
[0110] When data representing an arbitrary image (for example, each frame constituting a video) is input to the learning model that has undergone such learning, the data of each pixel output from the output layer 83 becomes closer to the data of pixels P1 to PN constituting the desired image as the output image more closely matches the desired image.
[0111] According to the first embodiment, a learned model can be obtained that can generate a desired image and can be used to detect a frame of a scene approximated to the scene of the desired image from among a plurality of frames constituting a video.
[0112] [Second Embodiment] FIGS. 12 to 16 show a second embodiment. The second embodiment searches for a frame having a high degree of coincidence with a desired image from among a plurality of frames constituting video data by using the learned learning model 80 generated with reference to FIG. 11. In the second embodiment, video data photographed by a photographing device is recorded, the recorded video data is read, and a frame having a high degree of coincidence with a desired image is searched for from among a plurality of frames constituting the read video data. Hereinafter, in order to detect a desired image, the acceleration detected by the acceleration sensor 18 is used, and a process of determining whether there is an object moving at a speed or acceleration equal to or higher than a certain level in the video is used in combination. However, at least as long as an image photographed by the drive recorder 1 is applied to the learned model generated by the method described in the first embodiment, any method for detecting a desired image may be used.
[0113] FIG. 12 is a flowchart showing a processing procedure of event recording that is constantly performed during recording by the drive recorder 1 shown in FIGS. 1 and 2.
[0114] A program for implementing the processing procedure shown in FIG. 12 may be installed in the controller 20 of the drive recorder, or may be stored in an SD card (an example of a recording medium storing the program) 23, or may be received by the communication circuit 19 from being transmitted via a network and installed in the drive recorder 1.
[0115] When the drive recorder 1 is mounted on a vehicle, when the engine of the vehicle is started, shooting by the camera 17 of the drive recorder 1 is started, and recording (constant recording) of the video data obtained by the shooting to the SD card 23 is started by the SD card reader / writer 14. The constant recording starts when the engine of the vehicle is started and ends when the engine of the vehicle is turned off. Also, even when the engine of the vehicle is not running, when a recording start command is given from the operation button 12 to the controller 20, the constant recording starts, and when a recording end command (an example of a recording stop command) is given from the operation button 12 to the controller 20, the constant recording ends. The period from the start to the end of the constant recording becomes one video of the constant recording. Also, in this embodiment, the video data obtained by shooting by the camera 17 is also given to and recorded in the RAM 20c. The video data recorded in the RAM 20c is periodically recorded for a certain period of time, and is repeatedly overwritten when that certain period of time has elapsed.
[0116] In this embodiment, when an event is detected by the acceleration sensor 18 or when an event is detected by a recording command being given by the operation button 12 (YES in step 91), event recording is performed in the event recording area of the SD card 23 (step 98).
[0117] When the acceleration sensor 18 or an event detection by a recording command is not detected (NO in step 91), the video data recorded in the RAM 20c (an example of a recording medium) by the controller 20 (an example of a first recording control means) is read, and the controller 20 determines whether there is an object moving at a speed or acceleration equal to or higher than a certain level in the video represented by the read video data (step 92). This object may or may not be defined in advance. It may be the relative speed or acceleration between the vehicle equipped with the drive recorder 1 and the object, or the absolute speed or acceleration of the object.
[0118] When there is an object moving at a speed or acceleration equal to or higher than a certain level in the video represented by the read video data (YES in step 92), the video data read from the RAM 20c is input to the learned learning model 80 one frame at a time (step 93). Learning processing is performed in the learning model 80, and image data representing a feature image is obtained from the output layer 83 of the learning model 80 one frame at a time (step 94). Although the learning model 80 is described as if it were hardware, in reality, it is implemented by software in the controller 20.
[0119] The controller 20 calculates the degree of coincidence between the obtained feature image and the desired image (step 95).
[0120] FIG. 13 shows a state in which video data is input to the learned learning model 80 one frame at a time.
[0121] Pixels P11 to P1N constituting one frame are input to the input layer 81 according to the arrangement of one frame, and data of pixels P21 to P2N is output as image data representing a feature image through the intermediate layer 82 and the output layer 83.
[0122] Data of these pixels P21 to P2N representing the feature image is compared with data of each pixel constituting the desired image, and a degree-of-match calculation process between the feature image and the desired image is performed in the controller 20.
[0123] FIG. 14 is a flowchart showing a processing procedure for calculating the degree of match.
[0124] First, a degree of match between the desired image and the feature image is calculated as a first degree of match (step 131). Subsequently, when the desired image is obtained from the drive recorder 1 mounted on a vehicle (an example of a second vehicle), a degree of match between vehicle information (an example of running information) and a frame input to the learning model 80 when the feature image was obtained is calculated as a second degree of match with vehicle information obtained from the drive recorder 1 mounted on a vehicle (an example of a first vehicle) (step 132). The running information includes, for example, physical quantities indicating the state of the vehicle such as the speed of the vehicle, the attitude of the vehicle body, and the acceleration of the vehicle. Such physical quantities are physical quantities that change as the vehicle runs. In addition to this, the running information may include information that changes as the vehicle runs, such as the location (position information) of the vehicle. Further, a comprehensive degree of match is calculated from the first degree of match and the second degree of match (step 133).
[0125] In the processing procedure shown in FIG. 14, since the vehicle information when the desired image was obtained is used, in the case where the desired image is a virtual target image generated using a learned learning model, the comprehensive degree of match will be the first degree of match without calculating the second degree of match. However, even if the desired image is a virtual target image, if the vehicle information about the image used to obtain the target image is known, the vehicle information may be used to calculate average vehicle information and calculate the second degree of match.
[0126] Returning to FIG. 12, when the calculated degree of coincidence (overall degree of coincidence) is equal to or greater than the threshold value (YES in step 96), among the frames constituting the video represented by the video data recorded in the RAM 20c, a frame representing a scene with a degree of coincidence with the desired image scene equal to or greater than the threshold value is detected, and an event recording command is generated from the controller 20 (an example of warning control means), and a warning is issued (step 97). The warning may be a voice warning output from the speaker 15 (an example of a warning device) of the drive recorder 1, a warning by characters, images, etc. displayed on the display 11 (an example of a warning device), a warning by the light emission, blinking, etc. of an LED (light emitting diode). Further, a warning may be given using a speaker, a display screen, etc. provided in the vehicle on which the drive recorder 1 is mounted. By the warning, the driver of the vehicle can know that the driving situation is a situation close to the scene of the desired image (is in the driving situation). For example, if the desired image is an image with a high possibility of leading to an accident, more careful driving can be taken. Also, when the event recording command is generated, event recording is performed in the event recording area of the SD card 23 (step 98).
[0127] In the above-described embodiment, when the object is moving at a speed or acceleration equal to or higher than a certain level (YES in step 92), the captured data is input to the learned learning model. However, the captured data may be input to the learned learning model regardless of whether the object is moving at a speed or acceleration equal to or higher than a certain level, or the captured data may be input to the learned learning model without determining whether the object is moving at a speed or acceleration equal to or higher than a certain level.
[0128] FIG. 15 is an example of one of the plurality of frames constituting the video represented by the video data read from the RAM 20c. FIG. 16 shows a plurality of frames constituting the video represented by the video data read from the RAM 20c.
[0129] As shown in FIG. 16, the video represented by the video data is composed of a plurality of frames FR1 to FRe, and it is assumed that the degree of coincidence (total degree of coincidence) obtained when the frame FRn shown in FIGS. 15 and 16 is input to the learning model 80 is equal to or greater than the threshold value. Then, data representing the frames representing the video shot from several seconds to several tens of seconds (or other time periods) before the frame FRn and the frames representing the video shot from several seconds to several tens of seconds (or other time periods) after the frame FRn are recorded in the event recording area of the SD card 23 by the SD card reader / writer 14.
[0130] The warning in the above-described embodiment may be changed in content according to the type of scene or the degree of danger. For example, in a scene with a high degree of danger, both voice and display are used to warn so that the driver can recognize it, or the volume is increased. In a scene with a low degree of danger, for example, when a scene similar to a favorite place is found, the driver is notified with a relatively quiet voice.
[0131] FIG. 17 shows a modified example of the learning model 80 shown in FIG. 13.
[0132] The learning model 105 shown in FIG. 17 includes an input layer 106, an intermediate layer 107, and an output layer 108, and the output layer 108 is a logistic regression layer.
[0133] A plurality of frames constituting a video are input one by one into a learning model 105. For example, data representing pixels P31 to P3N constituting one frame is input into an input layer 106 of the learning model 105 according to a pixel array. In such a case, when the frame has a high degree of coincidence with a desired image, identification data is output from an output layer 83 which is a logistic regression layer. When the identification data is output, it is determined that the degree of coincidence between the desired image and the frame input to the learning model 105 is equal to or higher than a threshold value, and an event recording command is generated from a controller 20. The learning model 105 is also a learned learning model generated by inputting image data representing a number of desired images from an input layer 106 and training so that identification data is output from an output layer 108 when the image data representing the desired image is input from the input layer 106.
[0134] Also, an event recording command may be generated when the degree of coincidence (the above-described second degree of coincidence) between the vehicle information when the frame input to the learning model 105 is obtained and the vehicle information corresponding to the desired image is equal to or higher than a threshold value when the identification data is output.
[0135] Event recording can be performed based on the captured video even without an event recording command from an acceleration sensor 18 or an operation button 12. Also, when calculating the overall degree of coincidence using vehicle information, not only the captured video but also vehicle information is used, so more appropriate event recording can be performed according to the situation of the vehicle.
[0136] In the above-described embodiment, when a frame approximated to a desired image is detected from the video, event recording is performed and the video data is recorded on an SD card 23. However, identification information for identifying a frame approximated to the desired image without recording the video data may be recorded on the SD card 23 or other recording media.
[0137] In the embodiments described above, the presence or absence of event recording was determined using the overall matching degree. However, the presence or absence of event recording may be determined based on whether the first matching degree is greater than or equal to a threshold value without using the second matching degree. However, by combining physical quantities indicating the driving state of the vehicle, it is possible to expect the effect of more accurately detecting frames of a desired scene.
[0138] [Third Embodiment] This embodiment proposes an automatic playback support function in viewer software for recorded data. In the video data of a drive recorder for management purposes, it is often the case that there are no problems basically for a long time. For events, the occurrence points are displayed in the viewer software. However, it has sometimes been painful for administrators or the like to discover and confirm the attention points that are not events. Therefore, the inventor has learned a desired video situation by deep learning or the like. When there is a video in the recorded data that is the same as or similar to the desired situation, it is identified and detected as an attention point separately from the event, and is displayed in viewer software or the like, or when fast-forwarding playback, it is returned to the normal speed at the event occurrence point and the attention point. The inventor has considered a solution method to make it easier for administrators or the like to confirm.
[0139] In addition, this embodiment proposes a recovery function for data such as video recorded on a recording medium. For example, when adopting the recording format described later, instead of the FAT format used in Windows (registered trademark), etc., data is written sequentially to the storage medium. Therefore, even if the video is in an abnormal state (even if there are missing parts), the data itself is recorded. However, when a special abnormality occurs, there is a problem that a predetermined sector of the video data is erased with 0 or other information is written, resulting in the inability to play the video with viewer software. In this case, the contents of the storage medium may be checked sector by sector and restored manually. However, the inventor considered a solution method in which the state of the storage medium is learned by a method such as deep learning, not only to restore the erased data from the data before and after the missing part, but also to restore it to a normal state as the above recording format so that it can be confirmed with viewer software or the like. Hereinafter, the details of this embodiment will be described.
[0140] Figures 18 to 29 show the third embodiment and relate to the playback of a video recorded by a recording device such as the drive recorder 1. In the second embodiment, when recording video data in a recording device such as the drive recorder 1, event recording is performed when a frame with a high degree of coincidence with a desired image is found. In the third embodiment, when playing back the continuously recorded video data recorded on the SD card 23 (or other recording medium), a frame with a high degree of coincidence with a desired image is found and event recording is performed. In the third embodiment, the continuously recorded video data is played back using a personal computer. However, when a recording device such as the drive recorder 1 has a playback function, playback can be performed on such a recording device and the next process can also be performed.
[0141] Figure 18 is a block diagram showing the electrical configuration of a personal computer 110 that plays back a video.
[0142] The overall operation of the personal computer 110 is coordinated by a control device 115.
[0143] The personal computer 110 is provided with a display device 111, and images and the like are displayed on the display screen of the display device 111 under the control of the display control device 112. A communication device 113 for communicating with the Internet or the like and a memory 114 are connected to the control device 115. The personal computer 110 is provided with an input device 116 such as a keyboard and a mouse, and commands given from the input device 116 are input to the control device 115.
[0144] Furthermore, the control device 115 is also connected with a hard disk drive 117 that reads data recorded on the hard disk 118 and records data on the hard disk 118, and an SD card reader / writer 119 that reads data recorded on the SD card and records data on the SD card 23.
[0145] FIG. 19 shows the recording format of the SD card 23.
[0146] In the recording area of the SD card 23, a file system area 121, a constant recording area 122, and an event recording area 125 are formed.
[0147] In the file system area 121, dedicated software for playing back data recorded in the constant recording area 122 and the event recording area 125 is recorded.
[0148] In the constant recording area 122, a management area 123 and a recording area 124 are formed. In the management area 123, various setting information is recorded by the dedicated software recorded in the file system area 121. In the recording area 124, image data representing frames constituting a video is recorded in frame number order. In the recording area 124, a header recording area 131, a frame image data recording area 132, and a footer recording area 133 are formed for each frame. Additional information such as a frame number, an address position, and a shooting time is recorded in the header, and additional information other than the additional information recorded in the header is recorded in the footer.
[0149] Frames 1 to E recorded in the recording area 124 of the always recording area 122 represent one video period of always recording.
[0150] Similar to the always recording area 122, the event recording area 125 also has a management area 126 and a recording area 127 formed. Various setting information is recorded in the management area 126 by the dedicated software recorded in the file system area 121. In the recording area 127, image data representing the frames constituting the video is recorded in frame number order. In the recording area 127, a header recording area 131, a frame image data recording area 132, and a footer recording area 133 are also formed for each frame. Additional information such as frame number, address position, and shooting time is recorded in the header, and additional information other than the additional information recorded in the header is recorded in the footer.
[0151] Frames 1 to E recorded in the recording area 127 of the event recording area 125 represent one video of event recording.
[0152] Figure 20 is an example of the playback window 150.
[0153] When the SD card 23 is loaded into the personal computer 110 and the icon of the SD card 23 appearing on the display screen of the display device 111 is clicked, the icon of the dedicated software appears. When the icon of the dedicated software is clicked, the dedicated software is launched, and the playback window 150 is displayed on the display screen of the display device 111.
[0154] In the playback window 150, there are formed a video display area 151 for displaying the selected video, an event recording list display area 152, a map image display area 153, a vehicle information etc. display area 160, an always recording list display area 161, and an acceleration display area 163.
[0155] The video display area 151 is an area for displaying videos (constant recording videos, event recording videos) recorded by a drive recorder 1 or the like. The event recording list display area 152 is an area for displaying an event recording list in the form of recording time (it may also be in a form other than recording time). The event recording list displayed in the event recording list display area 152 shows all the event recordings stored in the SD card 23. When a constant recording video is selected, the event recording videos recorded during the shooting of the selected constant recording video may be displayed in the event recording list display area 152.
[0156] The map image display area 153 is an area for displaying an image of a map in the vicinity of the shooting location of the video displayed in the video display area 151. By transmitting data representing the position of the vehicle from the personal computer 110 to a map server existing on the Internet and transmitting data representing a map in the vicinity of the vehicle's position from the map server to the personal computer 110, an image of the map is displayed in the map image display area 153. The vehicle information etc. display area 160 is an area for displaying information such as the speed of the vehicle at the time of video recording. The vehicle information etc. display area 160 includes a vehicle state display area 158 for displaying the state of the vehicle to which the drive recorder 1 that recorded the video displayed in the video display area 151 is attached, an operation button 157 for giving operation commands such as play and stop of the video, an interior enhancement button 154 for giving an instruction to perform a process of brightly displaying the video inside the vehicle, a normal display button 155 for giving an instruction to stop the process of brightly displaying the video inside the vehicle and perform normal display, a time display area 156 for displaying the recording time of the video displayed in the video display area 151, an acceleration graph display area 159 for graphically displaying the acceleration of the vehicle, etc.
[0157] The always-recorded list display area 161 is an area for displaying, in a list, the always-recorded videos recorded on the SD card 23. An image of a vehicle is displayed in the acceleration display area 162. In front of the vehicle image, the character “+X” is displayed together with an arrow pointing forward, the character “+Y” is displayed together with an arrow pointing in the horizontal direction beside the vehicle image, and the character “+Z” is displayed together with an arrow pointing upward above the vehicle image. The acceleration with respect to the front or rear of the vehicle is displayed in front of the vehicle image, the acceleration with respect to the lateral direction of the vehicle is displayed in the lateral direction of the vehicle image, and the acceleration with respect to the vertical direction of the vehicle is displayed in the upward direction of the vehicle image. Since the acceleration of the vehicle is displayed together with the appearance of the vehicle, it becomes relatively easy to understand in which direction the acceleration is applied to the vehicle.
[0158] FIG. 21 is a flowchart showing a reproduction processing procedure, and FIGS. 22 and 23 show an example of the reproduction window 150.
[0159] When a desired event-recorded video or always-recorded video is selected from the event-recorded list or always-recorded list displayed in the event-recorded list display area 152 or always-recorded list display area 161 of the reproduction window 150, the video data representing the selected video is read from the always-recorded area 122 of the SD card 23 by the SD card reader / writer 119 (step 141). In this embodiment, it is assumed that an always-recorded video is selected.
[0160] When the always-recorded video is selected, a representative image of the selected always-recorded video is displayed in the video display area 151 of the reproduction window 150 as shown in FIG. 22 (step 142). When the play button among the operation buttons 157 is pressed (YES in step 143), the reproduction process of the read video data is performed by the control device 115 (an example of a reproduction device and reproduction control means) (step 144). The selected always-recorded video is displayed in the video display area 151.
[0161] When the stop button among the operation buttons 157 is pressed during the playback of the always-recorded video (YES in step 145), the video being played is stopped by the control device 115 (step 146).
[0162] Even when a desired event recording video is selected from the event recording list, the video data representing the selected event recording video is read from the event recording area 125 of the SD card 23 by the SD card reader / writer 119 (step 141).
[0163] When an event recording video is selected, a representative image of the selected event recording video is displayed in the video display area 151 of the playback window 150 as shown in FIG. 23 (step 142). The representative image of the event recording video may be an image taken at the time when the event recording command was generated, or may be the first image of the event recording video.
[0164] When the play button among the operation buttons 157 is pressed (YES in step 143), the playback process of the read video data is performed by the control device 115 (step 144). The selected event recording video is displayed in the video display area 151. When the stop button among the operation buttons 157 is pressed during the playback of the event recording video (YES in step 145), the video being played is stopped by the control device 115 (step 146).
[0165] Even during the playback process of the continuously recorded video described above, the processes from step 92 to step 98 in FIG. 12 may be performed, and when a desired image is detected, it may be event-recorded in the event recording area 125 by the control device 115. When an event is recorded, it will be additionally displayed in the event recording list display area 152. Preferably, information indicating that it is an image related to the desired image is recorded in the header or the like of the frame constituting the continuously recorded video in which the desired image is detected during playback. When the video is played back, the control device 115 or the display control device 112 (an example of notification control means) controls the display device 111 so that the time of the found event recording or the like is displayed (an example of notification) in the playback window 150 of the display device 111 (an example of a notification device).
[0166] FIG. 24(A) and FIG. 24(B) are an example of a table showing the relationship between the frame numbers and address position information of the frames constituting one video.
[0167] Since the frame numbers and address position information are recorded in the header recording area 131 regardless of whether it is a continuously recorded video or an event-recorded video as described above, it is not necessary to create the tables shown in FIGS. 24(A) and 24(B). However, a table as shown in FIGS. 24(A) and 24(B) may be generated and recorded in the management area 123 of the continuously recorded area 122 or the management area 126 of the event recording area 125.
[0168] As described above, in the header recording area 131, frame numbers and address position information (information indicating the address at the head of the frame) are recorded corresponding to the frames constituting the continuously recorded video or the event-recorded video. The video data representing the continuously recorded video or the event-recorded video according to this embodiment is continuously recorded in the recording area 124 of the continuously recorded area 122 or the recording area 127 of the event recording area 125 in the order of frame numbers. For example, as shown in FIG. 24(A), one video composed of frames from frame 1 to frame E is recorded in the recording area 124 of the continuously recorded area 122 or the recording area 127 of the event recording area 125 for each frame.
[0169] In such a case, at least one of the frame number or the address position information may not be recorded normally and may be in an abnormal state. For example, although the frame numbers should be recorded continuously, they are not consecutive, the address position information is at an address position extremely larger than the previous address position information, the address position information indicates an address position before the previous address position information, etc. In such a case, although the video data itself is normal, the video of that part may not be playable.
[0170] For example, in FIG. 24(A), both the frame number and the address position information are recorded normally, but in FIG. 24(B), as shown by hatching, the frame numbers and the address position information from frame 11 to frame 20 have abnormal values, and the video represented by these frames from frame 11 to frame 20 often cannot be played back normally.
[0171] In this embodiment, when at least one of the frame number and the address position information is not recorded normally, rewrite them to normal values. Also, when the video data is not recorded normally such as when the video data is missing, generate the missing video data and write it to the frames that are not recorded normally.
[0172] FIG. 25 is a flowchart showing the repair processing procedure, and FIG. 26 shows the recording format of the SD card 23.
[0173] The repair processing may be started when the SD card 23 is loaded into the personal computer 110 and the playback software is started, or when a predetermined video is specified for video playback, the repair processing for that video may be started, or a repair mode button or the like may be provided in the playback window 150, and the repair processing may be started when the repair mode button is pressed.
[0174] When the restoration process of the desired video starts, the frame numbers recorded in the header recording area 131 of the frames constituting the video are read by the control device 115 (step 171). If the read frame numbers are not consecutive, the control device 115 determines that the non-consecutive frame numbers are not recorded correctly (YES in step 172). Also, as shown in Equation 1, since the sum of the address position information of a frame and the amount of video data of that frame is equal to the address position information of the next frame of that frame, if such a relationship is broken, the control device 115 determines that the address position information is not recorded correctly (YES in step 172).
[0175] Frame address position information + Video data amount of that frame = Address position information of the next frame... Equation 1
[0176] If at least one of the frame number or the address position information is not recorded correctly (NO in step 172), the control device 115 repairs the frame number or the address position information (or both the frame number and the address position information) that is not recorded correctly (step 173). As described above, if the frame number is not recorded correctly, the control device 115 rewrites the frame number so that it is consecutive in the order of the frames constituting the video. If the address position information is not recorded correctly, the control device 115 rewrites the address position information so that it has a normal value to establish the relationship shown in Equation 1.
[0177] Subsequently, the control device 115 determines whether the video data recorded in the recording area 124 or 127 is recorded normally (step 174). If the scenes represented by the frames are too different, for example, if an automobile was shown up to a certain frame, but the next frame is an image where the object cannot be recognized, such as completely black or completely white, then that frame is determined not to be recorded normally. However, even if it is a frame where the scene has suddenly changed, if it is determined to be a frame constituting the video of the same shooting scene based on the shooting position information or the like, it is determined to be recorded normally. For example, in the case of an accident, the scene will suddenly change, and this is to prevent such a situation from being determined not to be recorded normally. The user may visually check and determine frames that are not recorded normally.
[0178] If the video data is not recorded normally (NO in step 174), the abnormal video data is generated and the missing part (missing data) of the video is repaired (step 175). The method for repairing the missing part of the video will be described later. Referring to FIG. 26, as shown by hatching, if the video data recorded in the recording area 124 of the always recording area 122 from frame 10 to frame 20 is not recorded normally, the video data recorded in those frames from frame 10 to frame 20 is repaired. When the missing part of the video is repaired, it is preferable to record, as additional information, information indicating that the repaired frame is not a frame obtained by shooting but a repaired frame in the header recording area 131 or the footer recording area 131. This is to be able to distinguish between the actual scene obtained by shooting and the pseudo scene generated later.
[0179] When the missing part of the video is repaired, it is conceivable that the data amount of the frame changes, so the address position information of the repaired frame is likely to change. For this reason, the address position information is corrected according to Equation 1 (step 176).
[0180] In this embodiment, a learning model is used for repairing video data. However, without using the learning model, video data may be generated by image interpolation using frames before and after the frame to be repaired, etc.
[0181] FIGS. 27 and 28 show a method for generating a learning model. FIG. 27 is an example of a learning model, and FIG. 28 shows a state of training the learning model. FIG. 29 is a flowchart showing a learning model generation processing procedure.
[0182] In FIGS. 27 and 28, ten frames from frame 21 to frame 30 (an example of the second video data, teacher data, and the first part) are generated by machine learning using the frames from frame 1 to frame 20 before frame 21 and the frames from frame 31 to frame 50 after frame 30 (an example of the first video data, teacher data, and the second image). The frames from frame 1 to frame 50 are frames actually obtained by shooting in other vehicles. For example, the frames actually obtained by shooting in other vehicles are frames obtained by shooting with a drive recorder mounted on another vehicle that has traveled in the same place or another place where the same or similar scenes (e.g., scenery) can be shot. The machine learning is supervised learning that takes frames from frame 1 to frame 20 and frames from frame 31 to frame 50 as inputs and outputs frames from frame 21 to frame 30. Other frames other than frames 21 to 30 can be generated in the same way.
[0183] The learning model 180 includes an input layer 181, an intermediate layer 182, and an output layer 183. Each neuron in the input layer 181 is input in the order of the frame numbers from frame 1 to 20 and from frame 31 to 50 in the order of shooting. As shown in FIG. 28, data of each pixel such as pixel E1 data and pixel E2 data in the input layer 181 of the learning model 180 is input in the order of the pixel array for each pixel of each frame. In this way, video data (frames 1 to 20, frames 31 to 50) excluding some video parts (frames 21 to 30) is input to the input layer 181 (step 191).
[0184] The data input from the input layer 181 is output from the output layer 183 via the intermediate layer 182. The learning model 180 is trained so that data corresponding to frames 21 to 29 is output from each neuron in the output layer 183 according to the pixel array. As shown in FIG. 28, input is performed in the order of the frame numbers from frame 1 to 20 and from frame 31 to 50 in the order of shooting, and when the frames from frame 21 to 20 are output from the output layer 183, the learning of the learning model ends. In this way, learning is performed so that the output becomes a fixed video part (frames 21 to 30). Similarly, by performing a large number of learnings on other video data (step 192), it becomes possible to generate a learning model of the video of the defective video part using the frames of the video before and after the defective video part.
[0185] In FIGS. 27 to 29, the number of input frames is 40 and the number of output frames is 10, but the number of input frames and the number of output frames can be changed without being limited to these numbers of frames. Also, if both the number of input frames and the number of output frames are kept constant and the number of frames of the defective video to be generated is less than the number of output frames, frame thinning may be performed so that the number of frames of the video data representing the output frames becomes the number of frames of the defective video to be generated. If the number of frames of the defective video to be generated is more than the number of output frames, the number of frames can be adjusted by interpolation or the like using the video data representing the output frames.
[0186] Preferably, a plurality of learning models 180 suitable for different scenes are generated in advance, and the defective video portion is generated using the learning model corresponding to the scene. For example, the learning model 180 used for generating the defective video when the vehicle in front has an accident, the learning model 180 used for generating the defective video when a bird, a bicycle, a pedestrian, etc. suddenly appear in front, etc. Prepare the learning model 180 corresponding to the scene before or after the defective video portion, and use the learning model 180 corresponding to the scene to generate the defective video portion.
[0187] Furthermore, in the generation of the defective video, when the video is recorded by a drive recorder mounted on a vehicle, etc., the learning model 180 generated using a video with a high degree of match of the vehicle information of the vehicle on which the drive recorder that recorded the defective video is mounted may be used.
[0188] FIG. 30 is a flowchart showing the generation processing procedure of the defective portion of the video.
[0189] When the video data of the defective portion is found in the video data recorded on the SD card 23, the video data before and after the defective portion is input to the learned learning model 180 in the shooting order for each frame (step 201).
[0190] The input video data is learned in the learned learning model 180 (an example of the second learning model) (step 202), and the video data including the defective portion (an example of the third image) is input to the learning model 180 by the control device 115 (an example of the video data generation means), and the video data representing the video data of the defective portion (an example of the fourth image) is output from the learning model 180 in frame order (step 203). The video data of the defective portion obtained in this way is recorded in the defective portion of the recording areas 124 or 127 of the control device 115 (an example of the second recording control means) (step 204).
[0191] Even if there are recording errors such as frame numbers recorded on the memory card 23, they can be repaired. Even if there are missing parts in the video, they can be repaired.
[0192] In the above description, in response to "detecting an image representing the above scene", information regarding the image representing the scene is displayed in the event recording list display area 152, and when the displayed information is selected, an event recording video including the image representing the scene is displayed based on the video data representing the corresponding video. However, the present invention is not limited to this. During the playback of the video recorded on the SD card 23, in response to this selection, it may be possible to fast forward to the point in time when the image of the scene is displayed, or jump (seek) to that point in time during the playback of the image.
[0193] In each of the embodiments described above, various processes have been performed based on video data. However, processes may also be performed based on data that is not in video format. The data that is not in video format may be composed of a set of data in a plurality of still image formats indicating still images taken in chronological order.
[0194] In the above-described embodiments, the drive recorder 1 had an acceleration sensor 18 for detecting the state of the vehicle (such as the driving state). However, if the acceleration sensor 18 is not used for generating the learning model or detecting the scene, it may not be provided. In the second embodiment, the drive recorder 1 may not execute the process of step 91, or may not execute the process of detecting an event by the acceleration sensor 18 among the processes of step 91. In this case, without using the second degree of coincidence, a specific scene image is detected based on whether the first degree of coincidence is greater than or equal to a threshold value. Even in such a case, the drive recorder 1 may not have the acceleration sensor 18. Even when the drive recorder 1 uses a sensor other than the acceleration sensor 18 to detect the state of the vehicle, it can be configured not to have the sensor.
[0195] In the above-described embodiments, if there are other machine learning or other alternative means for the functions realized using deep learning technology, these may be used for realization. For example, pattern matching technology may be used as a means for detecting images with a high degree of similarity.
[0196] Note that the scope of the present invention is not limited to the configurations explicitly described in the specification or limited thereto, but also includes combinations of various aspects of the present invention disclosed in this specification within its scope. Among the present invention, the configuration for which a patent is sought is specified in the appended claims, but even a configuration that is not currently specified in the claims has the intention of making the configuration disclosed in this specification the claims in the future.
[0197] The present invention of the application is not limited to the configurations described in the above-described embodiments. The constituent elements of the above-described embodiments and modification examples may be arbitrarily selected and combined. Also, any constituent element of each embodiment and modification example may be arbitrarily combined with any constituent element described in the means for solving the invention or a constituent element embodying any constituent element described in the means for solving the invention. There is also an intention to obtain rights in the amendment or divisional application of the present application regarding these. Even if there are descriptions such as "in the case of ~" or "when ~", it is not described as a configuration limited to that case or that time. Configurations other than these cases and times are also disclosed and have the intention of obtaining rights. Also, the descriptions with an order are not limited to this order. Configurations with some parts deleted or the order changed are also disclosed and have the intention of obtaining rights.
[0198] In addition, due to the change to the design registration application, it has the intention of obtaining rights for the overall design or partial design. Although the drawings depict the entire device in solid lines, the drawings include not only the overall design but also partial designs claimed for some parts of the device. For example, not only can some members of the device be regarded as partial designs, but the drawings also include partial designs for some parts of the device regardless of their relationship with the members. As some parts of the device, they can be either some members of the device or parts of those members. Regarding the overall design, of course, it has the intention of obtaining rights for partial designs in which any part of the solid-line portion of the drawings is made into a dashed-line portion. Also, for modules, members, parts, etc. inside the housing of the device, those shown in the drawings are all independently subject to transactions, and similarly, it has the intention of making changes to the design registration application to obtain rights.
Explanation of Reference Signs
[0199] 1: Drive recorder, 3: Front glass, 4: Rearview mirror, 5: Cigarette socket, 6: Power cable, 10: SD card slot, 11: Display, 12: Operation button, 13: Joint rail, 14: SD card reader / writer, 15: Speaker, 16: GPS receiver, 17: Camera, 18: Acceleration sensor, 19: Communication circuit, 20: Controller, 20c: RAM, 20d: Timer, 21: LTE module, 22: Database, 23: SD card, 30: Server, 31: Control device, 33: Memory, 34: Hard disk drive, 35: Hard disk, 40: Learning model, 40A: First learning model, 40B: Second learning model, 40C: Third learning model, 41: Input layer, 42: Intermediate layer, 43: Output layer, 43a: Neuron, 50: Target image, 51: First target image, 52: Second target image, 53: Third target image, 71: Generator, 72: Discriminator, 73: Weight modifier, 80: Learning model, 81: Input layer, 82: Intermediate layer, 83: Output layer, 105: Learning model, 106: Input layer, 107: Intermediate layer, 108: Output layer, 110: Personal computer, 111: Display device, 112: Display control device, 113: Communication device, 114: Memory, 115: Control device, 116: Input device, 117: Hard disk drive, 118: Hard disk, 119: SD card reader / writer, 121: File system area, 122: Constant recording area, 123: Management area, 124: Recording area, 125: Event recording area, 126: Management area, 127: Recording area, 131: Header recording area, 132: Frame image data recording area, 133: Footer recording area, 150: Playback window, 151: Video display area, 152: Event recording list display area, 153: Map image display area, 154: Indoor enhancement button, 155: Normal display button, 156: Time display area, 157: Operation button, 158: Vehicle state display area, 159: Acceleration graph display area, 160: Vehicle information etc. display area, 161: Recording list display area, 162: Acceleration display area, 163: Acceleration display area, 180: Learning model, 181: Input layer, 182: Intermediate layer, 183: Output layer, 421: Input layer, E1: Pixel, E2: Pixel, FR1: Frame, FRn: Frame, P1: Pixel, P21: Pixel
Claims
1. A system having a function of detecting an image representing a specific image scene from images taken by a vehicle at each of a plurality of time points based on data defining the specific image scene that can be taken by a vehicle, the system comprising: The device is provided with a function of displaying information about an image that represents the detected scene of the specific image, and, depending on the selection of information about the image that represents the scene of the specific image that has been displayed, fast-forwarding to a point in time at which an image that represents the scene of the specific image is displayed during playback of images taken by a vehicle at each of the multiple points in time, or a function of jumping to the point in time. A system characterized by.
2. The data defining the scene of the specific image is a trained learning model that has been trained by machine learning using the specific image as training data, and the function of detecting an image representing the scene of the specific image is configured to detect an image representing the scene of the specific image by inputting the images taken at the multiple time points into the trained model. The system of claim 1 .
3. The information on the image representing the specific image scene is displayed in a list display area on the screen, and the information on the image representing the specific image scene is selected from the list displayed in the list display area.
3. The system according to claim 1 or 2, characterized in that
4. The list display area is an area for displaying a list of event records. The system of claim 3 .
5. The system is provided with a function for recording information indicating that the image represents a specific scene detected in the image in each frame constituting the images captured by the vehicle at each of a plurality of points in time.
5. A system according to claim 1, wherein:
6. A program for causing a computer to realize the functions of the system according to any one of claims 1 to 5.
Citation Information
Patent Citations
Danger prediction method
JP2017162438A
Driving information recording system, driving information recording method, and program
JP2018169667A
Drive recorder, display device for drive recorder, and program
JP2018196066A
Information processing method, information processing apparatus and information processing program
JP2019016227A
Display control device, display control method, and program
JP2020013070A