Information processing device, information processing method, and program

The system efficiently estimates vanishing points from dashcam footage to isolate and correct the desired area, addressing data acquisition challenges by reducing the impact of camera and vehicle variability.

JP2026083887APending Publication Date: 2026-05-20RECRUIT
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
RECRUIT
Filing Date
2024-11-08
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Existing systems face challenges in efficiently acquiring data from vehicle-mounted dashcams due to varying factors like camera type, field of view, installation angle, and vehicle specifications, making it difficult to consistently extract data from a desired area.

Method used

An information processing system that extracts feature points from video frames and estimates vanishing points to accurately determine the region of interest, generating mask data to isolate and correct the desired area for data acquisition.

Benefits of technology

This approach allows for efficient data acquisition from dashcam footage by reducing the influence of varying factors, improving the quality and reliability of data collected from vehicle-mounted cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026083887000001_ABST
    Figure 2026083887000001_ABST
Patent Text Reader

Abstract

To provide a mechanism for efficiently estimating vanishing points. [Solution] An information processing device comprising: an acquisition unit that acquires video footage captured by an in-vehicle camera; an extraction unit that extracts feature points from frames included in the video; and an estimation unit that estimates the vanishing point of the video based on a first feature point extracted from a first frame of the video and a second feature point corresponding to the first feature point, extracted from a second frame of the video.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] Vehicles are being equipped with devices such as dashcams to acquire various types of data.

[0003] For example, Patent Document 1 discloses a technique for measuring the number of people passing by at a specific location in the analysis of video data from a dashcam.

[0004] Patent Document 2 discloses a technique for detecting roadways and sidewalks by detecting the boundary lines of an image.

[0005] Patent Document 3 discloses a technique for determining the normality of the imaging direction based on whether or not a vanishing point is detected in the image acquired from a drive recorder. [Prior art documents] [Patent Documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2020-160811 [Patent Document 2] Japanese Patent Publication No. 2019-087166 [Patent Document 3] Japanese Patent Publication No. 2015-207888 [Overview of the project] [Problems that the invention aims to solve]

[0007] A particular aspect of the present invention aims to provide a mechanism for efficiently estimating vanishing points. [Means for solving the problem]

[0008] An information processing device according to one aspect of the present invention includes: an acquisition unit that acquires video footage captured by an in-vehicle camera; an extraction unit that extracts feature points from frames included in the video footage; and an estimation unit that estimates the vanishing point of the video footage based on a first feature point extracted from a first frame of the video footage and a second feature point extracted from a second frame of the video footage that corresponds to the first feature point. According to this embodiment, a mechanism for efficiently estimating vanishing points can be provided.

[0009] An information processing method according to one aspect of the present invention includes an acquisition step in which one or more computers acquire video footage captured by an in-vehicle camera; an extraction step in which feature points are extracted from frames included in the video footage; and an estimation step in which the vanishing points of the video footage are estimated based on a first feature point extracted from a first frame of the video footage and a second feature point extracted from a second frame of the video footage that corresponds to the first feature point. According to this embodiment, a mechanism for efficiently estimating vanishing points can be provided.

[0010] A program according to one aspect of the present invention causes one or more computers to perform the following steps: an acquisition step of acquiring video footage captured by an in-vehicle camera; an extraction step of extracting feature points from frames included in the video footage; and an estimation step of estimating the vanishing points of the video footage based on a first feature point extracted from a first frame of the video footage and a second feature point extracted from a second frame of the video footage that corresponds to the first feature point. According to this embodiment, a mechanism for efficiently estimating vanishing points can be provided. [Effects of the Invention]

[0011] According to a predetermined aspect of the present invention, a mechanism for efficiently estimating vanishing points can be provided. [Brief explanation of the drawing]

[0012] [Figure 1]It is a diagram for explaining the outline of an information processing system according to an embodiment of the present invention. [Figure 2] It is a diagram illustrating the configuration of an information processing system according to an embodiment of the present invention. [Figure 3] It is a diagram illustrating the hardware configuration of an information processing device and an in-vehicle device according to an embodiment of the present invention. [Figure 4] The functional configuration of an information processing device according to an embodiment of the present invention is illustrated. [Figure 5] The configuration of a database according to an embodiment of the present invention is illustrated. [Figure 6] It is a diagram for explaining the operation of an information processing device according to an embodiment of the present invention. [Figure 7] It is a diagram for explaining the operation of an information processing device according to an embodiment of the present invention. [Figure 8] It is a diagram for explaining the operation of an information processing device according to an embodiment of the present invention. [Figure 9] It is a diagram for explaining the operation of an information processing device according to an embodiment of the present invention. [Figure 10] It is a diagram for explaining the operation of an information processing device according to an embodiment of the present invention. [Figure 11] It is a diagram for explaining the operation of an information processing device according to an embodiment of the present invention. [Figure 12] It is a diagram for explaining the operation of an information processing device according to an embodiment of the present invention. [Figure 13] It is a diagram for explaining the operation of an information processing device according to an embodiment of the present invention. [Figure 14] It is a diagram for explaining the operation of an information processing device according to an embodiment of the present invention. [Figure 15] It is a flowchart illustrating the operation of an information processing device according to an embodiment of the present invention. [Figure 16] It is a flowchart illustrating the operation of an information processing device according to an embodiment of the present invention.

Embodiments for Carrying Out the Invention

[0013] Embodiments of the present invention will be described with reference to the attached drawings. In each drawing, components denoted by the same reference numerals have the same or similar configurations.

[0014] 1. Overview of the Information Processing System Disclosed An information processing system according to one embodiment of the present invention aims to improve the efficiency of acquiring data related to the activity along the road on which a vehicle equipped with a drive recorder is traveling, from video footage captured by the drive recorder.

[0015] Figure 1 illustrates the forward field of view e101 captured by a dashcam mounted in a vehicle. The field of view e101 includes the road r11 on which the vehicle is traveling and the sidewalk r12. Data related to the sidewalk r12 area may include information on the number of people and the amount of light at night. Such data is useful for analyzing the level of activity and congestion.

[0016] However, attempting to acquire data on a predetermined area based on videos recorded by a dashcam presented various problems. For example, the recorded videos varied due to a variety of fluctuating factors, such as the type of dashcam, field of view, installation angle, height, and position, as well as the specifications of the vehicle on which the dashcam is installed and the driving conditions. This made it difficult to efficiently acquire data on a predetermined area. Therefore, it is desirable to reduce the influence of the various factors mentioned above and efficiently acquire data on the desired area.

[0017] Therefore, the inventors focused on the fact that information regarding vanishing points can be used when acquiring data about a desired region from a video.

[0018] An information processing system according to one embodiment of the present invention extracts feature points from frames contained in a video and estimates vanishing points based on the extracted feature points. This makes it possible to efficiently estimate vanishing points. The estimated information regarding vanishing points can be used to obtain data concerning a desired region while reducing the influence of the various factors described above.

[0019] Furthermore, the information processing system according to one embodiment of the present invention generates mask data indicating a region to be masked from a part of the video based on the estimated vanishing point. This makes it possible to acquire data for a desired region while reducing the influence of the various factors mentioned above.

[0020] 2. Configuration of the Information Processing System Figure 2 illustrates the configuration of an information processing system 1 according to one embodiment of the present invention. The information processing system 1 comprises an information processing device 10 and an in-vehicle device 20. Each device constituting the information processing system 1 is connected to the other via a communication network N.

[0021] The communication network N is a wired or wireless network. Examples include the Internet, LAN (Local Area Network), dedicated lines, telephone lines, corporate networks, mobile communication networks, Bluetooth®, WiFi (Wireless Fidelity), other communication lines, and combinations thereof. Each component of the information processing system 1 may be located in the same vehicle or in different vehicles. Furthermore, the information processing device 10 may be installed in a predetermined facility, and the in-vehicle device 20 may be installed in a predetermined vehicle.

[0022] The information processing device 10 is a device that provides a function for estimating the vanishing point of a video acquired from the in-vehicle device 20. In one embodiment, the information processing device 10 is also a device that provides a function for generating mask data indicating a region to be masked from a part of the video acquired from the in-vehicle device 20. Examples of the information processing device 10 include desktop computers, laptop computers, and other dedicated or general-purpose computers. The information processing device 10 may consist of a single computer or multiple computers on a communication network N.

[0023] The in-vehicle device 20 is a device mounted in a vehicle that has the function of recording video to be transmitted to the information processing device 10. Examples include drive recorders, car cameras, and dash cameras. The vehicle is not particularly limited, but examples include cars, motorcycles, bicycles, trains, railways, monorails, and the like.

[0024] In this disclosure, the in-vehicle device 20 is described as being connected to a communication network N, but the in-vehicle device 20 is not necessarily required to be connected to the communication network N. For example, the in-vehicle device 20 may be connected to the communication network N via other devices connected to the in-vehicle device 20.

[0025] In this disclosure, the in-vehicle device 20 is described as being mounted on a vehicle, but the scope of application of this disclosure is not limited to this. For example, the mechanism provided by the information processing system of this disclosure may be used for video footage taken while mounted on a helicopter, airplane, ship, submarine, spacecraft, etc.

[0026] 3. Hardware Configuration Figure 3 illustrates the hardware configurations of the information processing device 10 and the in-vehicle device 20, respectively.

[0027] The information processing device 10 and the in-vehicle device 20 each include a processor 11, a storage device 12, a communication interface 13, an input device 14, and an output device 15.

[0028] Processor 11 includes CPUs (Central Processing Units), GPUs (Graphics Processing Units), etc.

[0029] The storage device 12 is a memory, HDD (Hard Disk Drive), and / or SSD (Solid State Drive), etc. The memory is, for example, RAM (Random Access Memory) and / or ROM (Read Only Memory).

[0030] Communication IF13 is a device that communicates data with other devices via a communication network N, either by wire or wirelessly.

[0031] The input device 14 accepts information input. Examples include a keyboard, touch panel, mouse, camera, and / or microphone. The input device 14 of the in-vehicle device 20 includes at least a camera capable of recording video.

[0032] The output device 15 outputs information. For example, it may be a display, a touch panel, and / or a speaker.

[0033] 4. Functional Configuration Figure 4 illustrates the functional configuration of the information processing device 10. The information processing device 10 comprises a storage unit 100 and a control unit 110. The control unit 110 comprises an acquisition unit 111, an extraction unit 112, an estimation unit 113, a generation unit 114, and a correction unit 115. The storage unit 100 can be realized using a storage device 12 of the information processing device 10. Each configuration of the control unit 110 can be realized by the processor 11 of the information processing device 10 executing a program stored in the storage device 12. The program can be stored in a storage medium. The storage medium storing the program may be a non-transitory computer-readable medium. The non-transitory storage medium may be, for example, a USB (Universal Serial Bus) memory or a CD-ROM (Compact Disc Read-Only Memory).

[0034] (Storage unit 100) The memory unit 100 stores various types of information necessary for the information processing device 10 to operate. For example, the memory unit 100 stores feature point DB 101.

[0035] Figure 5 illustrates the configuration of the feature point DB101. The feature point DB101 shown in Figure 5 stores frames identified by frame numbers and coordinates that identify the position of objects contained within each frame. For example, frame f1 and the coordinates 1a(x) corresponding to the position of "the top of streetlamp A" contained within frame f1. 10 ,y 10 ), coordinate 1a'(x) corresponds to the position of "bottom of street lamp A" included in frame f1. 20 ,y 20 ), and coordinate 1b(x) corresponding to the position of the "end of signboard B" included in frame f1. 30 ,y 30 ) are stored in correspondence with each other. Also, for example, frame f2 and the coordinates 2a(x) corresponding to the position of "the head of street lamp A" contained in frame f2 are stored. 11 ,y 11) The coordinates 2a’(x 21 , y 21 ) corresponding to the position of "the bottom of street lamp A" included in the frame f2, and the coordinates 2b(x 31 , y 31 ) corresponding to the position of "the end of signboard B" included in the frame f2 are stored in association with each other.

[0036] By referring to the feature point DB101, the coordinates of each feature point corresponding to the object common between the frames can be obtained. For example, by referring to the feature point DB101 shown in FIG. 5, as the coordinates of each feature point corresponding to the object of "the head of street lamp A" common between the frame f1 and the frame 2, the coordinates 1a(x 10 , y 10 ) and the coordinates 2a(x 11 , y 11 ) can be obtained.

[0037] (Acquisition unit 111) The acquisition unit 111 acquires the video captured by the in-vehicle device 20. For example, the acquisition unit 111 receives the video captured by the camera of the in-vehicle device 20. The video may be acquired sequentially or intermittently.

[0038] In one embodiment, the acquisition unit 111 may detect and acquire the video of the vehicle equipped with the in-vehicle device 20 moving straight based on GPS information or the like. (Extraction unit 112) The extraction unit 112 extracts feature points from the frames included in the video acquired by the acquisition unit 111. As the method for extracting feature points from the image of the frame, those well-known to those skilled in the art can be used. For example, the method disclosed in Japanese Patent No. 4689758 can be used.

[0039] Figure 6 illustrates an example of feature point extraction. Frames f1 and f2 shown in Figure 6 are examples of frames included in a video acquired from the in-vehicle device 20. Here, frames f1 and f2 are assumed to be consecutive frames. The extraction unit 112 extracts feature point p1a corresponding to the top of streetlamp A included in frame f1, and extracts feature point p1b corresponding to the top of streetlamp A included in frame f2. Since feature points p1a and p1b both correspond to the common object, the top of streetlamp A, they are in a corresponding relationship with each other.

[0040] In one embodiment, the extraction unit 112 may generate coordinate information corresponding to the position of a feature point for each frame. For example, the extraction unit 112 assigns coordinates on a pixel-by-pixel basis for each frame using a coordinate system such as a Cartesian coordinate system. The extraction unit 112 then identifies the coordinates corresponding to the position of a predetermined feature point based on this coordinate system. The extraction unit 112 also stores the identified coordinate information in the feature point DB 101 as coordinate information corresponding to the position of the feature point.

[0041] (Estimation part 113) The estimation unit 113 estimates the vanishing point of the video based on the feature points extracted by the extraction unit 112. For example, the estimation unit 113 estimates the vanishing point of the video based on the feature point p1a extracted from frame f1 of the video and the feature point p1b extracted from frame f2.

[0042] In one embodiment, the estimation unit 113 estimates the vanishing points of the video as coordinates where the number of intersections (sometimes referred to as the "number of intersections" in this disclosure) of the lines drawn passing through mutually corresponding feature points with respect to a plurality of feature points included in the video is equal to or greater than a reference value.

[0043] For example, the estimation unit 113 reads the coordinate information of the first and second feature points from the feature point DB 101. The estimation unit 113 also calculates a function that identifies a straight line passing through the read coordinates using a coordinate system. In this way, it calculates functions for multiple straight lines passing through mutually corresponding feature points. Then, based on each of the calculated functions, the estimation unit 113 identifies the coordinates corresponding to the intersections where the straight lines intersect when each of the straight lines is drawn using a coordinate system, and can use this to estimate the vanishing points.

[0044] Figure 7 illustrates an example of vanishing point estimation. Diagram e701 in Figure 7 illustrates the plotting of coordinate information for each corresponding feature point read from feature point DB101 in a coordinate system. Diagram e701 shows a line L1 passing through the coordinates of feature point p1a, which corresponds to the top of streetlamp A in frame f1, and feature point p1b, which corresponds to the top of streetlamp A in frame f2.

[0045] In one embodiment, the estimation unit 113 calculates a straight line passing through corresponding feature points between two consecutive frames. For example, it calculates a straight line L1 passing through feature point p1a, which corresponds to the top of streetlamp A in frame f1, and feature point p1b, which corresponds to the top of streetlamp A in frame f2. It also calculates a straight line L2 connecting feature point p1b, which corresponds to the top of streetlamp A in frame f2, and feature point p1c, which corresponds to the top of streetlamp A in frame f3. Furthermore, it calculates a straight line L3 connecting feature point p1c, which corresponds to the top of streetlamp A in frame f3, and feature point p1d, which corresponds to the top of streetlamp A in frame f4.

[0046] Diagram e702 in Figure 7 illustrates the plotting of multiple lines calculated by the estimation unit 113 in a coordinate system. Diagram e702 shows multiple lines and the intersections where the lines intersect. The estimation unit 113 identifies the coordinates corresponding to the location of the intersections and can be used to estimate the vanishing points.

[0047] For example, the estimation unit 113 sets the reference value for the number of intersections to 3 and estimates coordinates with 3 or more intersections as vanishing points. In the example of diagram e702, the number of intersections at coordinate a, which corresponds to the location of the intersection, is 3 (3 lines intersect at coordinate a), and the number of intersections at coordinate b is 2 (2 lines intersect at coordinate b). Therefore, coordinate a with 3 or more intersections is estimated as a vanishing point.

[0048] In one embodiment, the estimation unit 113 may estimate one or more vanishing points for a single video. For example, multiple vanishing points may be estimated for a single video, such as vanishing point v1 corresponding to frames f1 to f50 included in the video, vanishing point v2 corresponding to frames f51 to f80 included in the video, and vanishing point v3 corresponding to frames f81 to f100 included in the video.

[0049] In one embodiment, the number of frames corresponding to one vanishing point may be the number of consecutive frames in which the coordinates of the vanishing point are included in a predetermined coordinate region.

[0050] For example, the estimation unit 113 sets a reference value for the number of intersections used to estimate the vanishing point and a coordinate region used to estimate the vanishing point. Specifically, for example, a circle with radius r centered at coordinates (x1, y1) is set as the coordinate region used to estimate the vanishing point. The reference value for the number of intersections used to estimate the vanishing point is set to 1000. Then, when the coordinate region contains coordinates where the number of intersections of the calculated lines is 1000 or more, the estimation unit 113 sets the number of frames that formed the basis of the calculation as the number of frames corresponding to one vanishing point. For example, if the number of such frames is 50, one vanishing point is estimated for the portion of the video corresponding to 50 frames.

[0051] Furthermore, for example, the estimation unit 113 may further set a suitable number of coordinates having a predetermined number of intersections included in the coordinate region and use this for estimating the vanishing point. For example, the estimation unit 113 sets the suitable number of coordinates having a predetermined number of intersections included in the coordinate region to 3 or more. Then, when the coordinate region contains three coordinates with 1000 or more intersections of the calculated lines, the estimation unit 113 sets the number of frames used as the basis for the calculation to be the number of frames corresponding to one vanishing point.

[0052] In one embodiment, the estimation unit 113 may set a reference value for the number of intersections used to estimate the vanishing point based on the number of straight lines calculated. In another embodiment, the estimation unit 113 may set a reference value for the number of intersections used to estimate the vanishing point based on the number of intersections formed by the straight lines calculated.

[0053] In one embodiment, the estimation unit 113 may estimate vanishing points using a mathematical model such as a Bayesian network model. The mathematical model such as a Bayesian network model may be constructed by the estimation unit 113 based on information input to the information processing device 10, or it may be constructed by another device and referenced by the estimation unit 113 as appropriate. A method well known to those skilled in the art can be used to construct the Bayesian network model.

[0054] The estimation unit 113 may generate coordinate information corresponding to the estimated vanishing point location and store it in the storage unit 100. For example, the estimation unit 113 generates coordinate information corresponding to the estimated vanishing point for frames f1 to f50 and stores it in the feature point DB 101.

[0055] (Generation unit 114) The generation unit 114 may determine the area to mask the video and generate mask data based on the coordinates corresponding to the vanishing point estimated by the estimation unit 113 and the supplementary coordinates. The supplementary coordinates can be specified using the same coordinate system as the coordinates corresponding to the vanishing point and are used to assist in determining the area to mask using the vanishing point. The supplementary coordinates may be any position within the frame, but for example, the coordinates of the top left of the frame may be the first supplementary coordinate and the coordinates of the bottom left may be the second supplementary coordinate.

[0056] In one embodiment, the generation unit 114 determines the area to mask the video based on a coordinate region specified by the coordinates corresponding to the vanishing point and the supplementary coordinates. For example, the generation unit 114 specifies a predetermined coordinate region by a line segment passing through the coordinates corresponding to the vanishing point and the supplementary coordinates specified by the same coordinate system, and determines the area outside that coordinate region as the area to mask the video. Specifically, for example, the area outside the triangle formed by connecting the coordinates corresponding to the vanishing point and the first supplementary coordinate and the second supplementary coordinate may be determined as the area to mask the video. In this case, the area to mask the video may also be the portion within the predetermined region from the vanishing point.

[0057] In one embodiment, the generation unit 114 determines an area to mask the video based on the coordinates corresponding to the vanishing point and supplementary information applied to the same coordinate system. For example, the generation unit 114 identifies a coordinate region of a circle centered at the coordinates of the vanishing point based on the coordinates corresponding to the vanishing point and supplementary information relating to the radius of a circle centered at the coordinates corresponding to the vanishing point, and determines the area outside this coordinate region as the area to mask the video. The supplementary information is not particularly limited, but examples include predetermined coordinates (which can be used as the center coordinates of a circle or as the vertex coordinates of a polygon), angles (which can be applied to predetermined coordinates and line segments passing through those coordinates to determine other coordinate positions or determine the shape of an arc), radii when drawing circular figures, and functions for drawing figures specified on the coordinate system.

[0058] Mask data is data in which different information is assigned to the area to be masked and the other areas across a predetermined number of frames in a video. By applying mask data to the original video, the area to be masked in the video can be masked. In a masked video, the area to be masked is masked. The masking method can be any method well known to those skilled in the art. It is not particularly limited, but may include methods such as randomization, shuffling, encryption, hashing, tokenization, or nullification.

[0059] Figures 8-12 illustrate an example of mask data generation. Diagram e801 in Figure 8 illustrates the field of view captured by a dashcam. Diagram e801 shows the road r81 traveled by the vehicle equipped with the dashcam, and the sidewalk r82 along that road. An example of generating mask data to efficiently acquire data related to the area of ​​sidewalk r82 (corresponding to "Area to be counted Y80" in the figure) using the estimated vanishing point will be explained below with reference to Figures 9-12.

[0060] Diagram e901 in Figure 9 illustrates the plotting of the coordinate information of the vanishing point v91, read from feature point DB101, in a coordinate system.

[0061] Diagram e1001 in Figure 10 illustrates the process of determining supplemental coordinates. Diagram e1001 defines angles θ1 and θ2 used to determine supplemental coordinates, with reference to a baseline (horizontal line) L10 passing through vanishing point v91.

[0062] Diagram e1101 in Figure 11 illustrates the determination of supplementary coordinates. Diagram e1101 shows supplementary coordinates ws1 and ws2 determined based on the vanishing point v91, angle θ1, and angle θ2. In the example in Figure 11, the generation unit 114 determines supplementary coordinate ws1 on line segment L1, which is drawn by applying angle θ1 to the reference line L10 shown in Figure 10. Similarly, the generation unit 114 determines supplementary coordinate ws2 on line segment L2, which is drawn by applying angle θ2 to the reference line L10 shown in Figure 10. However, the method of determining supplementary coordinates ws1 and ws2 is not limited to this; for example, the coordinate corresponding to the upper left of the screen may be used as supplementary coordinate ws1, and the coordinate corresponding to the lower left of the screen may be used as supplementary coordinate ws2.

[0063] Diagram e1201 in Figure 12 illustrates a mask created using mask data generated from Figures 9-11. In the example in Figure 12, a triangular coordinate region is identified by a line segment passing through vanishing point v91 and supplemental coordinate ws1, a line segment passing through vanishing point v91 and supplemental coordinate ws2, and a line segment passing through supplemental coordinates ws1 and ws2. The area outside this coordinate region is determined to be the area to be masked. In this example, mask data is generated to efficiently acquire data related to the area Y80 to be counted in Figure 8.

[0064] (Correction section 115) The correction unit 115 may modify the area to be masked and correct the mask data based on further supplementary information (additional supplementary information) that can be applied to the same coordinate system as the coordinate system indicating the area to be masked. The additional supplementary information can be prepared in the same way as the supplementary information determined by the generation unit 114.

[0065] In one embodiment, the correction unit 115 modifies the area to be masked based on the coordinate region specified by the coordinates and supplementary coordinates corresponding to the vanishing point used to determine the area to be masked, and / or additional supplementary coordinates. The modification of the area to be masked can be performed in the same manner as the determination of the area to be masked by the generation unit 114. The modification of the area to be masked can be done by overwriting the mask data that defines the area to be masked determined by the generation unit 114. The modification of the area to be masked can be done by associating the area to be masked determined by the generation unit 114 with the modified area to be masked and storing it in the storage unit 100. The modification of the area to be masked can also be done by storing information regarding the difference between the area to be masked determined by the generation unit 114 and the modified area to be masked. As an example, such information regarding the difference can be used to provide information regarding the reliability of data obtained from video. For example, it can be used to provide information regarding the reliability of data obtained from video to which the uncorrected mask data has been applied and the reliability of data obtained from video to which the corrected mask data has been applied.

[0066] Figures 13-14 illustrate an example of mask data correction. Diagram e1301 in Figure 13 illustrates the field of view captured by the dashcam. Compared to Figure 8, in the example in Figure 13, the area closer to the vehicle equipped with the dashcam (corresponding to "Area to be counted Y130" in Figure 13) is the area from which data is to be acquired from the video. By using an area closer to the vehicle equipped with the dashcam, it is expected that errors in counting people will be less likely to occur when acquiring data on crowd levels, and the quality of the acquired data will be improved.

[0067] Diagram e1401 in Figure 14 illustrates the application of corrected mask data, using the mask data from diagram e1201 shown in Figure 12 as the uncorrected mask data. In the example in Figure 14, compared to the case in Figure 12, areas further away from the vehicle equipped with the drive recorder are included in the masked area, and it is expected that the quality of data acquired from the video to which the mask data has been applied will improve.

[0068] 5. Operation Figure 15 is a flowchart illustrating the operation of an information processing device 10 according to one embodiment of the present invention. The information processing device 10 acquires video footage captured by the in-vehicle camera of the in-vehicle device 20 using the acquisition unit 111 (step S150).

[0069] Next, the information processing device 10 extracts feature points from the frames contained in the video using the extraction unit 112 (step S151).

[0070] For example, if a video totaling 10 seconds is acquired, feature points are extracted from each of the frames f1 to f100 contained in the video, and coordinate information corresponding to the position of the feature points is generated and stored in the feature point DB101, as illustrated in Figure 5.

[0071] Furthermore, the information processing device 10 estimates the vanishing point of the video based on a first feature point extracted from the first frame of the video and a second feature point extracted from the second frame that corresponds to the first feature point (step S152).

[0072] For example, feature points with corresponding relationships in frames f1 to f100 are matched, and the calculation of lines connecting these corresponding feature points is repeated. Then, the vanishing points are estimated based on the number of intersections where these calculated lines intersect (intersection count).

[0073] For example, when a coordinate appears where the number of intersections exceeds a predetermined threshold, that coordinate is estimated as the vanishing point of the video corresponding to the frame processed up to that point.

[0074] For example, if a predetermined number of coordinates with a number of intersections equal to or greater than a predetermined standard value are included consecutively within a predetermined coordinate region, that coordinate region is estimated as the region of the vanishing point of the video corresponding to the frame for which the predetermined number of coordinates were calculated.

[0075] For example, if a predetermined number of coordinates with a number of intersections equal to or greater than a predetermined threshold are included consecutively within a predetermined coordinate region, the average of the coordinates included in that coordinate region is estimated as the vanishing point of the video corresponding to the frame for which the predetermined number of coordinates were calculated.

[0076] Figure 16 is a flowchart illustrating the operation of an information processing device 10 according to one embodiment of the present invention. The information processing device 10 reads coordinates corresponding to the vanishing points estimated by the estimation unit 113 from the feature point DB 101, and uses these vanishing points to generate mask data indicating a region to be masked from a part of the video.

[0077] The information processing device 10 estimates the vanishing point using the estimation unit 113 (step S160), and the generation unit 114 determines supplementary coordinates based on the coordinates corresponding to the estimated vanishing point's position (step S161).

[0078] Next, the information processing device 10, using the generation unit 114, determines the area to be masked based on the coordinates corresponding to the vanishing point and the supplementary coordinates, and generates mask data (step S162).

[0079] Each of the steps described above can be rearranged or performed in parallel as long as it does not create any inconsistencies in the content.

[0080] In this disclosure, the terms "first frame" and "second frame" simply mean different frames and are not necessarily intended to refer to consecutive frames. Similarly, the terms "first feature point" and "second feature point" simply mean different feature points and are not intended to provide any indication of a correspondence between these feature points.

[0081] 6. Other Embodiments In the embodiments described above, the purpose was to acquire data relating to the area of ​​the sidewalk along the road on which the vehicle equipped with the drive recorder is traveling, from video footage captured by the drive recorder. However, the scope of application of this disclosure is not limited to this. For example, even when the road is a single lane and data relating to the area of ​​the sidewalks on both sides of the road can be acquired in the same manner by estimating the vanishing point and generating mask data that indicates the area to be masked in part of the video based on the estimated vanishing point. Similarly, even when it is desired to acquire data relating to the area above the road, in front of the vehicle, behind the vehicle, etc., the vanishing point can be estimated and mask data that indicates the area to be masked in part of the video based on the estimated vanishing point.

[0082] In the above-described embodiment, a method was explained in which one coordinate whose number of intersections is greater than or equal to a reference value is estimated as the vanishing point. However, the method is not limited to this, and the coordinate with the largest number of intersections may be estimated as the vanishing point. For example, if there are multiple coordinates whose number of intersections is greater than or equal to a reference value, the estimation unit 113 selects the coordinate with the largest number of intersections and estimates the selected coordinate as the vanishing point of the video.

[0083] In the embodiments described above, a method for estimating a predetermined single coordinate as the vanishing point was explained, but the invention is not limited to this, and the vanishing point may be estimated based on multiple coordinates. For example, if there are multiple coordinates for which the number corresponding to an intersection is greater than or equal to a standard value, the estimation unit 113 estimates the average coordinate of these multiple coordinates as the vanishing point.

[0084] The estimation unit 113 may estimate vanishing points when it draws a straight line connecting a first feature point and a second feature point with respect to multiple feature points included in the video, and the number of intersections of the said straight line is equal to or greater than a reference value within a predetermined coordinate region. Furthermore, for example, the estimation unit 113 estimates the vanishing point of the video corresponding to a predetermined number of frames, in response to the confirmation that a predetermined number of frames containing coordinates with a number of intersections of the lines equal to or greater than a reference value within a predetermined coordinate region have been continuously identified. In this case, the average coordinate of the predetermined number of frames may be estimated as the vanishing point.

[0085] In the embodiments described above, a method of estimating a predetermined coordinate as the vanishing point was explained, but the invention is not limited to this, and a region indicated by a predetermined coordinate may be estimated as the region of the vanishing point.

[0086] In the above-described embodiment, a method for estimating the vanishing point based on feature points corresponding to an object common to two consecutive frames was explained. However, the method is not limited to this, and the vanishing point may be estimated based on three or more feature points corresponding to an object common to three or more consecutive frames. For example, the extraction unit 112 extracts a first feature point corresponding to the "head of streetlamp A" extracted from the first frame of the video, a second feature point corresponding to the "head of streetlamp A" extracted from the second frame, and a third feature point corresponding to the "head of streetlamp A" extracted from the third frame. The estimation unit 113 then estimates the vanishing point based on the first feature point, the second feature point, and the third feature point.

[0087] In the embodiments described above, a method for estimating vanishing points based on feature points corresponding to common objects between two consecutive frames was explained. However, the invention is not limited to this, and vanishing points may also be estimated based on feature points corresponding to common objects between two or more discontinuous frames. [Explanation of Symbols]

[0088] 1... Information processing system, 10... Information processing device, 11... Processor, 12... Memory device, 13... Communication interface, 14... Input device, 15... Output device, 20... In-vehicle device, 100... Storage unit, 101... Feature point DB, 110... Control unit, 111... Acquisition unit, 112... Extraction unit, 113... Estimation unit, 114... Generation unit, 115... Correction unit

Claims

1. An acquisition unit that acquires video footage captured by an in-vehicle camera, An extraction unit that extracts feature points from frames contained in the aforementioned video, An estimation unit estimates the vanishing point of the video based on a first feature point extracted from the first frame of the video and a second feature point corresponding to the first feature point, extracted from the second frame of the video. An information processing device equipped with the following features.

2. The estimation unit estimates the coordinates where the number of intersections of the lines is equal to or greater than a reference value when drawing a line connecting the first feature point and the second feature point with respect to the multiple feature points included in the video, as the vanishing points. The information processing apparatus according to claim 1.

3. The estimation unit estimates vanishing points using a Bayesian network model based on a straight line connecting the first feature point and the second feature point, with respect to a plurality of feature points included in the video. The information processing apparatus according to claim 1.

4. The estimation unit estimates the vanishing point of the video based on a series of consecutive frames in which, when a line is drawn connecting the first feature point and the second feature point with respect to a plurality of feature points included in the video, the number of intersections of the line is equal to or greater than a reference value, and the coordinates are included in a predetermined region. The information processing apparatus according to claim 1.

5. A generation unit generates mask data indicating the area to be masked in the video, based on the coordinates corresponding to the vanishing point and supplementary coordinates identified by the same coordinate system. The information processing apparatus according to claim 1, further comprising:

6. One or more computers, The acquisition step involves obtaining video footage captured by an in-car camera, and An extraction step to extract feature points from frames contained in the aforementioned video, An estimation step of estimating the vanishing point of the video based on a first feature point extracted from the first frame of the video and a second feature point corresponding to the first feature point extracted from the second frame of the video. An information processing method that performs [this action].

7. On one or more computers, The acquisition step involves obtaining video footage captured by an in-car camera, and An extraction step to extract feature points from frames contained in the aforementioned video, An estimation step of estimating the vanishing point of the video based on a first feature point extracted from the first frame of the video and a second feature point corresponding to the first feature point extracted from the second frame of the video. A program that executes something.