State estimation device, state estimation method, and state estimation program

The state estimation device uses machine learning to automate imaging device calibration, eliminating the need for manual input and specialized equipment, enhancing efficiency and deployment in traffic environments.

JP7850619B2Active Publication Date: 2026-04-23KYOCERA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
KYOCERA CORP
Filing Date
2022-07-21
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Conventional imaging device calibration requires a measurement vehicle and manual input of road features, which is laborious and inefficient.

Method used

A state estimation device using machine learning models to estimate the installation state of an imaging device by analyzing image data, eliminating the need for manual input and specialized jigs.

Benefits of technology

Enables automated calibration of imaging devices in traffic environments without human intervention, improving efficiency and enabling widespread adoption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850619000001
    Figure 0007850619000001
  • Figure 0007850619000002
    Figure 0007850619000002
  • Figure 0007850619000003
    Figure 0007850619000003
Patent Text Reader

Abstract

To enable estimation of an installation state of an imaging device which images a traffic environment without requiring human work and traffic regulation.SOLUTION: A state estimation device 100 comprises: a first state estimation unit which estimates first feature amount data from input image data; a second state estimation unit which estimates second feature amount data from the input image data; a feature estimation unit which estimates an installation state parameter of an imaging device that images the input image data from data obtained by combining the first feature amount data and the second feature amount data with a state estimation model that has performed machine-learning so as to estimate the installation state parameter of the imaging device that images the input image data by using third teacher data having the image data obtained by imaging a traffic environment by the imaging device and correct answer value data of the installation state parameter of the imaging device that images the image data; and a diagnostic unit which diagnoses the installation state of the imaging device on the basis of the estimated installation state parameter.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a state estimation device, a state estimation method, and a state estimation program.

Background Art

[0002] Cameras installed on roads, on the roadside of roads, etc. are known to perform calibration. Patent Document 1 discloses performing calibration using a measurement vehicle equipped with a GPS receiver, a data transmitter, a landmark, etc. Patent Document 2 discloses that in camera calibration, when the direction of a line existing on a road plane is input in a captured image, the road plane parameters are estimated based on the direction and the direction represented by an arithmetic expression including the road plane parameters.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In Patent Document 1, a measurement vehicle is required, and an operator is required when performing calibration. In Patent Document 2, it is necessary to manually input the lanes on the road into the image, and there is a problem that the work is laborious. For this reason, there is a need to estimate the installation state of an imaging device that captures the traffic environment without requiring human work or traffic control for a conventional imaging device that captures a road.

Means for Solving the Problems

[0005] A state estimation device according to one embodiment includes: a first state estimation unit trained to estimate first feature data from first image data including a moving object captured by an imaging device; a second state estimation unit trained to estimate second feature data from second image data including a road captured by the imaging device; and a feature estimation unit trained to estimate the installation state parameters of the imaging device that captured the input image data from the first feature data and the second feature data.

[0006] A state estimation device according to one embodiment includes a first object estimation model that has been machine-trained to estimate first feature data obtained by estimating the features of a first extractable object from input first image data, using first training data which includes first image data of a traffic environment captured by an imaging device and first ground truth data of a first extractable object including a moving object included in the first image data, and a first state estimation unit that estimates first feature data from input image data, and second training data which includes second image data of a traffic environment captured by an imaging device and second ground truth data of a second extractable object including a road included in the second image data, using second training data which estimates the features of a first extractable object from input first image data. The system includes: a second state estimation unit that estimates second feature data from input image data using a second object estimation model trained with machine learning to do so; a feature estimation unit that estimates the installation state parameters of the imaging device that captured the input image data using a third training data which has image data of the imaging device capturing the traffic environment and ground truth data of the installation state parameters of the imaging device that captured the image data, using data synthesized from the first feature data and the second feature data; and a diagnostic unit that diagnoses the installation state of the imaging device based on the estimated installation state parameters.

[0007] A state estimation method according to one embodiment is a first object estimation model in which a computer has been trained to estimate first feature data obtained by estimating the features of a first extractable object from input first image data, using first training data which has first image data of a traffic environment captured by an imaging device and first ground truth data of a first extractable object including a moving object contained in the first image data, and second feature data obtained by estimating the features of a first extractable object from input first image data, using second training data which has second image data of a traffic environment captured by an imaging device and second ground truth data of a second extractable object including a road contained in the second image data, and This includes: using a second object estimation model trained with machine learning to estimate quantitative data to estimate second feature data from the input image data; using a third training data consisting of image data of the imaging device capturing the traffic environment and ground truth data of the installation state parameters of the imaging device that captured the input image data to estimate the installation state parameters of the imaging device that captured the input image data from data obtained by combining the first feature data and the second feature data; and diagnosing the installation state of the imaging device based on the estimated installation state parameters.

[0008] A state estimation program according to one embodiment is a first object estimation model that has been trained by a computer to estimate first feature data obtained by estimating the features of a first extractable object from the input first image data, using first training data which has first image data of a traffic environment captured by an imaging device and first ground truth data of a first extractable object including a moving object included in the first image data, and second feature data obtained by estimating the features of a first extractable object from the input first image data, using second training data which has second image data of a traffic environment captured by an imaging device and second ground truth data of a second extractable object including a road included in the second image data, and A second object estimation model, trained using machine learning to estimate quantitative data, is used to estimate second feature data from the input image data. A state estimation model, trained using machine learning to estimate the installation state parameters of the imaging device that captured the input image data, is used to estimate the installation state parameters of the imaging device that captured the input image data, using third training data which includes image data of the imaging device capturing the traffic environment and ground truth data of the installation state parameters of the imaging device that captured the image data. This state estimation model is used to estimate the installation state parameters of the imaging device that captured the input image data from data obtained by combining the first feature data and the second feature data. The installation state of the imaging device is then diagnosed based on the estimated installation state parameters. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 is a diagram illustrating an example of the relationship between a learning device and a state estimation device according to an embodiment. [Figure 2] Figure 2 shows an example of image data captured by the imaging device shown in Figure 1. [Figure 3] Figure 3 shows an example of the configuration of a learning device according to this embodiment. [Figure 4] Figure 4 shows an example of nighttime image data. [Figure 5] Figure 5 shows an example of image data taken in the early morning. [Figure 6]Figure 6 shows an example of a CNN used for state estimation by the learning device shown in Figure 3. [Figure 7] Figure 7 shows an example of a CNN used for object detection by the learning device shown in Figure 3. [Figure 8] Figure 8 shows an example of the configuration of a state estimation device according to an embodiment. [Figure 9] Figure 9 shows an example of the configuration of the control unit of the state estimation device according to the embodiment. [Figure 10] Figure 10 is a flowchart showing an example of a state estimation method performed by a state estimation device. [Figure 11] Figure 11 is a flowchart showing an example of a state estimation method performed by the first processing unit. [Figure 12] Figure 12 is a flowchart showing an example of a state estimation method performed by the second processing unit. [Modes for carrying out the invention]

[0010] Multiple embodiments for implementing the state estimation device, state estimation method, and state estimation program according to this application will be described in detail with reference to the drawings. However, the present invention is not limited by the following description. Furthermore, the components described below include those easily conceivable by those skilled in the art, those substantially identical, and those within the so-called equivalent range. Similar components may be denoted by the same reference numerals in the following description. Furthermore, redundant descriptions may be omitted.

[0011] (System Overview) In conventional systems, linking captured images with the real world using information about the installation status of the imaging device required specialized jigs and procedures. Furthermore, conventional systems required road closure procedures because the imaging device was installed near roads. The state estimation device according to this embodiment eliminates the need for jig-based procedures and road closure procedures, contributing to the widespread adoption of the imaging device 10 in traffic environments.

[0012] Figure 1 is a diagram illustrating an example of the relationship between a learning device and a state estimation device according to an embodiment. Figure 2 is a diagram showing an example of image data captured by the imaging device shown in Figure 1. As shown in Figure 1, System 1 includes an imaging device 10 and a state estimation device 100. The imaging device 10 can acquire image data D10 of the traffic environment 1000. The state estimation device 100 has the function of acquiring image data D10 from the imaging device 10 and estimating the installation state of the imaging device 10 based on the image data D10. The imaging device 10 and the state estimation device 100 are configured to communicate by wire or wireless. In the example shown in Figure 1, for the sake of simplicity, the case where System 1 has one imaging device 10 and one state estimation device 100 is described, but the number of imaging devices 10 and state estimation devices 100 may be multiple.

[0013] The imaging device 10 is installed so as to be able to image a traffic environment 1000 including a road 1100 and traffic objects 1200 moving on the road 1100. The traffic objects 1200 moving on the road 1100 include, for example, vehicles, people, etc. that can move on the road 1100. The traffic objects 1200 include, for example, large automobiles, ordinary automobiles, large special automobiles, large two-wheelers, ordinary motorcycles, small special automobiles, etc. defined by the Road Traffic Law, but may also include other vehicles, moving bodies, etc. Note that large automobiles include those with a total vehicle weight of 8000 kg or more, a maximum loading capacity of 5000 kg or more, a seating capacity of 11 or more people (buses, trucks, etc.), etc. The imaging device 10 can electronically image an image, for example, using an image sensor such as a CCD (Charge Coupled Device Image Sensor) or a CMOS (Complementary Metal Oxide Semiconductor). The imaging direction of the imaging device 10 is installed in a state facing the road plane of the traffic environment 1000. The imaging device 10 can be installed, for example, on a road, an intersection, a parking lot, etc. The road 1100 imaged by the imaging device 10 may include shapes such as straight lines, curves, and gradients of the road, signs installed on the road, the shape of the median strip, lines, marks, and signs drawn on the road, guardrails, streetlights, trees, sidewalks, destination guide plates, advertisements, and phosphors that pay attention to road shapes such as curves.

[0014] In an example shown in FIG. 1, the imaging device 10 is installed on the roadside at an installation angle that can image the imaging area of the traffic environment 1000 including the road 1100 and its surroundings from above. The imaging device 10 acquires image data D10 by imaging the traffic environment 1000. The imaging device 10 may be provided so that the imaging direction is fixed, or may be provided so that the imaging direction can be changed by a movable mechanism at the same position. The image data D10 of the imaging device 10 is data indicating an image D11 having a first area D110 showing a plurality of roads 1100 and a second area D120 showing traffic objects 1200 passing through the road 1100, as shown in FIG. 2.

[0015] The imaging device 10 supplies the captured image data to the state estimation device 100. In the present embodiment, the image data includes, for example, two-dimensional images such as moving images and still images. The imaging device 10 of the present embodiment performs imaging at night and during the day, captures various image data, and supplies it to the state estimation device 100. A predetermined region D100 is set in advance for the image data D10 with respect to the image D11. The predetermined region D100 is a region that includes traffic objects 1200 that can be used for estimation, and can be set as appropriate based on the traffic environment 1000 to be imaged. The predetermined region D100 may be all regions of the image D11. The traffic objects 1200 that can be used for estimation include, for example, traffic objects 1200 used as correct values for machine learning of the state estimation model M1. The traffic objects 1200 that can be used for estimation are traffic objects 1200 suitable for the estimation of the state estimation model M1.

[0016] As shown in FIG. 1, the state estimation device 100 may be provided near the imaging device 10 or at a position away from the imaging device 10. In an example shown in FIG. 1, for the sake of simplicity of explanation, the case where the state estimation device 100 is supplied with the image data D10 from one imaging device 10 will be described. However, for example, the image data D10 may be supplied from each of a plurality of imaging devices 10. The traffic object 1200 is moving along the road direction C1 toward the imaging device 10 in the lane, and the road direction C2 indicates the direction of the oncoming lane.

[0017] The state estimation device 100 has a function to manage the installation state parameters of the imaging device 10. The installation state parameters include, for example, the installation angle and installation position of the imaging device 10. The installation state parameters may also include, for example, the number of pixels of the imaging device 10 and the size of the image D11. The state estimation device 100 can estimate the installation state parameters of the imaging device 10 that captured the image data D10 using the first state estimation model M1, the second state estimation model M2, and the feature estimation model M3, which have been machine-learned by the learning device 200. The state estimation device 100 identifies whether the multiple image data D10 were taken at night or during the day (early morning), inputs the nighttime image data into the first state estimation model M1, inputs the daytime image data into the second state estimation model M2, outputs the feature quantities of each image data, combines the first feature quantity data calculated using the first state estimation model M1 and the second feature quantity data calculated using the second state estimation model M2, inputs them into the feature estimation model M3, and estimates the output of the feature estimation model M3 as the installation state parameters of the imaging device 10.

[0018] The learning device 200 is, for example, a computer, a server, etc. The learning device 200 may or may not be included in the configuration of System 1. The learning device 200 acquires a plurality of first training data, each having image data D10 of a traffic environment 1000 including traffic objects 1200, and correct value data D21 of the installation state parameters of the imaging device 10 that captured the image data D10. The correct value data D21 includes, for example, data indicating correct values ​​such as the installation angle (α, β, γ) of the imaging device 10, the installation position (x, y, z), the number of pixels, and the size of the image D11. The correct value data D21 is an example of first correct value data. The installation angle includes, for example, the pitch angle α in the direction in which the imaging device 10 is tilted downwards, the yaw angle β that allows the imaging device 10 to swing the imaging direction laterally, and the roll angle γ in the direction in which the imaging device 10 is tilted. The installation location has, for example, a position (x,z) on the road surface and a height y. The ground truth data D21 may be, for example, a ground truth value combining two values ​​α and γ that can identify the orientation relative to the road surface. The ground truth data D21 may be, for example, a ground truth value combining three values ​​α, γ, and y that can identify the scale. The ground truth data D21 may be, for example, a ground truth value combining four values ​​α, β, γ, and y that can identify the main road direction. The ground truth data D21 may be, for example, a ground truth value combining six values ​​α, β, γ, x, y, and z that are used in general calibration.

[0019] The learning device 200 generates a first state estimation model M1, a second state estimation model M2, and a feature estimation model M3 using machine learning based on a combination of multiple image data captured from the same location and installation state parameters. The training data image data is classified into nighttime image data and daytime image data. The nighttime image data contains information that identifies the area of ​​traffic objects. The daytime image data is an image in which traffic objects are below a threshold and includes roads and fixed objects placed around the roads. The training data includes nighttime image data containing traffic objects as the first training data and daytime image data as the second training data.

[0020] The learning device 200 generates a first state estimation model M1 that estimates first feature data from input image data D10 using machine learning with multiple first training data sets. The learning device 200 generates a second state estimation model M2 that estimates second feature data from input image data D10 using machine learning with multiple second training data sets. The learning device 200 generates a feature estimation model M3 that estimates the installation state parameters of the imaging device 10 that has captured images, using machine learning with data that is a combination of the first and second feature data sets, from the first state estimation model M1 and the second state estimation model M2.

[0021] Supervised machine learning can utilize algorithms such as neural networks, linear regression, and logistic regression. The first state estimation model M1 is a model that has been trained using image data from multiple training datasets and ground truth data D21, which are feature data of traffic objects contained in the image data, to estimate the first feature from the input image data D10. When image data is input to the first state estimation model M1, it estimates and outputs the first feature data, which includes the feature data of traffic objects contained in the image data. The second state estimation model M2 is a model that has been trained using image data from multiple training datasets and ground truth data, which are feature data of traffic environment other than traffic objects such as roads contained in the image data, to estimate the second feature from the input image data D10. When image data is input to the second state estimation model M2, it estimates and outputs the second feature data, which includes the feature data of traffic environment other than traffic objects contained in the image data. The third feature estimation model M3 is a machine learning model that uses feature data from multiple training datasets and ground truth data D22, which is data on the installation state parameters of the imaging device 10 that captured the image data D10, to estimate the installation state parameters of the imaging device 10 that captured the image data D10 using data that is a composite of the input first feature data and second feature data. When data that is a composite of the first feature data and second feature data is input to the feature estimation model M3, it estimates the installation state parameters of the imaging device 10 that captured the image data D10 and outputs the estimation result. The learning device 200 contributes to eliminating the need for dedicated jigs and manual work in the calculation of the installation state of the imaging device 10 by providing the generated state estimation models M1, M2, and feature estimation model M3 to the state estimation device 100. An example of the learning device 200 will be described later.

[0022] The state estimation device 100 inputs image data D into the first state estimation model M1 and the second state estimation model M2 provided by the learning device 200, outputs first feature data and second feature data from the first state estimation model M1 and the second state estimation model M2, inputs feature data obtained by combining the first feature data and the second feature data into the feature estimation model M3, and can estimate the installation state parameters for the image data D10 based on the output of the feature estimation model M3. The state estimation device 100 can diagnose the installation state of the imaging device 10 based on the estimated installation state parameters. As a result, the state estimation device 100 can eliminate the need for special jigs and manual work to calculate the installation state of the imaging device 10 when installing or maintaining the imaging device 10 in relation to the traffic environment 1000. By eliminating the need for jigs and manual work, the state estimation device 100 can eliminate the need for traffic restrictions. As a result, the state estimation device 100 can contribute to the widespread adoption of imaging devices 10 installed in traffic environments 1000, and can also improve the efficiency of maintenance.

[0023] The learning device 200 can acquire multiple third training data sets, each containing image data D10 of the traffic environment 1000 including traffic objects 1200 captured by the imaging device 10, and ground truth data D22 for object detection in the image data D10. The ground truth data D22 includes, for example, data showing the ground truth values ​​of the traffic objects 1200 in image D11, such as position, size, type, and number of objects. The ground truth data D22 is an example of second ground truth data. The ground truth data D22 includes, for example, a total of five data points for each object in image D11: two data points for the object's position (x,y), two data points for the object's size (w,h), and one data point for the object's type. The object types include, for example, people, large vehicles, regular vehicles, large special vehicles, large motorcycles, regular motorcycles, small special vehicles, bicycles, etc.

[0024] The learning device 200 generates an object estimation model M4 that estimates at least one of the location, size, and type of traffic objects 1200 in the traffic environment 1000 indicated by the input image data D10 using machine learning with multiple third-party training data. The object estimation model M4 is a model that has been machine-learned using multiple training data image data D10 and ground truth data D22 to estimate the location, size, and type of traffic objects 1200 in the traffic environment 1000 indicated by the input image data D10. When the image data D10 is input, the object estimation model M2 estimates the location, size, type, and number of traffic objects 1200 in the traffic environment 1000 indicated by the image data D10, and outputs the estimation result. The learning device 200 can provide the generated object estimation model M4 to the state estimation device 100.

[0025] The state estimation device 100 has a function to process image data D10 captured by the imaging device 10 of the traffic environment 1000 so that it includes traffic objects 1200 used for estimation. The state estimation device 100 can, for example, use an object estimation model M4 to process image data D10 captured by the imaging device 10 of the traffic environment 1000 so that it includes traffic objects 1200 used for estimation. As a result, the state estimation device 100 can input image data D10 that can be used for estimating the installation state parameters of the imaging device 10 into the state estimation model M1, thereby improving the estimation accuracy of the state estimation model M1. Image data D10 that can be used for estimating the installation state parameters of the imaging device 10 is data that can improve the probability of the estimation result of the state estimation model M1.

[0026] System 1 can provide a function to manage the maintenance of one or more imaging devices 10 using the estimation results of the state estimation device 100. System 1 can also provide a function to instruct a change in the installation state of the imaging device 10 based on the installation state parameters and installation location estimated by the state estimation device 100.

[0027] (Learning device) Figure 3 shows an example of the configuration of the learning device 200 according to the embodiment. As shown in Figure 3, the learning device 200 includes a display unit 210, an operation unit 220, a communication unit 230, a storage unit 240, and a control unit 250. The control unit 250 is electrically connected to the display unit 210, the operation unit 220, the communication unit 230, the storage unit 240, etc. In this embodiment, an example of the learning device 200 performing machine learning using a CNN (Convolutional Neural Network), which is a type of neural network, will be described.

[0028] The display unit 210 is configured to display various types of information under the control of the control unit 250. The display unit 210 has a display panel such as a liquid crystal display or an organic EL display. The display unit 210 displays information such as characters, graphics, and images in response to signals input from the control unit 250.

[0029] The operation unit 220 has one or more devices for receiving user input. These devices include, for example, keys, buttons, touchscreens, mice, etc. The operation unit 220 can supply signals to the control unit 250 corresponding to the received input.

[0030] The communication unit 230 can communicate with, for example, the state estimation device 100, other communication devices, etc. The communication unit 230 can support various communication standards. The communication unit 230 can send and receive various types of data, for example, via wired or wireless networks. The communication unit 230 can supply received data to the control unit 250. The communication unit 230 can send data to a destination instructed by the control unit 250.

[0031] The storage unit 240 can store programs and data. The storage unit 240 is also used as a work area for temporarily storing the processing results of the control unit 250. The storage unit 240 may include any non-transitory storage medium, such as a semiconductor storage medium and a magnetic storage medium. The storage unit 240 may include multiple types of storage media. The storage unit 240 may include a combination of a portable storage medium, such as a memory card, optical disk, or magneto-optical disk, and a storage medium reader. The storage unit 240 may include a storage device used as a temporary storage area, such as RAM (Random Access Memory).

[0032] The memory unit 240 can store various data, such as the program 241, training data 242, the first state estimation model M1, the second state estimation model M2, the feature estimation model M3, and the object estimation model M4. The program 241 causes the control unit 250 to execute a function that generates a state estimation model using a CNN to estimate the installation state parameters of the imaging device 10 that captured the image data D10. The program 241 also causes the control unit 250 to execute a function that generates an object estimation model using a CNN to estimate information about the object represented by the image data D10.

[0033] The training data 242 consists of training data and other data used in machine learning. The training data 242 includes data that combines image data D10 used in machine learning for state estimation and ground truth data D21 associated with the image data D10. Image data D10 is input data for supervised learning. For example, image data D10 shows a color image of a traffic environment 1000 including traffic objects 1200, with a pixel count of 1280 × 960. The ground truth data D21 includes data that shows the installation state parameters of the imaging device 10 that captured the image data D10. The ground truth data D21 is ground truth data for supervised machine learning. For example, the ground truth data D21 includes data that shows six parameters (values) of the installation angle (α, β, γ) and installation position (x, y, z) of the imaging device 10.

[0034] The training data 242 further includes data that combines image data D10 used for machine learning object estimation and ground truth data D22 associated with the image data D10. For example, image data D10 represents a color image of a traffic environment 1000 including traffic objects 1200, with a pixel count of 1280 × 960. The ground truth data D22 includes data indicating the object position, object size, object type, etc., of the objects (traffic objects 1200) shown in the image data D10, for each traffic object 1200 (object) included in the image. The object position includes, for example, the coordinates (x, y) in the associated image data D10. The object size includes, for example, the width, height, etc., of the object shown in the associated image data D10.

[0035] Image data D10 includes images taken at night and in the early morning. Figure 4 shows an example of nighttime image data. Figure 5 shows an example of early morning image data. As shown in Figure 4, the image data taken at night clearly shows the location of lighting devices such as the headlights of vehicles, which are traffic objects, and it is difficult to identify traffic environment other than traffic objects that are not emitting light, such as road lanes. As shown in Figure 5, the image data taken in the early morning has low traffic volume and few traffic objects, and it is easy to identify traffic environment other than traffic objects.

[0036] The first state estimation model M1 is a learning model generated by extracting features, regularities, patterns, etc. from the image data D10 of the training data 242, using nighttime image data (first image data) and ground truth data D21, and performing machine learning on the relationship between the image and the features corresponding to the ground truth data D21. When image data D10 is input to the first state estimation model M1, it predicts training data 242 that is similar to the features of the image data D10, estimates the first feature data, and outputs it. The second state estimation model M2 is a learning model generated by extracting features, regularities, patterns, etc. from the image data D10 of the training data 242, using daytime image data (second image data) and ground truth data D21, and performing machine learning on the relationship between the image and the features corresponding to the ground truth data D21. When image data D10 is input to the second state estimation model M2, it predicts training data 242 that is similar to the features of the image data D10, estimates the second feature data, and outputs it. Here, daytime image data (second image data) is an image with fewer moving objects compared to nighttime image data (first image data). The third feature estimation model M3 is a learning model generated by using image data D10 and ground truth data D21 from the training data 242 to extract features, regularities, patterns, etc. from image data D10, and by machine learning the relationship between the image features and the ground truth data D21. When data synthesized from the first feature data and the second feature data is input to the third feature estimation model M3, it predicts training data 242 that are similar to the features of the image data D10, estimates the installation state parameters of the imaging device 10 that captured the image data D10, and outputs the estimation result.

[0037] The object estimation model M4 is a learning model generated by using image data D10 and ground truth data D22 from the training data 242 to extract the features, regularities, patterns, etc. of objects in image data D10, and then learning the relationship with the ground truth data D22. When image data D10 is input to the object estimation model M4, it predicts training data 242 that are similar to the features, etc. of objects in the image data D10, estimates the position, size, type, etc. of objects in the image shown by image data D10 based on the ground truth data D22, and outputs the estimation result.

[0038] The control unit 250 is an arithmetic processing unit. The arithmetic processing unit includes, but is not limited to, a CPU (Central Processing Unit), SoC (System-on-a-Chip), MCU (Micro Control Unit), FPGA (Field-Programmable Gate Array), and coprocessor. The control unit 250 can comprehensively control the operation of the learning device 200 to realize various functions.

[0039] Specifically, the control unit 250 can execute instructions included in the program 241 stored in the storage unit 240, while referring to the information stored in the storage unit 240 as needed. The control unit 250 then controls the functional units according to the data and instructions, thereby realizing various functions. The functional units include, for example, the display unit 210 and the communication unit 230, but are not limited to these.

[0040] The control unit 250 has functional units such as the first acquisition unit 251, the first machine learning unit 252, the second acquisition unit 253, the second machine learning unit 254, the third acquisition unit 255, the third machine learning unit 256, the fourth acquisition unit 257, and the fourth machine learning unit 258. The control unit 250 realizes the functions of the first acquisition unit 251, the first machine learning unit 252, the second acquisition unit 253, the second machine learning unit 254, the third acquisition unit 255, the third machine learning unit 256, the fourth acquisition unit 257, and the fourth machine learning unit 258 by executing the program 241. The program 241 is a program that causes the control unit 250 of the learning device 200 to function as the first acquisition unit 251, the first machine learning unit 252, the second acquisition unit 253, the second machine learning unit 254, the third acquisition unit 255, the third machine learning unit 256, the fourth acquisition unit 257, and the fourth machine learning unit 258.

[0041] The first acquisition unit 251 acquires image data D10, which is an image of the traffic environment 1000 including traffic objects 1200, and feature data corresponding to the ground truth value data D21 of the installation state parameters of the imaging device 10 that captured the image data D10, as training data. The first acquisition unit 251 acquires nighttime image data from the image data. The first acquisition unit 251 acquires feature data corresponding to the image data D10 and the ground truth value data D21 from a pre-set storage location, a storage location selected by the operation unit 220, etc., and stores them in association with the training data 242 in the storage unit 240. The first acquisition unit 251 acquires feature data corresponding to multiple image data D10 and ground truth value data D21 to be used for machine learning.

[0042] The first machine learning unit 252 generates a first state estimation model M1 that estimates the features of image data (first image data) using machine learning with multiple training data 242 (first training data) acquired by the first acquisition unit 251. For example, the first machine learning unit 252 constructs a CNN based on the training data 242. The CNN is constructed to take image data D10 as input and output a classification result for image data D10. The classification result is feature data that includes the features of traffic objects contained in image data D10.

[0043] The second acquisition unit 253 acquires image data D10 of the traffic environment 1000 and feature data corresponding to the ground truth value data D21 of the installation state parameters of the imaging device 10 that captured the image data D10, as training data. The second acquisition unit 253 acquires image data from the image data, particularly from the early morning. The second acquisition unit 253 acquires feature data corresponding to the image data D10 and the ground truth value data D21 from a preset storage location, a storage location selected by the operation unit 220, etc., and stores them in association with the training data 242 in the storage unit 240. The second acquisition unit 253 acquires feature data corresponding to multiple image data D10 and ground truth value data D21 to be used for machine learning.

[0044] The second machine learning unit 254 generates a second state estimation model M1 that estimates the features of image data (second image data) using machine learning with multiple training data 242 (second training data) acquired by the second acquisition unit 253. The second machine learning unit 252 constructs a CNN, for example, based on the training data 242. The CNN is constructed to take image data D10 as input and output a classification result for image data D10. The classification result is feature data that includes features of fixed objects in the image data D10 other than traffic objects, such as roads, signs, traffic lights, and road shoulders.

[0045] The third acquisition unit 255 acquires feature data acquired by the first machine learning unit 252 and the second machine learning unit 254 based on image data D10 of the traffic environment 1000 including traffic objects 1200, and acquires feature data corresponding to the ground truth value data D21 of the installation state parameters of the imaging device 10 that captured the image data D10, as training data. The third acquisition unit 255 acquires feature data and ground truth value data D21 from a preset storage location, a storage location selected by the operation unit 220, etc., and stores them in association with the training data 242 in the storage unit 240. The third acquisition unit 255 acquires multiple feature data and ground truth value data D21 to be used for machine learning.

[0046] The third machine learning unit 256 generates a feature estimation model M3 that estimates the installation state parameters of the imaging device 10 that captured the input image data D10, using machine learning with multiple training data 242 (feature data) acquired by the third acquisition unit 255. For example, the third machine learning unit 256 constructs a CNN based on the training data 242. The CNN is constructed to take feature data as input and output an identification result for the image data D10. The identification result contains information for estimating the installation state parameters of the imaging device 10 that captured the image data D10.

[0047] Figure 6 shows an example of a CNN used for state estimation by the learning device 200 shown in Figure 3. The first machine learning unit 252, the second machine learning unit 254, and the third machine learning unit 256 construct the CNN shown in Figure 6 based on the acquired training data 242. In this embodiment, machine learning is performed in the first machine learning unit 252, the second machine learning unit 254, and the third machine learning unit 256, and the learning results of the first machine learning unit 252 and the second machine learning unit 254 are supplied to the third machine learning unit 256. Note that these may be performed as separate processes, or the learning in the first machine learning unit 252, the second machine learning unit 254, and the third machine learning unit 256 may be performed as a single learning process. Feature data that will serve as the ground truth data for the first machine learning unit 252 and the second machine learning unit 254 can be generated by various methods. As is well known, the CNN has an input layer, an intermediate layer, and an output layer.

[0048] As shown in Figure 6, the learning device 200 includes a first learning unit 400, a second learning unit 410, a third learning unit 420, and output layers 430, 440, and 450. The first learning unit 400 executes the processing of the first machine learning unit 252. The first learning unit 400 includes an input layer 500 and an intermediate layer 510, and outputs the results processed in the intermediate layer 510 to the output layer 430. The input layer 500 receives image data taken at night from among the image data. The input layer 500 outputs the input data to the intermediate layer 510. The input image data D10 is, for example, data representing a 640 × 640 × 3 color image.

[0049] The intermediate layer 510 has a plurality of feature extraction layers and a coupling layer. Each of the plurality of feature extraction layers extracts different features of image D11 represented by image data D10. The features of image D11 to be extracted include, for example, features related to traffic objects in the image. The feature extraction layer has, for example, one or more convolution layers and a pooling layer, and extracts desired features from the input image data D10. The convolution layer of the feature extraction layer is a layer that extracts parts of image D11 that are similar in shape to the filter (weights) by performing a convolution operation on the input data. The convolution layer is configured to apply an activation function to the feature map which is the result of the operation. In this embodiment, the ReLU (Rectified linear unit) function is applied as the activation function, but a sigmoid function or the like may also be applied. The pooling layer of the feature extraction layer summarizes the features of image data D10 obtained by convolution into a maximum value or average value, so that even if the position of the extracted features changes, they are considered to be the same feature. The feature extraction layer 2210 can extract more advanced and complex features by increasing the number of convolutional and pooling layers in order to learn the optimal output. The combined layer combines the features extracted by multiple feature extraction layers and outputs them to the output layer 430. The hidden layer 510 outputs data representing the feature quantities to the output layer 430.

[0050] The second learning unit 410 executes the processing of the second machine learning unit 254. The second learning unit 410 includes an input layer 520 and an intermediate layer 530, and outputs the results processed in the intermediate layer 530 to the output layer 440. The input layer 520 receives image data taken in the early morning and during the day. The input layer 520 outputs the input data to the intermediate layer 530. The input image data D10 is, for example, data representing a 640×640×3 color image.

[0051] The intermediate layer 530 has multiple feature extraction layers and a coupling layer. Each of the multiple feature extraction layers extracts different features of image D11 represented by image data D10. The features of image D11 to be extracted include, for example, features related to the traffic environment other than traffic objects in the image. The feature extraction layer has, for example, one or more convolution layers and a pooling layer, and extracts desired features from the input image data D10. The convolution layer of the feature extraction layer is a layer that extracts parts of image D11 that are similar in shape to the filter (weights) by performing a convolution operation on the input data. The convolution layer is configured to apply an activation function to the feature map which is the result of the operation. In this embodiment, the ReLU (Rectified linear unit) function is applied as the activation function, but a sigmoid function or the like may also be applied. The pooling layer of the feature extraction layer summarizes the features of the image data D10 obtained by convolution into maximum and average values, so that even if the position of the extracted features changes, they are considered to be the same feature. The feature extraction layer 2210 can extract more advanced and complex features by increasing the number of convolutional and pooling layers in order to learn the desired optimal output. The combined layer combines the features extracted by multiple feature extraction layers and outputs them to the output layer 440. The hidden layer 530 outputs data representing the feature quantities to the output layer 440. The second learning unit 410 outputs data in the same data format as the first learning unit 400, that is, data with the same number of pixels in the feature quantity data.

[0052] The learning device 200 combines the output layer 430 and the feature data output by the output layer 440 in the synthesis unit 540. The synthesis unit 540 selects one image's feature data from the feature data of multiple images output by the output layer 430, selects one image's feature data from the feature data of multiple images output by the output layer 440, and synthesizes them to generate one feature data. The synthesis unit 540 performs synthesis processing for the number of images of feature data output by the output layer 430 and the output layer 440, generating a predetermined number of image feature data. The method of image selection by the synthesis unit 540 is not particularly limited. Also, one image data may be used multiple times.

[0053] The third learning unit 420 executes the processing of the third machine learning unit 256. The third learning unit 420 includes an intermediate layer 530 and outputs the results processed in the intermediate layer 550 to the output layer 450. The intermediate layer 530 is supplied with feature data synthesized in the synthesis unit 540.

[0054] The intermediate layer 530 has multiple feature extraction layers and a coupling layer. Each of the multiple feature extraction layers extracts distinct features of image D11 represented by image data D10 included in the feature data. The feature extraction layer has, for example, one or more convolution layers and extracts desired features from the input image data D10. The convolution layer of the feature extraction layer is a layer that extracts parts of image D11 that are similar in shape to the filter (weights) by performing a convolution operation on the input data. The convolution layer is configured to apply an activation function to the feature map which is the result of the operation. In this embodiment, the ReLU (Rectified linear unit) function is applied as the activation function, but a sigmoid function or the like may also be applied. The pooling layer of the feature extraction layer summarizes the features of the image data D10 obtained by convolution into a maximum value or average value, so that even if the position of the extracted features changes, they are considered to be the same feature. The feature extraction layer 2210 can extract more advanced and complex features by increasing the number of convolutional and pooling layers in order to learn the optimal output. The combined layer combines the features extracted by multiple feature extraction layers and outputs them to the output layer 450.

[0055] The output layer 450 estimates the installation status parameters of the imaging device 10 that captured the image data D10, based on the features extracted by the intermediate layer 550 and the ground truth data D21. The output layer 450 identifies the ground truth data D21 associated with features similar to those output by the combined layer and outputs the installation status parameters indicated by the ground truth data D21.

[0056] The first machine learning unit 252, the second machine learning unit 254, and the third machine learning unit 256 perform machine learning using multiple training data 242 to determine the weights of the intermediate layers and set them up in a CNN to estimate the installation state parameters of the imaging device 10 that captured the input image data D10, thereby generating the first state estimation model M1, the second state estimation model M2, and the feature estimation model M3. The first machine learning unit 252, the second machine learning unit 254, and the third machine learning unit 256 store the generated first state estimation model M1, second state estimation model M2, and feature estimation model M3 in the storage unit 240. As a result, when image data D10 is input, the first state estimation model M1 can output feature data of the image data. When image data D10 is input, the second state estimation model M2 can output feature data of the image data. When data with synthesized feature data is input, the feature estimation model M3 can output the result of estimating the installation state parameters of the imaging device 10 that captured the image data D10.

[0057] The fourth acquisition unit 257 shown in Figure 3 acquires image data D10 captured by the imaging device 10 of the traffic environment 1000 including traffic objects 1200, and the correct value data D22 for object detection in the image data D10, as training data 242. The second acquisition unit 253 acquires the image data D10 and the correct value data D22 from a preset storage location, a storage location selected by the operation unit 220, etc., and stores them in association with the training data 242 in the storage unit 240. The fourth acquisition unit 257 acquires multiple image data D10 and correct value data D22 to be used for machine learning.

[0058] The fourth machine learning unit 258 generates an object estimation model M2 that estimates at least one of the location, size, and type of traffic object 1200 in the traffic environment 1000 indicated by the input image data D10 using machine learning with the training data 242 (fourth training data). The fourth machine learning unit 258 constructs a CNN that corresponds to the detection of traffic object 1200 (object) based on the training data 242. The CNN is constructed to take the image data D10 as input and output estimation results that estimate the location, size, and type of traffic object 1200 in the traffic environment 1000 indicated by the image data D10. The identification result has the location, size, and type of traffic object 1200 indicated by the image data D10.

[0059] The fourth machine learning unit 258 constructs the CNN shown in Figure 7 based on the acquired training data 242. The CNN has an input layer 2100, an intermediate layer 2200, and an output layer 2300. The input layer 2100 supplies the input image data D10 to the intermediate layer 2200. The input image data D10 is, for example, data representing a 640×640×3 color image. The intermediate layer 2200 has multiple feature extraction layers 2210 and a combined layer 2220. The feature extraction layer 2210 extracts traffic objects 1200 (features) from the image data D10. The feature extraction layer 2210 has, for example, multiple convolutional layers and a pooling layer, and extracts traffic objects 1200 as features from the input image data D10. The convolutional layer of the feature extraction layer 2210 extracts parts of image D11 that resemble the shape of the filter (weights) by performing a convolution operation on the input data. The convolutional layer is configured to apply an activation function to the feature map, which is the result of the operation. The pooling layer of the feature extraction layer 2210 summarizes the features of the image data D10 obtained by convolution into maximum and average values, so that even if the position of the extracted features changes, they are considered to be the same feature. To learn the optimal output, the feature extraction layer 2210 can extract more advanced and complex features by increasing the number of convolutional and pooling layers. The coupling layer 2220 combines the features extracted by multiple feature extraction layers 2210 and outputs them to the output layer 2300.

[0060] The output layer 2300 estimates the traffic objects 1200 in the image shown by the image data D10 based on the features extracted by the intermediate layer 2200 and the ground truth data D22. The output layer 2300 identifies the ground truth data D22 associated with features similar to those output by the fully connected layer 2220, and outputs the location, size, and type of the traffic objects 1200 shown by the ground truth data D22, as well as the estimated number of traffic objects 1200.

[0061] The fourth machine learning unit 258 uses multiple training data 242 (fourth training data) acquired by the fourth acquisition unit 257 to perform machine learning, determining the weights of the intermediate layer 2200 and setting it up as a CNN to estimate the traffic objects 1200 in image D11 shown by the input image data D10, thereby generating an object estimation model M2. The fourth machine learning unit 258 stores the generated object estimation model M2 in the storage unit 240. As a result, when image data D10 is input, the object estimation model M4 can output the location, size, and type of the traffic objects 1200 shown in the image data D10, for the number of estimated traffic objects 1200.

[0062] The above describes an example of the functional configuration of the learning device 200 according to this embodiment. Note that the above configuration described using Figure 3 is merely an example, and the functional configuration of the learning device 200 according to this embodiment is not limited to this example. The functional configuration of the learning device 200 according to this embodiment can be flexibly modified according to specifications and operation.

[0063] (State estimation device) Figure 8 shows an example of the configuration of the state estimation device 100 according to the embodiment. Figure 9 shows an example of the configuration of the control unit of the state estimation device according to the embodiment. As shown in Figure 8, the state estimation device 100 comprises an input unit 110, a communication unit 120, a storage unit 130, and a control unit 140. The control unit 140 is electrically connected to the input unit 110, the communication unit 120, the storage unit 130, etc.

[0064] The input unit 110 receives image data D10 captured by the imaging device 10. The input unit 110 has a connector that can be electrically connected to the imaging device 10, for example, via a cable. The input unit 110 supplies the image data D10 input from the imaging device 10 to the control unit 140.

[0065] The communication unit 120 can communicate with, for example, a learning device 200, a management device that manages the imaging device 10, etc. The communication unit 120 can support various communication standards. The communication unit 120 can send and receive various types of information, for example, via a wired or wireless network. The communication unit 120 can supply received data to the control unit 140. The communication unit 120 can send data to a destination instructed by the control unit 140.

[0066] The storage unit 130 can store programs and data. The storage unit 130 is also used as a work area for temporarily storing the processing results of the control unit 140. The storage unit 130 may include any non-transient storage medium such as a semiconductor storage medium and a magnetic storage medium. The storage unit 130 may include multiple types of storage media. The storage unit 130 may include a combination of a portable storage medium such as a memory card, optical disk, or magneto-optical disk and a storage medium reader. The storage unit 130 may include a storage device used as a temporary storage area such as RAM.

[0067] The memory unit 130 can store, for example, a program 131, setting data 132, feature data (feature data storage unit) 133, image data D10, state estimation model M1, object estimation model M2, etc. The program 131 causes the control unit 140 to execute various control functions for operating the state estimation device 100. The setting data 132 includes data such as various settings related to the operation of the state estimation device 100 and settings related to the installation status of the imaging device 10 to be managed. The feature data 133 includes feature data calculated during the processing of multiple image data D10. The memory unit 130 can store multiple image data D10 in chronological order. The first state estimation model M1, the second state estimation model M2, the feature estimation model M3, and the object estimation model M4 are machine learning models generated by the learning device 200.

[0068] The control unit 140 is an arithmetic processing unit. The arithmetic processing unit includes, but is not limited to, a CPU, SoC, MCU, FPGA, and coprocessor. The control unit 140 comprehensively controls the operation of the state estimation device 100 to realize various functions.

[0069] Specifically, the control unit 140 executes instructions included in the program 131 stored in the storage unit 130, while referring to the data stored in the storage unit 130 as needed. The control unit 140 then controls the functional units according to the data and instructions, thereby realizing various functions. The functional units include, for example, the input unit 110 and the communication unit 120, but are not limited to these.

[0070] The control unit 140 has functional units such as a processing unit 141, an estimation unit 142, and a diagnostic unit 143. The control unit 140 realizes these functional units by executing program 131. Program 131 is a program that causes the control unit 140 of the state estimation device 100 to function as a processing unit 141, an estimation unit 142, and a diagnostic unit 143. As shown in Figure 8, the processing of the processing unit 141 and the estimation unit 142 is performed by the first processing unit 150, the second processing unit 160, and the third processing unit 170, respectively. The first processing unit 150 includes a model acquisition unit 152 included in the processing unit 141, a preprocessing unit 154, and a used image determination unit 156 and a state estimation unit 158 ​​included in the estimation unit 142. The second processing unit 160 includes a model acquisition unit 162 included in the processing unit 141, a preprocessing unit 164, and a state estimation unit 166 included in the estimation unit 142. The third processing unit 170 includes a model acquisition unit 172 included in the processing unit 141, a feature synthesis unit 174, and a feature estimation unit 176 included in the estimation unit 142.

[0071] The processing unit 141 acquires the models to be used by the estimation unit 142. The model acquisition unit 152 acquires the first state estimation model M1 and the object estimation model M4. The model acquisition unit 162 acquires the second state estimation model M2. The model acquisition unit 172 acquires the feature estimation model M3.

[0072] The processing unit 141 acquires image data D10 captured by the imaging device 10. The processing unit 141 preprocesses the image data D10 to be used by the estimation unit 142 and supplies the preprocessed image data D10 to the estimation unit 142. The preprocessing unit 154 performs various processes on the acquired image data. The preprocessing unit 154 uses the object estimation model M2 to extract traffic objects included in the acquired image data. The preprocessing unit 154 may also process the image data D10 so that traffic objects 1200 that can be used to estimate the installation state of the imaging device 10 are included in the image. Traffic objects 1200 that can be used for estimation include vehicles and the like that appear in a manner suitable for estimating the installation state parameters. Traffic objects 1200 that can be used for estimation include, for example, vehicles or people present in a predetermined region D100 of image D11, vehicles or people heading towards the imaging device 10, etc. In this embodiment, the predetermined region D100 includes, for example, a region set in advance in image D11, the central region of image D11, etc. Traffic objects 1200 unsuitable for estimation include, for example, large vehicles such as trucks, buses, and construction machinery located in a predetermined area D100 of image D11 shown by image data D10. The processing unit 141 processes the image data D10 to include at least one of the traffic objects 1200 located in the predetermined area D100 and the traffic objects 1200 facing the imaging device 10. The processing unit 141 may also have a function to process the image data D10 to delete or modify traffic objects 1200 that are unnecessary for estimating the installation status parameters of the imaging device 10 from image D11.

[0073] The preprocessing unit 164 performs various processing on the acquired image data. The preprocessing unit 164 uses the object estimation model M2 to extract traffic objects included in the acquired image data, and if the number of traffic objects is below a threshold, it selects the image data to be used. The preprocessing unit 164 may also perform processing such as brightness adjustment and edge detection to improve the accuracy of identifying the traffic environment.

[0074] The feature synthesis unit 174 synthesizes the feature data processed by the first processing unit 150 and the second processing unit 160, which are stored in the feature storage unit 133. The feature synthesis unit 174 then supplies the synthesized feature data to the feature estimation unit 176.

[0075] The estimation unit 142 performs estimation processing using the first state estimation model M1, the second state estimation model M2, and the feature estimation model M3 generated by the learning device 200. The image usage determination unit 156 selects image data to be used for estimation processing from the image data processed by the preprocessing unit 154. The image usage determination unit 156 selects a set number of image data based on criteria such as images of traffic objects extracted using the object estimation model M4 that are larger than a predetermined size, and images with a high estimation angle. The state estimation unit 158 ​​inputs the image data processed by the preprocessing unit 154 and determined to be used by the image usage determination unit 156 into the first state estimation model M1, estimates feature data (first feature data), and outputs it. The state estimation unit 166 inputs the image data processed by the preprocessing unit 164 into the second state estimation model M2, estimates feature data (second feature data), and outputs it. The feature estimation unit 176 inputs the data synthesized by the feature synthesis unit 174 into the feature estimation model M3, and estimates the installation status parameters of the imaging device 10 based on the output of the feature estimation model M3.

[0076] The diagnostic unit 143 can provide a function to diagnose the installation status of the imaging device 10 based on the installation status parameters estimated by the estimation unit 142. The diagnostic unit 143 can diagnose whether the estimated results of the installation status parameters are appropriate. The diagnostic unit 143 can diagnose the installation status of the imaging device 10 based on the installation status parameters estimated by the estimation unit 142 and the overhead view of the traffic object 1200 shown in the image data D10. The diagnostic unit 143 compares the orientation of the traffic object 1200 shown in the image data D10 with the orientation of the traffic object 1200 calculated based on the installation status parameters estimated by the estimation unit 142, and can diagnose the installation status of the imaging device 10 if the degree of agreement is higher than the judgment threshold. The diagnostic unit 143 can compare the installation status parameters estimated by the estimation unit 142 with pre-set installation status parameters and diagnose the installation status of the imaging device 10 based on the comparison result.

[0077] The control unit 140 can provide a function to supply installation status parameters estimated by the estimation unit 142, diagnostic results from the diagnostic unit 143, etc., to external devices, databases, etc. For example, the control unit 140 can perform control to supply installation status parameters estimated by the estimation unit 142, diagnostic results from the diagnostic unit 143, etc., via the communication unit 120.

[0078] The above describes an example of the functional configuration of the state estimation device 100 according to this embodiment. Note that the above configuration described with reference to Figure 8 is merely an example, and the functional configuration of the state estimation device 100 according to this embodiment is not limited to this example. The functional configuration of the state estimation device 100 according to this embodiment can be flexibly modified according to specifications and operation.

[0079] In this embodiment, the state estimation device 100 is described as having a control unit 140 that functions as a processing unit 141, an estimation unit 142, and a diagnostic unit 143. However, for example, the control unit 140 may be configured to include an estimation unit 142 and a diagnostic unit 143, but without a processing unit 141. In this case, the state estimation device 100 only needs to input the image data D10 captured by the imaging device 10 to the state estimation model M1 without preprocessing the image data D10. Furthermore, the system 1 may configure the processing unit 141 of the state estimation device 100 as part of the imaging device 10.

[0080] Figure 10 is a flowchart illustrating an example of a state estimation method performed by the state estimation device 100. The state estimation device 100 performs the method shown in Figure 10 at execution timings such as when the imaging device 10 is installed, during maintenance, or when it is instructed to run from an external source. For example, after installation and maintenance are performed at night, the state estimation device 100 acquires image data between night and early morning and performs the processing shown in Figure 10. Alternatively, image data may be acquired for processing between daytime and nighttime.

[0081] The state estimation device 100 performs estimation processing in the first processing unit (step S12). Figure 11 is a flowchart of an example of the state estimation method performed by the first processing unit. The first processing unit 150 acquires image data (step S32). The first processing unit 150 detects the time the image data was captured (step S34). The first processing unit 150 may obtain the time the image was captured from the imaging device 10, or it may perform image analysis and obtain the time the image was captured from the brightness and illuminance of the image. The first processing unit 150 determines whether the image is from night (step S36). If the first processing unit 150 determines that the image is not from night (No in step S36), it proceeds to step S44.

[0082] If the first processing unit 150 determines that it is a night image (first image data) (Yes in step S36), it analyzes the image data (step S38). Specifically, the first processing unit 150 uses the object estimation model M4 to detect traffic objects. The first processing unit 150 determines whether there are traffic objects (step S40). The determination criteria are not limited to the presence or absence of traffic objects; criteria may also include whether there are traffic objects above a threshold, or the size and position of the detected traffic objects. If the first processing unit 150 determines that there are no traffic objects (No in step S40), it proceeds to step S44. If the first processing unit 150 determines that there are traffic objects (Yes in step S40), it selects the image data to be analyzed (step S42).

[0083] The first processing unit 150 determines whether it has completed acquiring the required number of image data if it determined No in step S36, if it determined No in step S40, or if it executed the process in step S42 (step S44).

[0084] If the first processing unit 150 determines that it has not yet acquired the required number of image data (No in step S44), it returns to step S32. As a result, the first processing unit 150 repeats the process from step S32 to step S44 until it has acquired the required number of image data.

[0085] If the first processing unit 150 determines that it has completed acquiring the required number of image data (Yes in step S44), it processes the selected image data to create first feature data (step S46). The first processing unit 150 inputs the selected image data into the first state estimation model M1, estimates the feature data, and outputs the estimated feature data.

[0086] The state estimation device 100 stores the output first feature data in the feature storage unit 133 (step S14). Next, the state estimation device 100 performs estimation processing in the second processing unit (step S16). Figure 12 is a flowchart showing an example of the state estimation method performed by the second processing unit. The second processing unit 160 acquires image data (step S52). The second processing unit 160 detects the time the image data was captured (step S54). The second processing unit 160 may acquire the time the image was captured from the imaging device 10, or it may perform image analysis and acquire the time the image was captured from the brightness and illuminance of the image. The second processing unit 160 determines whether the image was taken in the early morning (step S56). If the second processing unit 160 determines that the image is not taken in the early morning (No in step S56), it proceeds to step S64.

[0087] If the second processing unit 160 determines that the image is from early morning (second image data) (Yes in step S56), it analyzes the image data (step S58). Specifically, the second processing unit 160 uses the object estimation model M4 to detect traffic objects. The second processing unit 160 determines whether the number of traffic objects is below a predetermined level (step S60). The determination criterion may be the presence or absence of traffic objects, or the size and position of the detected traffic objects. If the second processing unit 160 determines that the number of traffic objects is greater than the predetermined level (No in step S60), it proceeds to step S64. If the second processing unit 160 determines that the number of traffic objects is below a predetermined level (Yes in step S60), it selects the image data to be analyzed (step S62).

[0088] The second processing unit 160 determines whether the acquisition of the required number of image data has been completed (step S64) if it determined No in step S56, if it determined No in step S60, or if it executed the process in step S62.

[0089] If the second processing unit 160 determines that it has not yet acquired the required number of image data (No in step S64), it returns to step S52. As a result, the second processing unit 160 repeats the process from step S52 to step S64 until it has acquired the required number of image data.

[0090] If the second processing unit 160 determines that it has completed acquiring the required number of image data (Yes in step S64), it processes the selected image data to create first feature data (step S66). The second processing unit 160 inputs the selected image data into the second state estimation model M2, estimates the feature data, and outputs the estimated feature data.

[0091] Next, the state estimation device 100 combines the first feature data and the second feature data (step S18). The state estimation device 100 selects one image data from the first feature data output by the first processing unit 150 and one from the second feature data output by the second processing unit 160, combines them, and creates feature data for one image data. This generates feature data for image data that includes both the features of traffic objects extracted from the first feature data and the features of the traffic environment other than traffic objects extracted from the second feature data.

[0092] Next, the state estimation device 100 calculates an evaluation value (step S20). The state estimation device 100 inputs the synthesized feature data from the third processing unit 170 into the feature estimation model M3 to estimate the installation state parameters of the imaging device 10 that captured the input image data D10. Based on the installation state parameters estimated by the feature estimation model M3, the state estimation device 100 estimates the road area of ​​the road 1100 in the image D11 shown by the image data D10. The state estimation device 100 stores the estimated installation state parameters and road area in the storage unit 130, associating them with the image data D10.

[0093] The state estimation device 100 performs a diagnosis based on the evaluation value (step S22). The state estimation device 100 diagnoses whether the estimated installation state parameters of the imaging device 10 are appropriate. The diagnosis of whether the installation state parameters of the imaging device 10 are appropriate includes diagnosing that the installation state parameters of the imaging device 10 are appropriate if the installation state parameters do not require resetting, adjustment, etc. of the imaging device 10. The state estimation device 100 stores the diagnosis result in the storage unit 130 in association with the imaging device 10. If the state estimation device 100 diagnoses that the installation state parameters of the imaging device 10 are appropriate, it can supply the diagnosis result, installation state parameters, etc. to subsequent processing. If the state estimation device 100 diagnoses that the installation state parameters of the imaging device 10 are not appropriate, it can take an image with the imaging device 10 again and estimate the installation state parameters from the captured image data D10.

[0094] The state estimation device 100 can estimate installation state parameters with high accuracy by performing the above processing. Specifically, by using the first state estimation model M1 to extract features of traffic objects from nighttime image data, it is possible to identify traffic objects that have their lighting devices on in dark surroundings with high accuracy. Furthermore, by using the second state estimation model M2 to extract features of the traffic environment other than traffic objects from early morning image data, it is possible to extract features of the traffic environment with high accuracy using images that are not obstructed by moving objects such as vehicles. In addition, the state estimation device 100 can estimate installation state parameters with high accuracy by synthesizing these features and estimating the features. Moreover, since work is frequently performed at night, processing nighttime and early morning images allows for the estimation of installation state parameters in a short period of time after the work has been performed.

[0095] The state estimation device 100 described above has been described as being located outside the imaging device 10, but it is not limited to this. For example, the state estimation device 100 may be incorporated into the imaging device 10 and implemented by the control unit, module, etc., of the imaging device 10. For example, the state estimation device 100 may be incorporated into traffic signals, lighting equipment, communication equipment, etc., installed in the traffic environment 1000.

[0096] The state estimation device 100 described above may be implemented as a server device or the like. For example, the state estimation device 100 can be a server device that acquires image data D10 from each of the multiple imaging devices 10, estimates installation state parameters from the image data D10, and provides the estimation results.

[0097] The learning device 200 described above describes a case in which it generates a first state estimation model M1, a second state estimation model M2, a feature estimation model M3, and an object estimation model M4, but is not limited to this. For example, the learning device 200 may consist of two devices: a first device that generates the first state estimation model M1, the second state estimation model M2, and the feature estimation model M3, and a second device that generates the object estimation model M2. The first state estimation model M1, the second state estimation model M2, and the feature estimation model M3 may be represented by separate devices.

[0098] Furthermore, this disclosure may include not only cases where the first state estimation model M1, the second state estimation model M2, and the feature estimation model M3 are implemented as separate models and separate learning units, but also as an integrated model combining both models, with machine learning performed in a single integrated machine learning unit. In other words, this disclosure may also include examples where it is executed with a single model and a single learning unit.

[0099] Characteristic embodiments have been described in order to fully and clearly disclose the technology relating to the attached claims. However, the attached claims should not be limited to the above embodiments, but should be configured to embody all modifications and alternative configurations that a person skilled in the art may create within the scope of the fundamental matters presented herein. The contents of this disclosure can be modified in various ways by a person skilled in the art. Therefore, these modifications and adaptations are within the scope of this disclosure. For example, in each embodiment, each functional part, each means, each step, etc. can be added to or replaced with each functional part, each means, each step, etc. of other embodiments in a logically consistent manner. Also, in each embodiment, multiple functional parts, each means, each step, etc. can be combined into one or divided into two. Furthermore, each embodiment of this disclosure described above is not limited to being implemented strictly according to the respective embodiments described, but can be implemented by combining or omitting features as appropriate.

[0100] [Note] (Note 1) A first object estimation model is machine-trained to estimate first feature data obtained by estimating the features of a first extractable object from input first image data, using first training data which includes first image data of a traffic environment captured by an imaging device and first ground truth data of a first extractable object including a moving object contained in the first image data, and comprises a first state estimation unit that estimates the first feature data from input image data, A second object estimation model is machine-trained to estimate second feature data obtained by estimating the features of a first object from input first image data, using second training data which includes second image data of a traffic environment captured by an imaging device and second ground truth data of a second object to be extracted, including roads, included in the second image data, and a second state estimation unit that estimates the second feature data from input image data, A state estimation model trained using machine learning to estimate the installation state parameters of an imaging device that captured input image data, using third training data which includes image data of the traffic environment captured by the imaging device and ground truth data of the installation state parameters of the imaging device that captured the image data, a feature estimation unit that estimates the installation state parameters of the imaging device that captured input image data from data synthesized from the first feature data and the second feature data, A diagnostic unit that diagnoses the installation status of the imaging device based on the estimated installation status parameters, Equipped with, State estimation device. (Note 2) In the state estimation device described in (Appendix 1), A first preprocessing unit processes the first image data obtained by the imaging device from the traffic environment so that it includes a first target object that can be used for estimation. The imaging device further comprises a second preprocessing unit that processes the second image data, which is an image of the traffic environment, so that it includes a second object to be extracted that can be used for estimation. The first processing unit inputs the first image data processed by the processing unit into the first state estimation model and estimates the first feature data. The second processing unit inputs the first image data processed by the processing unit into the second state estimation model and estimates the second feature data. State estimation device. (Note 3) In the state estimation device described in (Appendix 1), The first image data is an image taken at night, The second image data mentioned above is an image taken during the daytime. State estimation device. (Note 4) In the state estimation device described in (Appendix 1), The system includes a feature storage unit that stores the aforementioned first feature data. State estimation device. (Note 5) In the state estimation device described in (Appendix 1), The first state estimation unit, the second state estimation unit, and the feature estimation unit are located on a cloud server. State estimation device. (Note 6) In the state estimation device described in (Appendix 1), The diagnostic unit diagnoses the installation status of the imaging device based on the installation status parameters estimated by the estimation unit and the overhead view of the traffic object shown in the image data. State estimation device. (Note 7) In the state estimation device described in (Appendix 6), The diagnostic unit compares the orientation of the traffic object shown in the image data with the orientation of the traffic object calculated based on the installation state parameters estimated by the estimation unit, and diagnoses the installation state of the imaging device if the degree of agreement is higher than the judgment threshold. State estimation device. (Note 8) Computers A first object estimation model, trained using machine learning to estimate first feature data obtained by estimating the features of a first object from input first image data, uses first training data comprising first image data captured by an imaging device of the traffic environment and first ground truth data of a first object to be extracted, including a moving object included in the first image data. The first object estimation model estimates the first feature data from input image data, A second object estimation model, trained using machine learning to estimate second feature data obtained by estimating the features of a first object from input first image data, uses second training data, which includes second image data of a traffic environment captured by an imaging device and second ground truth data of a second object to be extracted, including roads, contained in the second image data. A state estimation model, trained using machine learning to estimate the installation state parameters of an imaging device that captured input image data, using third training data comprising image data of the traffic environment captured by the imaging device and ground truth data of the installation state parameters of the imaging device that captured the image data, estimates the installation state parameters of the imaging device that captured input image data from data synthesized from the first feature data and the second feature data. To diagnose the installation status of the imaging device based on the estimated installation status parameters, A state estimation method comprising: (Note 9) On the computer, A first object estimation model, trained using machine learning to estimate first feature data obtained by estimating the features of a first object from input first image data, uses first training data comprising first image data captured by an imaging device of the traffic environment and first ground truth data of a first object to be extracted, including a moving object included in the first image data. The first object estimation model estimates the first feature data from input image data, A second object estimation model, trained using machine learning to estimate second feature data obtained by estimating the features of a first object from input first image data, uses second training data, which includes second image data of a traffic environment captured by an imaging device and second ground truth data of a second object to be extracted, including roads, contained in the second image data. A state estimation model, trained using machine learning to estimate the installation state parameters of an imaging device that captured input image data, using third training data comprising image data of the traffic environment captured by the imaging device and ground truth data of the installation state parameters of the imaging device that captured the image data, estimates the installation state parameters of the imaging device that captured input image data from data synthesized from the first feature data and the second feature data. To diagnose the installation status of the imaging device based on the estimated installation status parameters, A state estimation program that executes the following. (Note 10) A first state estimation unit trained to estimate first feature data from first image data including a moving object captured by an imaging device, A second state estimation unit trained to estimate second feature data from second image data including roads captured by the aforementioned imaging device, The system comprises a feature estimation unit trained to estimate the installation state parameters of the imaging device that captured the input image data, based on the first feature data and the second feature data. State estimation device. (Note 11) In the state estimation device described in (Appendix 10), The state estimation device wherein the second image data is an image with fewer moving objects compared to the first image data. [Explanation of Symbols]

[0101] 1 System 10 Imaging device 100 State Estimation Device 110 Input Section 120 Communications Department 130 Storage section 131 Programs 132 Configuration Data 133 Feature Memory Unit 140 Control Unit 141 Processing Unit 142 Estimation Department 143 Diagnostic Department 150 First Processing Unit 152, 162, 172 Model acquisition section 154, 164 Pre-processing section 156 Image detection unit 158, 166 State Estimation Unit 160 Second Processing Unit 170 Third Processing Unit 174 Feature Synthesis Section 176 Feature Estimation Unit 200 Learning Devices 210 Display section 220 Operation section 230 Communications Department 240 Storage section 241 Programs 242 training data 250 Control Unit 251 First acquisition part 252 First Machine Learning Department 253 Second Acquisition Department 254 Second Machine Learning Department 255 Third acquisition part 256 Third Machine Learning Department 257 4th acquisition part 258 4th Machine Learning Department 1000 Traffic environment 1100 Road 1200 Transportation 500, 520, 540, 2100 Input Layers 510, 530, 530, 2200 Middle layer 2210 Feature Extraction Layer 2220 bonding layer 2300 Output Layer D10 Image Data D11 Image D21 Correct Value Data D22 Correct Value Data D100 Predetermined area M1 First State Estimation Model M2 Second State Estimation Model M3 Feature Estimation Model M4 Object Estimation Model

Claims

1. A first state estimation unit trained to estimate first feature data from first image data including a moving object captured by an imaging device, A second state estimation unit trained to estimate second feature data from second image data including roads captured by the aforementioned imaging device, The system comprises a feature estimation unit trained to estimate installation state parameters, including at least one of the installation angle and installation position of the imaging device that captured the input image data, from the first feature data and the second feature data, The first image data is an image taken at night, The second image data mentioned above is an image taken during the daytime. State estimation device.

2. In the state estimation device according to claim 1, The second image data is an image with fewer moving objects compared to the first image data. State estimation device.

3. A first object estimation model is machine-trained to estimate first feature data obtained by estimating the features of a first object from input first image data, using first training data which includes first image data of a traffic environment captured by an imaging device and first ground truth data of a first object to be extracted, including a moving object included in the first image data, and a first state estimation unit that estimates the first feature data from input image data, A second object estimation model is machine-trained to estimate second feature data obtained by estimating the features of a first object from input first image data, using second training data which includes second image data of a traffic environment captured by an imaging device and second ground truth data of a second object to be extracted, including roads, included in the second image data. The second state estimation unit estimates the second feature data from input image data. A state estimation model trained using machine learning to estimate the installation state parameters of an imaging device that captured input image data, using third training data which includes image data of the traffic environment captured by the imaging device and ground truth data of the installation state parameters of the imaging device that captured the image data, and a feature estimation unit that estimates the installation state parameters of the imaging device that captured input image data using the first feature data and the second feature data, A diagnostic unit that diagnoses the installation status of the imaging device based on the estimated installation status parameters, Equipped with, The first image data is an image taken at night, The second image data mentioned above is an image taken during the daytime. State estimation device.

4. In the state estimation device according to claim 3, A first preprocessing unit processes the first image data captured by the imaging device of the traffic environment so that it includes a first target object that can be used for estimation. The imaging device further comprises a second preprocessing unit that processes the second image data, which is an image of the traffic environment, so that it includes a second object to be extracted that can be used for estimation. The first preprocessing unit inputs the processed first image data into the first state estimation model and estimates the first feature data. The second preprocessing unit inputs the processed first image data into the second state estimation model and estimates the second feature data. State estimation device.

5. In the state estimation device according to claim 3, The system includes a feature storage unit that stores the aforementioned first feature data. State estimation device.

6. In the state estimation device according to claim 3, The first state estimation unit, the second state estimation unit, and the feature estimation unit are located on a cloud server. State estimation device.

7. In the state estimation device according to claim 3, The diagnostic unit diagnoses the installation status of the imaging device based on the installation status parameters and the overhead view of the moving object shown in the image data. State estimation device.

8. In the state estimation device according to claim 7, The diagnostic unit compares the posture of the moving object shown in the image data with the posture of the moving object calculated based on the installation status parameters, and diagnoses the installation status of the imaging device if the degree of agreement is higher than the judgment threshold. State estimation device.

9. Computers A first object estimation model, trained through machine learning to estimate first feature data obtained by estimating the features of a first extracted object from input first image data, uses first training data comprising first image data captured by an imaging device of the traffic environment and first ground truth data of a first extracted object including a moving object contained in the first image data, thereby estimating first feature data from input image data, A second object estimation model, trained through machine learning to estimate second feature data obtained by estimating the features of a first object from input first image data, uses second training data, which includes second image data of a traffic environment captured by an imaging device and second ground truth data of a second object to be extracted, including roads, included in the second image data. Machine learning is used to estimate the installation state parameters of the imaging device that captured the input image data, using third training data which includes image data of the traffic environment captured by the imaging device and ground truth data of the installation state parameters of the imaging device that captured the image data. Using the state estimation model described above, the installation state parameters of the imaging device that captured the input image data are estimated from the data using the first feature data and the second feature data. to do, To diagnose the installation status of the imaging device based on the estimated installation status parameters, Equipped with, The first image data is an image taken at night, The second image data mentioned above is an image taken during the daytime. State estimation method.

10. On the computer, A first object estimation model, trained through machine learning to estimate first feature data obtained by estimating the features of a first extracted object from input first image data, uses first training data comprising first image data captured by an imaging device of the traffic environment and first ground truth data of a first extracted object including a moving object contained in the first image data, thereby estimating first feature data from input image data, A second object estimation model, trained through machine learning to estimate second feature data obtained by estimating the features of a first object from input first image data, uses second training data, which includes second image data of a traffic environment captured by an imaging device and second ground truth data of a second object to be extracted, including roads, included in the second image data. A state estimation model, trained using machine learning to estimate the installation state parameters of an imaging device that captured input image data, using third training data comprising image data of the traffic environment captured by the imaging device and ground truth data of the installation state parameters of the imaging device that captured the image data, estimates the installation state parameters of the imaging device that captured input image data using the first feature data and the second feature data. To diagnose the installation status of the imaging device based on the estimated installation status parameters, Make it run, The first image data is an image taken at night, The second image data mentioned above is an image taken during the daytime. State estimation program.

Citation Information

Patent Citations

  • Camera calibration system, and measuring vehicle and roadside device for the same

    JP2012010036A

  • Information processing apparatus, information processing method, and program

    JP2017129942A

  • Calibration apparatus and calibration method

    JP2021174044A

  • Method and device for calibrating pitch of camera on vehicle and method and device for continual learning of vanishing point estimation model to be used for calibrating the pitch

    US11080544B1