State estimation device, state estimation method, and state estimation program

The state estimation device automates imaging device calibration using machine-learning, addressing the labor-intensive manual calibration of conventional systems, enabling efficient and unrestricted deployment in traffic environments.

JP7739190B2Active Publication Date: 2025-09-16KYOCERA CORP
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2022011231
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-09-16
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

Conventional imaging devices require manual input of lane marks and dedicated tools for calibration, which is labor-intensive and restricts their widespread use in traffic environments.

Method used

A state estimation device and method that utilizes machine-learning to estimate the installation state parameters of imaging devices using image data, eliminating the need for manual labor and dedicated tools by employing a state estimation model trained on teacher data to diagnose the installation state of imaging devices.

Benefits of technology

Enables the widespread use of imaging devices in traffic environments by automating the calibration process, improving maintenance efficiency and eliminating the need for traffic restrictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007739190000001
    Figure 0007739190000001
  • Figure 0007739190000002
    Figure 0007739190000002
  • Figure 0007739190000003
    Figure 0007739190000003
Patent Text Reader

Abstract

To enable estimation of an installation state of an imaging device for imaging a traffic environment without requiring human work or traffic regulation.SOLUTION: For example, a state estimation device 100 includes an estimation unit 142 that estimates an installation state parameter of the imaging device that captures input image data using a state estimation model M1 that has undergone machine learning so as to estimate the installation state parameter of the imaging device that captures the input image data using first teacher data having image data of a traffic environment captured by the imaging device and first correct value data of an installation state parameter of the imaging device that captures the image data, and a diagnosis unit 143 that diagnoses the installation state of the imaging device on the basis of the estimated installation state parameter.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to a state estimation device, a state estimation method, and a state estimation program. [Background technology]

[0002] It is known that cameras installed on roads, roadsides, etc. are calibrated. Patent Document 1 discloses that calibration is performed using a measurement vehicle equipped with a GPS receiver, a data transmitter, landmarks, etc. Patent Document 2 discloses that in camera calibration, when the direction of a line existing on the road plane is input in a captured image, road plane parameters are estimated based on the direction and a direction expressed by an arithmetic expression including road plane parameters. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-10036 [Patent Document 2] Japanese Patent Application Laid-Open No. 2017-129942 Summary of the Invention [Problem to be solved by the invention]

[0004] In Patent Document 1, a measurement vehicle is required, and an operator is required for calibration. In Patent Document 2, the lane marks on the road must be manually input into the image, which is a labor-intensive task. For this reason, there is a need to estimate the installation status of an imaging device that captures an image of a traffic environment without requiring manual work or traffic regulations for conventional imaging devices that capture images of roads. [Means for solving the problem]

[0005] A state estimation device according to one aspect is a state estimation model that has been machine-trained to estimate the installation state parameters of the imaging device that captured input image data using first teacher data that includes image data of a traffic environment captured by an imaging device and first correct answer value data of the installation state parameters of the imaging device that captured the image data, and includes an estimation unit that estimates the installation state parameters of the imaging device that captured the input image data, and a diagnosis unit that diagnoses the installation state of the imaging device based on the estimated installation state parameters.

[0006] A state estimation method according to one aspect includes a computer estimating the installation state parameters of the imaging device that captured the input image data using a state estimation model that has been machine-learned to estimate the installation state parameters of the imaging device that captured the input image data, using first teacher data that includes image data of a traffic environment and first correct value data of the installation state parameters of the imaging device that captured the image data, and diagnosing the installation state of the imaging device based on the estimated installation state parameters.

[0007] A state estimation program according to one aspect causes a computer to estimate the installation state parameters of the imaging device that captured the input image data using a state estimation model that has been machine-learned to estimate the installation state parameters of the imaging device that captured the input image data, using first teacher data that includes image data of a traffic environment and first correct answer value data of the installation state parameters of the imaging device that captured the image data, and to diagnose the installation state of the imaging device based on the estimated installation state parameters. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of the relationship between a learning device and a state estimation device according to an embodiment. [Figure 2] FIG. 2 is a diagram showing an example of image data captured by the imaging device shown in FIG. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of a learning device according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a CNN used by the learning device illustrated in FIG. [Figure 5] FIG. 5 is a diagram illustrating an example of the configuration of a state estimation device according to the embodiment. [Figure 6] FIG. 6 is a flowchart showing an example of a state estimation method executed by the state estimation device. [Figure 7] FIG. 7 is a flowchart illustrating an example of a processing procedure of an object detection step executed by the state estimation device. [Figure 8] FIG. 8 is a diagram for explaining an example of a tracking process for image data including a traffic object. [Figure 9] FIG. 9 is a diagram for explaining an example of estimation of installation state parameters of the state estimating device according to the embodiment. [Figure 10] FIG. 10 is a diagram for explaining an example of estimation of installation state parameters of the state estimating device according to the embodiment. [Figure 11] FIG. 11 is a diagram for explaining an example of estimation of installation state parameters of the state estimating device according to the embodiment. [Figure 12] FIG. 12 is a diagram for explaining another example of the result diagnosis executed by the state estimating device. DETAILED DESCRIPTION OF THE INVENTION

[0009] A number of embodiments for implementing a state estimation device, a learning device, a state estimation method, a state estimation program, and the like according to the present application will be described in detail with reference to the drawings. Note that the following description does not limit the present invention. Furthermore, the components in the following description include those that can be easily imagined by a person skilled in the art, those that are substantially the same, and those that are within the so-called equivalent scope. In the following description, similar components may be assigned the same reference numerals. Furthermore, duplicated descriptions may be omitted.

[0010] (System Overview) In conventional systems, dedicated tools and work were required to link captured images with the real world using information on the installation state of the imaging device. Furthermore, in conventional systems, road control work was required because the imaging device was installed near a road. The state estimation device according to this embodiment eliminates the need for work using tools and road control work, contributing to the widespread use of imaging devices 10 in traffic environments.

[0011] FIG. 1 is a diagram illustrating an example of the relationship between a learning device and a state estimation device according to an embodiment. FIG. 2 is a diagram illustrating an example of image data captured by the imaging device illustrated in FIG. 1. As illustrated in FIG. 1, a system 1 includes an imaging device 10 and a state estimation device 100. The imaging device 10 can acquire image data D10 of an image of a traffic environment 1000. The state estimation device 100 has a function of acquiring the image data D10 from the imaging device 10 and estimating the installation state of the imaging device 10 based on the image data D10. The imaging device 10 and the state estimation device 100 are configured to be able to communicate with each other via wired or wireless communication. In the example illustrated in FIG. 1, for simplicity of explanation, a case will be described in which the system 1 includes one imaging device 10 and one state estimation device 100. However, the system 1 may include a plurality of imaging devices 10 and one state estimation device 100.

[0012] The imaging device 10 is installed so as to be able to capture an image of a traffic environment 1000 including a road 1100 and traffic objects 1200 moving on the road 1100. The traffic objects 1200 moving on the road 1100 include, for example, vehicles, people, etc. that can move on the road 1100. The traffic objects 1200 include, for example, large automobiles, standard automobiles, large special-purpose automobiles, large motorcycles, standard motorcycles, small special-purpose automobiles, etc., as defined by the Road Traffic Act, but may also include other vehicles, moving bodies, etc. Note that large automobiles include automobiles with a gross vehicle weight of 8,000 kg or more, those with a maximum load capacity of 5,000 kg or more, and those with a passenger capacity of 11 or more (such as buses and trucks). The imaging device 10 can capture images electronically using an image sensor, for example, a CCD (Charge Coupled Device Image Sensor) or a CMOS (Complementary Metal Oxide Semiconductor). The imaging device 10 is installed with its imaging direction facing the road plane of the traffic environment 1000. The imaging device 10 can be installed, for example, on a road, at an intersection, in a parking lot, or the like.

[0013] In the example shown in FIG. 1 , the imaging device 10 is installed on the roadside at an installation angle that allows it to capture an image of a traffic environment 1000 including a road 1100 and its surroundings from above. The imaging device 10 acquires image data D10 by capturing an image of the traffic environment 1000. The imaging device 10 may be installed so that its imaging direction is fixed, or may be installed so that its imaging direction can be changed at the same position using a movable mechanism. As shown in FIG. 2 , the image data D10 of the imaging device 10 is data representing an image D11 having a first region D110 representing multiple roads 1100 and a second region D120 representing traffic objects 1200 passing through the road 1100. The imaging device 10 supplies the captured image data D10 to the state estimation device 100. In this embodiment, the image data D10 includes, for example, a two-dimensional image such as a moving image or a still image. In the image data D10, a predetermined region D100 is set in advance for the image D11. The predetermined area D100 is an area that includes traffic objects 1200 that can be used for estimation, and can be set appropriately based on the traffic environment 1000 to be captured. The predetermined area D100 may be the entire area of ​​the image D11. The traffic objects 1200 that can be used for estimation include, for example, traffic objects 1200 that have been used as correct answers in the machine learning of the state estimation model M1. The traffic objects 1200 that can be used for estimation are traffic objects 1200 that are suitable for estimation by the state estimation model M1.

[0014] As shown in Fig. 1, the state estimation device 100 may be provided near the imaging device 10, or may be provided at a location distant from the imaging device 10. In the example shown in Fig. 1, for the sake of simplicity, a case will be described in which image data D10 is supplied to the state estimation device 100 from a single imaging device 10, but image data D10 may also be supplied from, for example, each of a plurality of imaging devices 10. A traffic object 1200 is moving along a lane toward the imaging device 10 in a road direction C1, and a road direction C2 indicates the direction of an oncoming lane.

[0015] The state estimation device 100 has a function of managing installation state parameters of the imaging device 10. The installation state parameters include, for example, the installation angle and installation position of the imaging device 10. The installation state parameters may also include, for example, the number of pixels of the imaging device 10 and the size of the image D11. The state estimation device 100 can estimate the installation state parameters of the imaging device 10 that captured the image data D10 using a state estimation model M1 that has been machine-learned by the learning device 200. The state estimation device 100 inputs the image data D10 to the state estimation model M1 and can estimate the output of the state estimation model M1 as the installation state parameters of the imaging device 10.

[0016] The learning device 200 is, for example, a computer, a server device, or the like. The learning device 200 may or may not be included in the configuration of the system 1. The learning device 200 acquires a plurality of first teacher data including image data D10 of a traffic environment 1000 including a traffic object 1200 and correct value data D21 of installation state parameters of the imaging device 10 that captured the image data D10. The correct value data D21 includes data indicating correct values ​​of, for example, the installation angle (α, β, γ) of the imaging device 10, the installation position (x, y, z), the number of pixels, the size of the image D11, and the like. The correct value data D21 is an example of first correct value data. The installation angle includes, for example, a pitch angle α in the direction in which the imaging device 10 heads downward, a yaw angle β at which the imaging device 10 can swing horizontally in the imaging direction, and a roll angle γ in the direction in which the imaging device 10 tilts. The installation position has, for example, a position (x, z) on the road surface and a height y. The correct value data D21 may be, for example, a correct value obtained by combining two values, α and γ, that can identify the orientation relative to the road surface. The correct value data D21 may be, for example, a correct value obtained by combining three values, α, γ, and y, that can identify the scale. The correct value data D21 may be, for example, a correct value obtained by combining four values, α, β, γ, and y, that can identify the main road direction. The correct value data D21 may be, for example, a correct value obtained by combining six values, α, β, γ, x, y, and z, that are used in general calibration.

[0017] The learning device 200 generates a state estimation model M1 that estimates installation state parameters of the imaging device 10 that captured input image data D10 through machine learning using multiple pieces of first training data. Supervised machine learning can use algorithms such as neural networks, linear regression, and logistic regression. The state estimation model M1 is a model that is machine-learned from multiple pieces of training data, image data D10 and ground truth data D21, so as to estimate installation state parameters of the imaging device 10 that captured the input image data D10. When the image data D10 is input, the state estimation model M1 estimates installation state parameters of the imaging device 10 that captured the image data D10 and outputs the estimation results. By providing the generated state estimation model M1 to the state estimation device 100, the learning device 200 can contribute to eliminating the need for dedicated tools or manual labor when the state estimation device 100 calculates the installation state of the imaging device 10. An example of the learning device 200 will be described later.

[0018] The state estimation device 100 inputs image data D to the state estimation model M1 provided by the learning device 200, and can estimate installation state parameters of the captured image data D10 based on the output of the state estimation model M1. The state estimation device 100 can diagnose the installation state of the imaging device 10 based on the estimated installation state parameters. This allows the state estimation device 100 to eliminate the need for dedicated jigs or human labor to calculate the installation state of the imaging device 10 when installing or maintaining the imaging device 10 in the traffic environment 1000. By eliminating the need for jigs and human labor, the state estimation device 100 can eliminate the need for traffic restrictions. As a result, the state estimation device 100 can contribute to the widespread use of imaging devices 10 installed in the traffic environment 1000 and can improve the efficiency of maintenance.

[0019] The learning device 200 can acquire a plurality of second teacher data sets including image data D10 of a traffic environment 1000 including traffic objects 1200 captured by the imaging device 10 and correct answer value data D22 for object detection in the image data D10. The correct answer value data D22 includes, for example, data indicating correct answer values ​​for the position, size, type, number of objects, etc. of the traffic objects 1200 in the image D11. The correct answer value data D22 is an example of second correct answer value data. The correct answer value data D22 includes, for example, two data sets for the position (x, y) of the object in the image D11, two data sets for the size (w, h) of the object, and one data set for the object type, for a total of five data sets, the number of which corresponds to the number of objects in the image D11. The object types include, for example, a person, a large vehicle, a standard vehicle, a large special-purpose vehicle, a large motorcycle, a standard motorcycle, a small special-purpose vehicle, a bicycle, etc.

[0020] The learning device 200 generates, through machine learning using a plurality of second teacher data, an object estimation model M2 that estimates at least one of the position, size, and type of a traffic object 1200 (object) in the traffic environment 1000 represented by the input image data D10. The object estimation model M2 is a model obtained by machine learning of the image data D10 and correct answer data D22 of the plurality of teacher data so as to estimate the position, size, and type of the traffic object 1200 in the traffic environment 1000 represented by the input image data D10. When the image data D10 is input, the object estimation model M2 estimates the position, size, type, and number of objects of the traffic object 1200 in the traffic environment 1000 represented by the image data D10, and outputs the estimation result. The learning device 200 can provide the generated object estimation model M2 to the state estimation device 100.

[0021] The state estimation device 100 has a function of processing image data D10 obtained by capturing an image of a traffic environment 1000 by the imaging device 10 so that the image data D10 includes a traffic object 1200 used for estimation. The state estimation device 100 can process the image data D10 obtained by capturing an image of the traffic environment 1000 by the imaging device 10, for example, using an object estimation model M2 so that the image data D10 includes the traffic object 1200 used for estimation. This enables the state estimation device 100 to input image data D10 that can be used to estimate installation state parameters of the imaging device 10 to the state estimation model M1, thereby improving the estimation accuracy of the state estimation model M1. The image data D10 that can be used to estimate installation state parameters of the imaging device 10 is data that can improve the probability of the estimation results of the state estimation model M1.

[0022] The system 1 can provide a function for managing the maintenance of one or more image capture devices 10 using the estimation results of the state estimation device 100. The system 1 can provide a function for instructing changes to the installation state of the image capture device 10 based on the installation state parameters and installation positions estimated by the state estimation device 100.

[0023] (Learning device) Fig. 3 is a diagram illustrating an example of the configuration of a learning device 200 according to an embodiment. As shown in Fig. 3, the learning device 200 includes a display unit 210, an operation unit 220, a communication unit 230, a storage unit 240, and a control unit 250. The control unit 250 is electrically connected to the display unit 210, the operation unit 220, the communication unit 230, the storage unit 240, and the like. In this embodiment, an example will be described in which the learning device 200 performs machine learning using a convolutional neural network (CNN), which is a type of neural network.

[0024] The display unit 210 is configured to be able to display various types of information under the control of the control unit 250. The display unit 210 has a display panel such as a liquid crystal display, an organic EL display, etc. The display unit 210 displays information such as characters, figures, and images in response to a signal input from the control unit 250.

[0025] The operation unit 220 has one or more devices for accepting user operations. The devices for accepting user operations include, for example, keys, buttons, a touch screen, a mouse, etc. The operation unit 220 can supply a signal corresponding to the accepted operation to the control unit 250.

[0026] The communication unit 230 can communicate with, for example, the state estimation device 100, other communication devices, etc. The communication unit 230 can support various communication standards. The communication unit 230 can send and receive various types of data via, for example, a wired or wireless network, etc. The communication unit 230 can supply the received data to the control unit 250. The communication unit 230 can send data to a destination instructed by the control unit 250.

[0027] The storage unit 240 can store programs and data. The storage unit 240 is also used as a working area for temporarily storing processing results of the control unit 250. The storage unit 240 may include any non-transitory storage medium, such as a semiconductor storage medium or a magnetic storage medium. The storage unit 240 may include multiple types of storage media. The storage unit 240 may include a combination of a portable storage medium, such as a memory card, an optical disk, or a magneto-optical disk, and a storage medium reader. The storage unit 240 may include a storage device used as a temporary storage area, such as a RAM (Random Access Memory).

[0028] The storage unit 240 can store various data such as a program 241, training data 242, a state estimation model M1, an object estimation model M2, etc. The program 241 causes the control unit 250 to execute a function of generating a state estimation model using CNN to estimate installation state parameters of the imaging device 10 that captured the image data D10. The program 241 causes the control unit 250 to execute a function of generating an object estimation model using CNN to estimate information about an object indicated by the image data D10.

[0029] The teacher data 242 is learning data, training data, etc. used in machine learning. The teacher data 242 includes data combining image data D10 used in machine learning for state estimation and correct value data D21 associated with the image data D10. The image data D10 is input data for supervised learning. For example, the image data D10 represents a color image of a traffic environment 1000 including a traffic object 1200, and has a pixel count of 1280 × 960. The correct value data D21 includes data indicating installation state parameters of the imaging device 10 that captured the image data D10. The correct value data D21 is correct data for supervised machine learning. The correct value data D21 includes data indicating six parameters (values), for example, the installation angle (α, β, γ) and installation position (x, y, z) of the imaging device 10.

[0030] The training data 242 further includes data combining image data D10 used for machine learning of object estimation and correct answer data D22 associated with the image data D10. For example, the image data D10 indicates a color image of a traffic environment 1000 including a traffic object 1200, and has a pixel count of 1280 x 960. The correct answer data D22 includes data indicating the object position, object size, object type, etc. of the object (traffic object 1200) indicated by the image data D10, the number of which corresponds to the number of traffic objects 1200 (objects) included in the image. The object position includes, for example, coordinates (x, y) in the associated image data D10. The object size includes, for example, the width, height, etc. of the object indicated by the associated image data D10.

[0031] The state estimation model M1 is a learning model generated by extracting features, regularities, patterns, etc. of the image data D10 using the image data D10 and correct value data D21 contained in the teacher data 242, and by machine learning the relationship with the correct value data D21. When image data D10 is input, the state estimation model M1 predicts teacher data 242 similar to the features, etc. of the image data D10, estimates installation state parameters of the imaging device 10 that captured the image data D10 based on the correct value data D21, and outputs the estimation results.

[0032] The object estimation model M2 is a learning model generated by extracting features, regularities, patterns, etc. of objects in the image data D10 using the image data D10 and correct value data D22 contained in the teacher data 242, and by machine learning the relationship with the correct value data D22. When image data D10 is input, the object estimation model M2 predicts teacher data 242 that is similar to the features, etc. of the objects in the image data D10, and estimates the position, size, type, etc. of the object in the image represented by the image data D10 based on the correct value data D22, and outputs the estimation results.

[0033] The control unit 250 is an arithmetic processing device. Examples of arithmetic processing devices include, but are not limited to, a central processing unit (CPU), a system-on-a-chip (SoC), a microcontrol unit (MCU), a field-programmable gate array (FPGA), and a coprocessor. The control unit 250 can comprehensively control the operation of the learning device 200 to realize various functions.

[0034] Specifically, the control unit 250 can execute instructions contained in the program 241 stored in the storage unit 240 while referring to information stored in the storage unit 240 as necessary. The control unit 250 then controls the functional units in accordance with the data and instructions, thereby realizing various functions. The functional units include, but are not limited to, the display unit 210 and the communication unit 230, for example.

[0035] The control unit 250 has functional units such as a first acquisition unit 251, a first machine learning unit 252, a second acquisition unit 253, and a second machine learning unit 254. The control unit 250 realizes the functions of the first acquisition unit 251, the first machine learning unit 252, the second acquisition unit 253, and the second machine learning unit 254 by executing a program 241. The program 241 is a program for causing the control unit 250 of the learning device 200 to function as the first acquisition unit 251, the first machine learning unit 252, the second acquisition unit 253, and the second machine learning unit 254.

[0036] The first acquisition unit 251 acquires, as teacher data, image data D10 obtained by capturing an image of a traffic environment 1000 including a traffic object 1200 and correct value data D21 of an installation state parameter of the imaging device 10 that captured the image data D10. The first acquisition unit 251 acquires the image data D10 and the correct value data D21 from a preset storage destination, a storage destination selected by the operation unit 220, or the like, and stores the image data D10 and the correct value data D21 in association with the teacher data 242 in the storage unit 240. The first acquisition unit 251 acquires a plurality of image data D10 and correct value data D21 to be used for machine learning.

[0037] The first machine learning unit 252 generates a state estimation model M1 that estimates installation state parameters of the imaging device 10 that captured the input image data D10, through machine learning using a plurality of training data 242 (first training data) acquired by the first acquisition unit 251. The first machine learning unit 252 constructs a CNN based on the training data 242, for example. The CNN receives the image data D10 as input, and a network is constructed to output a classification result for the image data D10. The classification result includes information for estimating installation state parameters of the imaging device 10 that captured the image data D10.

[0038] FIG. 4 is a diagram showing an example of a CNN used by the learning device 200 shown in FIG. 3. The first machine learning unit 252 constructs the CNN shown in FIG. 4 based on the acquired teacher data 242. As is well known, the CNN has an input layer 2100, an intermediate layer 2200, and an output layer 2300. The input layer 2100 supplies input image data D10 to the intermediate layer 2200. The input image data D10 is, for example, data representing a 640×640×3 color image.

[0039] The intermediate layer 2200 includes multiple feature extraction layers 2210 and a combination layer 2220. Each of the multiple feature extraction layers 2210 extracts a different feature from the image D11 represented by the image data D10. The extracted features of the image D11 include, for example, features related to the road 1100, lanes, etc. in the image. The feature extraction layer 2210 includes, for example, one or more convolution layers and a pooling layer, and extracts desired features from the input image data D10. The convolution layer of the feature extraction layer 2210 extracts portions of the image D11 that resemble the shape of a filter (weight) by performing a convolution operation on the input data. The convolution layer is configured to apply an activation function to a feature map, which is the operation result. In this embodiment, the activation function applied is a Relu (Rectified Linear Unit) function, but a sigmoid function or the like may also be applied. The pooling layer of the feature extraction layer 2210 summarizes the features of the image data D10 obtained by convolution into maximum and average values, so that the extracted features are considered to be the same even if their positions change. The feature extraction layer 2210 learns the optimal output, so by increasing the number of convolutional layers and pooling layers, it is possible to extract more advanced and complex features. The combination layer 2220 combines the features extracted by multiple feature extraction layers 2210 and outputs them to the output layer 2300.

[0040] The output layer 2300 estimates installation state parameters of the imaging device 10 that captured the image data D10, based on the features extracted in the intermediate layer 2200 and the correct value data D21. The output layer 2300 identifies the correct value data D21 associated with features similar to the features output by the combined layer 2220, and outputs the installation state parameters indicated by the correct value data D21.

[0041] The first machine learning unit 252 performs machine learning using the plurality of teacher data 242 (first teacher data) acquired by the first acquisition unit 251, thereby determining weights and the like of the intermediate layer 2200 and setting them in the CNN so as to estimate installation state parameters of the imaging device 10 that captured the input image data D10, thereby generating a state estimation model M1. The first machine learning unit 252 stores the generated state estimation model M1 in the storage unit 240. As a result, when image data D10 is input, the state estimation model M1 can output a result of estimating installation state parameters of the imaging device 10 that captured the image data D10.

[0042] 3 acquires image data D10 of a traffic environment 1000 including a traffic object 1200 captured by the imaging device 10, and correct answer value data D22 for object detection in the image data D10, as teacher data 242. The second acquisition unit 253 acquires the image data D10 and the correct answer value data D22 from a preset storage destination, a storage destination selected by the operation unit 220, or the like, and stores the image data D10 and the correct answer value data D22 in association with the teacher data 242 in the storage unit 240. The second acquisition unit 253 acquires a plurality of image data D10 and correct answer value data D22 to be used for machine learning.

[0043] The second machine learning unit 254 generates an object estimation model M2 that estimates at least one of the position, size, and type of a traffic object 1200 in the traffic environment 1000 represented by the input image data D10 through machine learning using the training data 242 (second training data). The second machine learning unit 254 constructs a CNN capable of detecting the traffic object 1200 (object) based on, for example, the training data 242. The CNN receives the image data D10 as input, and constructs a network to output an estimation result that estimates the position, size, and type of the traffic object 1200 in the traffic environment 1000 represented by the image data D10. The identification result includes the position, size, and type of the traffic object 1200 represented by the image data D10.

[0044] The second machine learning unit 254 constructs the CNN shown in FIG. 4 based on the acquired training data 242. The CNN has an input layer 2100, an intermediate layer 2200, and an output layer 2300. The input layer 2100 supplies input image data D10 to the intermediate layer 2200. The intermediate layer 2200 has multiple feature extraction layers 2210 and a combination layer 2220. The feature extraction layer 2210 extracts traffic objects 1200 (features) in the image represented by the image data D10. The feature extraction layer 2210 has, for example, multiple convolutional layers and a pooling layer, and extracts the traffic objects 1200 as features from the input image data D10. The convolutional layer of the feature extraction layer 2210 performs a convolutional operation on the input data to extract portions of the image D11 that resemble the shape of the filter (weight). The convolutional layer is configured to apply an activation function to a feature map, which is the result of the operation. The pooling layer of the feature extraction layer 2210 summarizes the features of the image data D10 obtained by convolution into maximum and average values, so that the extracted features are considered to be the same even if their positions change. The feature extraction layer 2210 learns the optimal output, so by increasing the number of convolutional layers and pooling layers, it is possible to extract more advanced and complex features. The combination layer 2220 combines the features extracted by multiple feature extraction layers 2210 and outputs them to the output layer 2300.

[0045] The output layer 2300 estimates the traffic objects 1200 in the image represented by the image data D10 based on the features extracted in the intermediate layer 2200 and the correct answer data D22. The output layer 2300 identifies the correct answer data D22 associated with features similar to the features output by the fully connected layer 2220, and outputs the positions, sizes, and types of the traffic objects 1200 indicated by the correct answer data D22, as well as the number of estimated traffic objects 1200.

[0046] The second machine learning unit 254 performs machine learning using the plurality of teacher data 242 (second teacher data) acquired by the second acquisition unit 253, thereby determining weights and the like of the intermediate layer 2200 and setting them in the CNN so as to estimate traffic objects 1200 in image D11 indicated by the input image data D10, thereby generating an object estimation model M2. The second machine learning unit 254 stores the generated object estimation model M2 in the storage unit 240. As a result, when image data D10 is input, the object estimation model M2 can output results regarding the positions, sizes, and types of the traffic objects 1200 indicated by the image data D10, for the number of estimated traffic objects 1200.

[0047] An example of the functional configuration of the learning device 200 according to this embodiment has been described above. Note that the configuration described above using Fig. 3 is merely an example, and the functional configuration of the learning device 200 according to this embodiment is not limited to this example. The functional configuration of the learning device 200 according to this embodiment can be flexibly modified according to specifications and operations.

[0048] (State Estimation Device) Fig. 5 is a diagram illustrating an example of the configuration of a state estimation device 100 according to an embodiment. As illustrated in Fig. 5, the state estimation device 100 includes an input unit 110, a communication unit 120, a storage unit 130, and a control unit 140. The control unit 140 is electrically connected to the input unit 110, the communication unit 120, the storage unit 130, etc.

[0049] The input unit 110 receives image data D10 captured by the imaging device 10. The input unit 110 has a connector that can be electrically connected to the imaging device 10 via a cable, for example. The input unit 110 supplies the image data D10 input from the imaging device 10 to the control unit 140.

[0050] The communication unit 120 can communicate with, for example, the learning device 200, a management device that manages the imaging device 10, etc. The communication unit 120 can support various communication standards. The communication unit 120 can send and receive various information via, for example, a wired or wireless network. The communication unit 120 can supply the received data to the control unit 140. The communication unit 120 can send data to a destination instructed by the control unit 140.

[0051] The storage unit 130 can store programs and data. The storage unit 130 is also used as a working area for temporarily storing processing results of the control unit 140. The storage unit 130 may include any non-transitory storage medium, such as a semiconductor storage medium or a magnetic storage medium. The storage unit 130 may include multiple types of storage media. The storage unit 130 may include a combination of a portable storage medium, such as a memory card, an optical disk, or a magneto-optical disk, and a storage medium reader. The storage unit 130 may include a storage device, such as RAM, that is used as a temporary storage area.

[0052] The storage unit 130 can store, for example, a program 131, setting data 132, image data D10, a state estimation model M1, an object estimation model M2, etc. The program 131 causes the control unit 140 to execute functions related to various controls for operating the state estimation device 100. The setting data 132 includes data such as various settings related to the operation of the state estimation device 100 and settings related to the installation state of the imaging device 10 to be managed. The storage unit 130 can store multiple pieces of image data D10 in chronological order. The state estimation model M1 and the object estimation model M2 are machine learning models generated by the learning device 200.

[0053] The control unit 140 is a processing unit. Examples of the processing unit include, but are not limited to, a CPU, an SoC, an MCU, an FPGA, and a coprocessor. The control unit 140 comprehensively controls the operation of the state estimation device 100 to realize various functions.

[0054] Specifically, the control unit 140 executes instructions contained in the program 131 stored in the storage unit 130 while referring to the data stored in the storage unit 130 as necessary. The control unit 140 then controls the functional units in accordance with the data and instructions, thereby realizing various functions. The functional units include, but are not limited to, the input unit 110 and the communication unit 120, for example.

[0055] The control unit 140 has functional units such as a processing unit 141, an estimation unit 142, and a diagnosis unit 143. The control unit 140 realizes the functional units such as the processing unit 141, the estimation unit 142, and the diagnosis unit 143 by executing the program 131. The program 131 is a program for causing the control unit 140 of the state estimation device 100 to function as the processing unit 141, the estimation unit 142, and the diagnosis unit 143.

[0056] The processing unit 141 acquires image data D10 captured by the imaging device 10. The processing unit 141 preprocesses the image data D10 to be used by the estimation unit 142 and supplies the preprocessed image data D10 to the estimation unit 142. The processing unit 141 processes the image data D10 of the traffic environment 1000 so that it includes traffic objects 1200 that can be used to estimate installation state parameters. The traffic objects 1200 that can be used to estimate installation state parameters include, for example, traffic objects 1200 facing the imaging device 10 and traffic objects 1200 included in the machine learning training data 242. The processing unit 141 processes the image data D10 so that the image includes traffic objects 1200 that can be used to estimate the installation state of the imaging device 10. The traffic objects 1200 that can be used for estimation include vehicles, people, and the like that appear in a manner suitable for estimating installation state parameters. Traffic objects 1200 that can be used for estimation include, for example, vehicles or people present in a predetermined area D100 of the image D11, and vehicles or people heading toward the imaging device 10. In this embodiment, the predetermined area D100 includes, for example, a predetermined area in the image D11, a central area of ​​the image D11, etc. Traffic objects 1200 that are not suitable for estimation include, for example, large vehicles such as freight trucks, buses, and construction machinery present in the predetermined area D100 of the image D11 shown by the image data D10. The processing unit 141 processes the image data D10 so as to include at least one of the traffic objects 1200 present in the predetermined area D100 and the traffic objects 1200 facing forward toward the imaging device 10.

[0057] The processing unit 141 estimates at least one of the position, size, and type of the traffic object 1200 in the traffic environment 1000 indicated by the input image data D10, using the object estimation model M2 generated by the learning device 200. Based on the estimation result of the object estimation model M2, the processing unit 141 processes the image data D10 of the captured traffic environment 1000 so that it includes the traffic object 1200 that can be used for estimation.

[0058] The processing unit 141 can provide a function for processing the image data D10 so as to delete or change the traffic event 1200 unnecessary for estimating the installation state parameters of the imaging device 10 from the image D11. The processing unit 141 can provide a function for selecting, from multiple pieces of image data D10 captured in chronological order, image data D10 that includes the traffic event 1200 used for estimating the installation state parameters of the imaging device 10 but does not include traffic events 1200 unnecessary for the estimation. The processing unit 141 can provide a function for adding the traffic event 1200 usable for estimating the installation state parameters to a preset setting area of ​​the image data D10. The setting area can be set as appropriate to, for example, the entire area of ​​the predetermined area D100, a partial area of ​​the predetermined area D100, or the like. The processing unit 141 can provide a function for determining the movement direction of the traffic event 1200 based on multiple pieces of image data D10 captured in chronological order, and processing the image data D10 so that the traffic event 1200 used for estimation is included based on the determination result.

[0059] The estimation unit 142 can provide a function of estimating installation state parameters of the imaging device 10 that captured input image data D10 using the state estimation model M1 generated by the learning device 200. The estimation unit 142 can input the image data D10 processed by the processing unit 141 to the state estimation model M1 and estimate the installation state parameters of the imaging device 10 based on the output of the state estimation model M1. The estimation unit 142 can estimate a road area corresponding to the road plane in the image D11 based on the estimated installation state parameters.

[0060] The diagnosis unit 143 can provide a function of diagnosing the installation state of the imaging device 10 based on the installation state parameters estimated by the estimation unit 142. The diagnosis unit 143 can diagnose whether the estimation results of the installation state parameters are appropriate. The diagnosis unit 143 can diagnose the installation state of the imaging device 10 based on the installation state parameters estimated by the estimation unit 142 and the overhead view state of the traffic object 1200 indicated by the image data D10. The diagnosis unit 143 compares the attitude of the traffic object 1200 indicated by the image data D10 with the attitude of the traffic object 1200 calculated based on the installation state parameters estimated by the estimation unit 142, and can diagnose the installation state of the imaging device 10 if the degree of match is higher than a determination threshold. The diagnosis unit 143 can compare the installation state parameters estimated by the estimation unit 142 with preset installation state parameters and diagnose the installation state of the imaging device 10 based on the comparison result.

[0061] The control unit 140 can provide a function of supplying the installation state parameters estimated by the estimation unit 142, the diagnosis results of the diagnosis unit 143, etc. to an external device, a database, etc. For example, the control unit 140 controls the supply of the installation state parameters estimated by the estimation unit 142, the diagnosis results of the diagnosis unit 143, etc. via the communication unit 120.

[0062] An example of the functional configuration of the state estimation device 100 according to this embodiment has been described above. Note that the above configuration described using Fig. 5 is merely an example, and the functional configuration of the state estimation device 100 according to this embodiment is not limited to this example. The functional configuration of the state estimation device 100 according to this embodiment can be flexibly modified according to specifications and operations.

[0063] In this embodiment, the state estimation device 100 will be described as having a control unit 140 functioning as a processing unit 141, an estimation unit 142, and a diagnosis unit 143. However, for example, the control unit 140 may be configured to include the estimation unit 142 and the diagnosis unit 143, but not the processing unit 141. In this case, the state estimation device 100 may input image data D10 captured by the imaging device 10 to the state estimation model M1 without preprocessing the image data D10. Furthermore, the system 1 may have the processing unit 141 of the state estimation device 100 as a component of the imaging device 10.

[0064] Fig. 6 is a flowchart showing an example of a state estimation method executed by the state estimation device 100. When diagnosing the installation state of the imaging device 10, the state estimation device 100 sequentially executes a preprocessing step S100, a state estimation step S200, and a result diagnosis step S300 shown in Fig. 6. The state estimation device 100 executes the method shown in Fig. 6 at execution times such as when the imaging device 10 is installed, when maintenance is performed, when an execution command is received from outside, etc.

[0065] The preprocessing step S100 is a step of processing the image data D10 so that the image includes a traffic object 1200 necessary for estimating the installation state of the imaging device 10. The preprocessing step S100 is implemented by the processing unit 141 of the control unit 140. The preprocessing step S100 processes the image data D10 so that an image D11 including a traffic object 1200 facing forward toward the imaging device 10 is preferentially supplied to the state estimation step S20. The preprocessing step S10 processes the image data D10 so that an image D11 including a large vehicle or the like in a predetermined region D100 of the image D11 is not supplied to the state estimation step S20. The preprocessing step S10 processes the image data D10 so as to process information about the traffic object 1200 in the image D11 in order to improve the reliability of estimating the installation state of the imaging device 10.

[0066] In this embodiment, the preprocessing step S100 includes an object detection step S100A and a tracking step S100B. The object detection step S100A is a step of detecting a traffic object 1200 (object) in the image D11. The tracking step S100B is a step of tracking the movement direction (orientation) of the traffic object 1200 in the image D11. Although the preprocessing step S100 will be described as including the object detection step S100A, the tracking step S100B, etc., it may also be configured to include, for example, only the object detection step S100A. In the preprocessing step S100, the state estimation device 100 may execute the tracking step S100B after executing the object detection step S100A, or may execute them in parallel.

[0067] 7 is a flowchart showing an example of the processing procedure of the object detection step S100A executed by the state estimation device 100. The processing procedure shown in FIG.

[0068] As shown in FIG. 7, the state estimation device 100 detects a traffic object 1200 from image data D10 (step S101). For example, the state estimation device 100 inputs image data D10 captured by the imaging device 10 into the object estimation model M2, and detects the traffic object 1200 included in the image data D10 based on the estimation result output by the object estimation model M2. The state estimation device 100 detects one or more traffic objects 1200 included in a predetermined area D100 of an image D11 represented by the image data D10, based on the position, size, and type of the traffic object output by the object estimation model M2. The state estimation device 100 detects that the traffic object 1200 is not included in the predetermined area D100 of the image D11, based on the output result of the object estimation model M2. After storing the estimation result of the object estimation model M2 and the detection result of the traffic object 1200 in the storage unit 130, the state estimation device 100 proceeds to step S102.

[0069] The state estimation device 100 determines whether the traffic object 1200 is included in the image D11 (step S102). For example, if the detection result in step S101 indicates that the traffic object 1200 has been detected, the state estimation device 100 determines that the traffic object 1200 is included in the image D11. If the state estimation device 100 determines that the traffic object 1200 is not included in the image D11 (No in step S102), the state estimation device 100 proceeds to step S103.

[0070] The state estimation device 100 determines whether the background image of the predetermined area D100 has been registered (step S103). For example, if the background image of the predetermined area D100 captured by the imaging device 10 has already been registered in the setting data 132, a database, or the like, the state estimation device 100 determines that the background image of the predetermined area D100, etc. has been registered. The background image is an image D11 that includes information for masking areas of large vehicles, standard-sized vehicles, etc. from the image data D10 and shows the traffic environment 1000 without including the traffic object 1200. If the state estimation device 100 determines that the background image of the predetermined area D100 has been registered (Yes in step S103), the state estimation device 100 returns the process to step S101, which has already been described, and continues the process. If the state estimation device 100 determines that the background image of the predetermined area D100 has not been registered (No in step S103), the state estimation device 100 proceeds to step S104.

[0071] The state estimation device 100 registers the background image of the image data D10 (step S104). For example, the state estimation device 100 registers an image D11 including a predetermined area D100 of the image data D10 in the setting data 132, a database, or the like as the background image of the imaging device 10. When the processing of step S104 ends, the state estimation device 100 returns the processing to step S101, which has already been described, and continues the processing.

[0072] Furthermore, if the state estimation device 100 determines that the traffic object 1200 is included in the image D11 (Yes in step S102), the state estimation device 100 proceeds to step S105. The state estimation device 100 determines whether a large vehicle is present in the predetermined area D100 of the image D11 (step S105). For example, if the type of the traffic object 1200 present in the predetermined area D100 of the image data D10 estimated by the object estimation model M2 is a large vehicle, a large special-purpose vehicle, or the like, the state estimation device 100 determines that a large vehicle is present in the predetermined area D100 of the image D11. If the state estimation device 100 determines that a large vehicle is not present in the predetermined area D100 of the image D11 (No in step S105), the state estimation device 100 proceeds to step S108, which will be described later.

[0073] Furthermore, if the state estimation device 100 determines that a large vehicle is present in the predetermined area D100 of the image D11 (Yes in step S105), the state estimation device 100 proceeds to step S106. As in step S103, the state estimation device 100 determines whether the background image of the predetermined area D100 has already been registered (step S106). If the state estimation device 100 determines that the background image of the predetermined area D100 has not already been registered (No in step S106), the state estimation device 100 returns to the process at step S101, which has already been described, and continues the process. If the state estimation device 100 determines that the background image of the predetermined area D100 has already been registered (Yes in step S106), the state estimation device 100 proceeds to step S107.

[0074] The state estimation device 100 masks the large vehicle in the image D11 (step S107). For example, the state estimation device 100 masks the large vehicle by replacing the area in the image showing the large vehicle with a registered background image. When the process of step S107 ends, the state estimation device 100 proceeds to step S108.

[0075] The state estimation device 100 determines whether a traffic object 1200 facing forward is present in the predetermined area D100 of the image D11 (step S108). For example, if the traffic object 1200 present in the predetermined area D100 of the image D11 represented by the image data D10 faces forward, the state estimation device 100 determines that the traffic object 1200 facing forward is present in the predetermined area D100 of the image D11. If the direction of movement of the traffic object 1200 in the predetermined area D100 tracked in the tracking step S100B is a direction toward the imaging device 10, the state estimation device 100 determines that the traffic object 1200 facing forward is present in the predetermined area D100 of the image D11. If the state estimation device 100 determines that the front-facing traffic object 1200 does not exist in the predetermined area D100 of the image D11 (No in step S108), the process returns to step S101, which has already been described, and continues. On the other hand, if the state estimation device 100 determines that the front-facing traffic object 1200 exists in the predetermined area D100 of the image D11 (Yes in step S108), the process proceeds to step S109.

[0076] The state estimation device 100 supplies the image data D10 to the estimation unit 142 (step S109). For example, the state estimation device 100 stores the image data D10, in which a traffic object 1200 facing forward is included in a predetermined area D100 of the image D11, in association with the imaging device 10 that captured the image data D10 and the estimation result, in the storage unit 130, thereby supplying the image data D10 to the estimation unit 142. Upon completing the processing of step S109, the state estimation device 100 ends the processing procedure shown in FIG. 7 and returns to the object detection step S100A of the preprocessing step S10 shown in FIG. 6.

[0077] After the object detection step S100A is completed, the state estimation device 100 executes a tracking step S100B. The state estimation device 100 traces the image data D10 back a predetermined time period and tracks the moving direction of the traffic object 1200 detected in the object detection step S100A.

[0078] Fig. 8 is a diagram illustrating an example of a tracking process for image data D10 including a traffic object 1200. In the example shown in Fig. 8, a traffic object 1200A and a traffic object 1200B are detected from the image data D10 by an object detection step S100A. The traffic object 1200 is a car facing forward with respect to the image capture device 10. The traffic object 1200B is a car facing away from the image capture device 10.

[0079] The state estimation device 100 executes a tracking step S100B to track the movement directions (orientations) of the traffic objects 1200A and 1200B. For example, the state estimation device 100 uses a tracking process using a known Kalman filter to obtain a trajectory L1 of the traffic object 1200A and a trajectory L2 of the traffic object 1200B from a plurality of consecutive image data D10. The trajectory L1 of the traffic object 1200A and the trajectory L2 of the traffic object 1200B continuously indicate the directions of the trajectories at times t1, t2, and t3.

[0080] 8, the imaging device 10 is installed so as to be able to capture an image of a two-lane road 1100 extending along a road direction M. Image data D10 indicates that the direction of trajectory L1 of a traffic object 1200A is the same as the road direction M, and that the traffic object 1200A faces forward relative to the imaging device 10. Image data D10 indicates that the direction of trajectory L2 of a traffic object 1200B is opposite to the road direction M, and that the traffic object 1200B does not face forward relative to the imaging device 10.

[0081] When tracking in the tracking step S100B is completed, the state estimation device 100 associates the movement direction of the tracked traffic object 1200 with the image data D10 and stores the image data D10 in the storage unit 130. This allows the state estimation device 100 to associate information about the orientation of the traffic object 1200 with the image data D10 supplied to the state estimation step S200, thereby assisting in estimating the installation state of the imaging device 10, which is capable of switching the imaging direction.

[0082] As shown in Fig. 6, upon completion of the preprocessing step S100, the state estimation device 100 executes a state estimation step S200. The state estimation step S200 includes a step of estimating installation state parameters of the imaging device 10 that captured input image data D10 using a machine-learned state estimation model M1 generated by the learning device 200. The state estimation step S200 estimates the road area of ​​the road 1100 in the image D11 represented by the image data D10, based on the installation state parameters estimated by the state estimation model M1. The state estimation device 100 associates the estimated installation state parameters and road area with the image data D10 and stores them in the storage unit 130.

[0083] Upon completion of the state estimation step S200, the state estimation device 100 executes a result diagnosis step S300. The result diagnosis step S300 includes a step of diagnosing whether the installation state parameters of the imaging device 10 estimated in the state estimation step S200 are appropriate. Diagnosing whether the installation state parameters of the imaging device 10 are appropriate includes diagnosing that the installation state parameters of the imaging device 10 are appropriate if the installation state parameters do not require resetting, adjustment, or the like of the imaging device 10. The state estimation device 100 associates the diagnosis result with the imaging device 10 and stores it in the storage unit 130. If the state estimation device 100 diagnoses that the installation state parameters of the imaging device 10 are appropriate, it can supply the diagnosis result, the installation state parameters, and the like to subsequent processing. Furthermore, if the state estimation device 100 diagnoses that the installation state parameters of the imaging device 10 are inappropriate, it can capture an image with the imaging device 10 again and estimate the installation state parameters using the captured image data D10.

[0084] 9 to 11 are diagrams illustrating an example of estimation of installation state parameters by the state estimation device 100 according to the embodiment. In a scene ST10 shown in FIGS. 9 to 11, image data D10-0 of a traffic environment 1000 captured by the imaging device 10 is data representing a single image D11 having a road 1100 and traffic objects 1200A, 1200B, and 1200C. The image data D10-0 is data before the state estimation device 100 executes the preprocessing step S100. The traffic object 1200A is a car facing forward with respect to the imaging device 10. The traffic object 1200B is a large car facing forward with respect to the imaging device 10. The traffic object 1200C is a car facing away from the imaging device 10.

[0085] 9, when the state estimation device 100 acquires image data D10-0 captured by the imaging device 10, it executes a preprocessing step S100 on the image data D10-0. The state estimation device 100 inputs the image data D10-0 to the object estimation model M2, and determines that three objects, traffic objects 1200A, 1200B, and 1200C, are included in the image D11 based on the estimation result of the object estimation model M2.

[0086] The state estimation model M1 according to this embodiment may have reduced accuracy in estimating the installation state parameters of the image capture device 10 when a large vehicle, construction machinery, or the like is present near the center of the image D11. Therefore, in scene ST11, the state estimation device 100 performs a process of masking a portion of the image D11 showing the large vehicle traffic object 1200B facing forward with a background image, and supplies the resulting image data D10 to the state estimation step S200. The state estimation device 100 inputs the processed image data D10 to the state estimation model M1, and obtains the installation state parameters of the image capture device 10 estimated by the state estimation model M1. Based on the obtained installation state parameters, the state estimation device 100 estimates a road region D500 of the road 1100 in the image D11 shown by the image data D10, and stores the image data D10, in which the road region D500 is added to the image D11, in the storage unit 130. This allows the state estimation device 100 to input image data D10, which includes the traffic event 1200A that can be used for estimation and in which the traffic event 1200B that is not required for estimation has been deleted from the image D11, to the state estimation model M1. As a result, the state estimation device 100 can improve the estimation accuracy of the installation state parameters compared to inputting unprocessed image data D10-0 to the state estimation model M1.

[0087] In the present embodiment, a case will be described in which the state estimation device 100 deletes the unnecessary traffic object 1200 from the image D11, but the present invention is not limited to this. The traffic object 1200 may disappear from the image capture area of ​​the image capture device 10 as it moves. For this reason, instead of deleting the unnecessary traffic object 1200, the state estimation device 100 may perform a process of extracting image data D10 that does not include the unnecessary traffic object 1200 from multiple image data D10 captured in chronological order.

[0088] Furthermore, in the scene ST10 shown in FIG. 10 , even if the traffic object 1200 that is not facing forward and is located near the edge of the image D11 is deleted, the accuracy with which the state estimation model M1 estimates the installation state parameters of the image capture device 10 may not change. Therefore, in the scene ST12 shown in FIG. 10 , the state estimation device 100 performs a process of masking the traffic object 1200B from the image data D10-0 shown in the scene ST10 and masking the portion of the image D11 showing the traffic object 1200C located near the left edge with a background image. The state estimation device 100 supplies the processed image data D10 to the state estimation step S200. In this case, the state estimation device 100 inputs the processed image data D10 to the state estimation model M1 and obtains the installation state parameters of the image capture device 10 estimated by the state estimation model M1. The state estimation device 100 estimates a road area D500 of the road 1100 in an image D11 represented by image data D10 based on the obtained installation state parameters, and stores image data D10 in which the road area D500 is added to image D11 in the storage unit 130. In this way, the state estimation device 100 can input, to the state estimation model M1, image data D10 that includes a traffic event 1200A that can be used for estimation and in which traffic events 1200B and 1200C that are not required for estimation have been deleted from image D11. As a result, the state estimation device 100 can improve the estimation accuracy of the installation state parameters compared to inputting unprocessed image data D10-0 to the state estimation model M1.

[0089] Furthermore, in an example shown in scene ST10 in FIG. 11 , the state estimation model M1 improves the accuracy of estimating the installation state parameters of the image capture device 10 when multiple vehicles facing forward are present near the center of image D11. Therefore, in scene ST13 shown in FIG. 11 , the state estimation device 100 performs a process of adding an image of a traffic object 1200D to a set region D140 near the center of image D11 for image data D10-0 shown in scene ST10, and supplies the image data D10 to the state estimation step S200. The image to be added to the set region D140 includes, for example, an image of the traffic object 1200 captured by the image capture device 10 in the past, an image of the traffic object 1200 pre-registered in a database, etc. The state estimation device 100 acquires an image of the traffic object 1200D from the database, past image data D10, etc., and masks the portion of the road 1100 indicated by the image data D10-0 in the traffic object image. This allows the state estimation device 100 to input, to the state estimation model M1, image data D10 that includes the traffic event 1200A that can be used for estimation and the added traffic event 1200D and that has deleted the traffic events 1200B and 1200C that are not required for estimation from the image D11. As a result, the state estimation device 100 can improve the estimation accuracy of the installation state parameters compared to inputting unprocessed image data D10-0 to the state estimation model M1.

[0090] Fig. 12 is a diagram illustrating another example of a result diagnosis performed by the state estimation device 100. The state estimation device 100 shown in Fig. 12 can use an object estimation model M2 that has been machine-learned to further estimate the orientation of a traffic object 1200 in a traffic environment 1000 represented by input image data D10.

[0091] The object estimation model M2 is a learning model generated by using the image data D10 and correct value data D22 contained in the teacher data 242 to extract features, regularities, patterns, bird's-eye view states, etc. of the traffic object 1200 in the image data D10, and by machine learning the relationship with the correct value data D22. The bird's-eye view state of the traffic object 1200 includes the object position (θx, θy), height, etc. of the traffic object 1200 when the traffic object 1200 is imaged from an angle looking down from the imaging device 10. When image data D10 is input, the object estimation model M2 predicts data similar to the image data D10 from the teacher data 242 and compares it with the correct value data D22 to estimate the position, size, type, bird's-eye view state, etc. of the traffic object 1200 (object) in the image D11 represented by the image data D10, and outputs the estimation result.

[0092] In the state estimation device 100, the processing unit 141 inputs image data D10 to the object estimation model M2, and based on the estimation result output by the object estimation model M2, the bird's-eye view state of the traffic object 1200 in the image D11 is stored in the storage unit 130. After the processing unit 141 processes the image data D10, the state estimation device 100 inputs the image data D10 to the state estimation model M1 in the estimation unit 142. In the state estimation device 100, the estimation unit 142 obtains installation state parameters of the imaging device 10 from the estimation result output by the state estimation model M1.

[0093] The state estimation device 100 applies the estimated installation state parameters to a calculation formula, a conversion table, etc., to calculate the orientation of the traffic object 1200 in the image D11. The state estimation device 100 compares, in the diagnosis unit 143, the orientation estimated from the actual appearance of the traffic object 1200 in the image data D10 with the overhead view of the traffic object 1200, and calculates the error between the state of the traffic object 1200 when the imaging device 10 captures the image using the installation state parameters and the overhead view. The setting data 132 in the storage unit 130 includes a judgment threshold for determining the error of the estimation result. The judgment threshold is a threshold set for determining that the error in reliability is small. If the calculated error is equal to or greater than the judgment threshold set in the setting data 132, the state estimation device 100 determines that the reliability of the estimation result by the estimation unit 142 is low. In this case, the state estimation device 100 does not supply the installation state parameters of the imaging device 10 estimated from the image data D10 to subsequent processing. The subsequent processing includes, for example, processing related to installation, maintenance, etc. of the imaging device 10 based on the installation state parameters. The state estimation device 100 acquires new image data D10 from the imaging device 10 as needed, and estimates the installation state parameters of the imaging device 10 using the image data D10.

[0094] Furthermore, if the calculated error is smaller than the determination threshold set in the setting data 132, the state estimation device 100 determines that the reliability of the estimation result of the estimation unit 142 is high. In this case, the state estimation device 100 supplies the installation state parameters of the imaging device 10 estimated from the image data D10 to subsequent processing. This eliminates the need for a jig used to calculate the installation state of the imaging device 10, and allows the state estimation device 100 to estimate the installation state parameters using the image data D10 captured by the imaging device 10. As a result, the state estimation device 100 eliminates the need for road regulation work such as installing and checking the imaging device 10, thereby contributing to the spread of the system 1.

[0095] Although the above-described state estimation device 100 is provided outside the imaging device 10, the present invention is not limited to this. For example, the state estimation device 100 may be incorporated into the imaging device 10 and realized by a control unit, module, or the like of the imaging device 10. For example, the state estimation device 100 may be incorporated into a traffic light, lighting equipment, communication equipment, or the like installed in the traffic environment 1000.

[0096] The above-described state estimation device 100 may be realized by a server device, etc. For example, the state estimation device 100 may be a server device that acquires image data D10 from each of the multiple imaging devices 10, estimates installation state parameters from the image data D10, and provides the estimation results.

[0097] The above-described learning device 200 generates a state estimation model M1 and an object estimation model M2, but is not limited to this. For example, the learning device 200 may be configured with two devices: a first device that generates the state estimation model M1 and a second device that generates the object estimation model M2.

[0098] Furthermore, the present disclosure may be applied not only to cases where the state estimation model M1 and the object estimation model M2 are implemented using separate models and separate learning units, but also to an embodiment where both models are combined into an integrated model and machine learning is performed using a single integrated machine learning unit. In other words, the present disclosure may also include an embodiment where a single model and a single learning unit are used.

[0099] Although specific embodiments have been described to fully and clearly disclose the claimed technology, the appended claims should not be limited to the above-described embodiments, but should be construed to embody all modifications and alternative arrangements that may be made by those skilled in the art within the scope of the basic concept presented herein.

[0100] [Note] (Appendix 1: State inference section + diagnosis section) an estimation unit that estimates the installation state parameters of the imaging device that captured the input image data using a state estimation model that has been machine-learned to estimate the installation state parameters of the imaging device that captured the input image data, using first teacher data that includes image data of a traffic environment captured by an imaging device and first correct answer value data of installation state parameters of the imaging device that captured the image data; a diagnosis unit that diagnoses the installation state of the imaging device based on the estimated installation state parameters; A state estimation device comprising: (Appendix 2: Appendix 1 + preprocessing section) 2. The state estimation device according to claim 1, a processing unit that processes the image data of the traffic environment captured by the imaging device so that the image data includes traffic objects that can be used for estimation; The estimation unit inputs the image data processed by the processing unit into the state estimation model and estimates the installation state parameters of the imaging device. State estimator. (Appendix 3) 3. The state estimation device according to claim 2, The processing unit performs processing so that the image data includes at least one of an image in which the traffic object is present in a predetermined area and an image in which the traffic object is facing the imaging device. State estimator. (Note 4: The preprocessing unit uses the second machine learning method.) 4. The state estimation device according to claim 2, The processing unit uses second teacher data including the image data and second correct answer data for object detection in the image data to perform machine learning on an object estimation model to estimate at least one of the position, size, and type of the traffic object in the traffic environment indicated by the input image data, and processes the image data of the traffic environment captured by the imaging device so as to include the traffic object used for estimation, based on an estimation result of at least one of the position, size, and type of the traffic object in the traffic environment indicated by the input image data. State estimator. (Appendix 5) 5. The state estimation device according to claim 4, the object estimation model is a model that has been machine-learned to further estimate the orientation of the traffic object in the traffic environment indicated by the input image data, The processing unit processes the image data so that the image data includes the traffic object in a direction approaching the imaging device using the object estimation model. State estimator. (Appendix 6) 6. The state estimation device according to claim 2, The processing unit processes the image data so as to delete or change the traffic object from the image, the traffic object being unnecessary for estimating the installation state parameters of the imaging device. State estimator. (Appendix 7) 7. The state estimation device according to claim 2, The traffic object that is not necessary for estimating the installation state parameters of the imaging device is at least one of a freight vehicle, a passenger car, and a construction machine that is present in a predetermined area of ​​the image shown by the image data. State estimator. (Appendix 8) 8. The state estimation device according to claim 2, The processing unit selects, from the plurality of image data captured in chronological order, image data that includes the traffic object used to estimate the installation state parameters of the imaging device and does not include the traffic object unnecessary for the estimation. State estimator. (Appendix 9) 9. The state estimation device according to claim 2, The processing unit adds the traffic object that can be used to estimate the installation state parameters to a preset setting area of ​​the image data. State estimator. (Appendix 10) 10. The state estimation device according to claim 2, The war history processing unit determines the direction of movement of the traffic object based on the plurality of image data captured in chronological order, and processes the image data based on the determination result so that the traffic object used for estimation is included. State estimator. (Appendix 11) 11. The state estimation device according to claim 10, The processing unit determines the direction of movement of the traffic object using a tracking process, and processes the image data based on the determination result so that the image data includes the traffic object used for estimation. State estimator. (Appendix 12) 12. The state estimation device according to claim 2, The diagnosing unit diagnoses an installation state of the imaging device based on the installation state parameters estimated by the estimating unit and an overhead view state of the traffic object indicated by the image data. State estimator. (Appendix 13) 13. The state estimation device according to claim 12, The diagnosing unit compares the posture of the traffic object indicated by the image data with the posture of the traffic object calculated based on the installation state parameters estimated by the estimating unit, and diagnoses the installation state of the imaging device when the degree of coincidence is higher than a judgment threshold. State estimator. (Appendix 14) 14. The state estimation device according to claim 13, When the degree of coincidence is equal to or less than a determination threshold, the diagnosis unit causes the estimation unit to estimate the installation state parameters based on the next image data captured by the imaging device. State estimator. (Appendix 15: Hereafter, a separate appendix on machine learning in the installation state) a first acquisition unit that acquires first teacher data having image data of a traffic environment including traffic objects and first correct value data of installation state parameters of an imaging device that captured the image data; a first machine learning unit that generates, by machine learning using the first teacher data, a state estimation model that estimates the installation state parameters of the imaging device that captured the input image data; A learning device comprising: (Appendix 16) 16. The learning device according to claim 15, The first correct value data includes a combination of at least two of the installation angle, height, and installation position of the imaging device. Learning device. (Appendix 17: Machine learning for object detection) 17. The learning device according to claim 15 or 16, a second acquisition unit that acquires second teacher data including image data of the traffic environment captured by the imaging device and second correct answer data for object detection in the image data; a second machine learning unit that generates, by machine learning using the second teacher data, an object estimation model that estimates at least one of the position, size, and type of the traffic object in the traffic environment indicated by the input image data; Equipped with The second correct answer data includes data indicating at least one of the position, size, and type of the traffic object and the number of the traffic objects. Learning device. (Appendix 18: Separate appendix on methods) The computer estimating the installation state parameters of the imaging device that captured the input image data using a state estimation model that has been machine-learned to estimate the installation state parameters of the imaging device that captured the input image data, using first teacher data that includes image data of an image of a traffic environment and first correct answer value data of installation state parameters of the imaging device that captured the image data; diagnosing the installation state of the imaging device based on the estimated installation state parameters; A state estimation method comprising: (Appendix 19: Program-specific appendix) On the computer, estimating the installation state parameters of the imaging device that captured the input image data using a state estimation model that has been machine-learned to estimate the installation state parameters of the imaging device that captured the input image data, using first teacher data that includes image data of an image of a traffic environment and first correct answer value data of installation state parameters of the imaging device that captured the image data; diagnosing the installation state of the imaging device based on the estimated installation state parameters; A state estimation program that executes the above. [Explanation of symbols]

[0101] 1 System 10. Imaging device 100 State Estimation Device 110 Input section 120 Communications Department 130 Storage section 131 Programs 132 Setting data 140 Control Unit 141 Processing section 142 Estimation Department 143 Diagnostic Department 200 Learning Device 210 Display section 220 Operation section 230 Communications Department 240 Storage section 241 Programs 242 training data 250 control section 251 First acquisition part 252 Machine Learning Department 1 253 Second Acquisition Department 254 Second Machine Learning Department 1000 Traffic environment 1100 Road 1200 Transportation 2100 input layer 2200 Middle Class 2210 Feature Extraction Layer 2220 bonding layer 2300 output layer D10 image data D11 Images D21 Correct answer data D22 Correct answer data D100 Predetermined area M1 State estimation model M2 object estimation model

Claims

1. an estimation unit that estimates the installation state parameters of the imaging device that captured the input image data using a state estimation model that has been machine-learned to estimate the installation state parameters of the imaging device that captured the input image data, using first teacher data that includes image data of a traffic environment captured by an imaging device and first correct answer value data of installation state parameters of the imaging device that captured the image data; a diagnosis unit that diagnoses the installation state of the imaging device based on the estimated installation state parameters; Equipped with The diagnosis of the installation state of the imaging device by the diagnosing unit diagnoses whether the installation state parameters are appropriate.

2. 2. The state estimation device according to claim 1, a processing unit that processes the image data of the traffic environment captured by the imaging device so that the image data includes traffic objects that can be used for estimation; The estimation unit inputs the image data processed by the processing unit into the state estimation model and estimates the installation state parameters of the imaging device. State estimator.

3. 3. The state estimation device according to claim 2, The processing unit performs processing so that the image data includes at least one of an image in which the traffic object is present in a predetermined area and an image in which the traffic object is facing the imaging device. State estimator.

4. 3. The state estimation device according to claim 2, The processing unit uses second teacher data including the image data and second correct answer data for object detection in the image data to perform machine learning on an object estimation model to estimate at least one of the position, size, and type of the traffic object in the traffic environment indicated by the input image data, and processes the image data of the traffic environment captured by the imaging device so as to include the traffic object used for estimation, based on an estimation result of the position, size, and type of the traffic object in the traffic environment indicated by the input image data. State estimator.

5. 5. The state estimation device according to claim 4, the object estimation model is a model that has been machine-learned to further estimate the orientation of the traffic object in the traffic environment indicated by the input image data, The processing unit processes the image data so that the image data includes the traffic object in a direction approaching the imaging device using the object estimation model. State estimator.

6. 3. The state estimation device according to claim 2, The processing unit processes the image data so as to delete or change the traffic object from the image, the traffic object being unnecessary for estimating the installation state parameters of the imaging device. State estimator.

7. 3. The state estimation device according to claim 2, The traffic object that is not necessary for estimating the installation state parameters of the imaging device is at least one of a freight vehicle, a passenger car, and a construction machine that is present in a predetermined area of ​​the image shown by the image data. State estimator.

8. 3. The state estimation device according to claim 2, The processing unit selects, from the plurality of image data captured in chronological order, image data that includes the traffic object used to estimate the installation state parameters of the imaging device and does not include the traffic object unnecessary for the estimation. State estimator.

9. 3. The state estimation device according to claim 2, The processing unit adds the traffic object that can be used to estimate the installation state parameters to a preset setting area of ​​the image data. State estimator.

10. 3. The state estimation device according to claim 2, The war history processing unit determines the direction of movement of the traffic object based on the plurality of image data captured in chronological order, and processes the image data based on the determination result so that the traffic object used for estimation is included. State estimator.

11. The state estimation device according to claim 10, The processing unit determines the direction of movement of the traffic object using a tracking process, and processes the image data based on the determination result so that the image data includes the traffic object used for estimation. State estimator.

12. 3. The state estimation device according to claim 2, The diagnosing unit diagnoses an installation state of the imaging device based on the installation state parameters estimated by the estimating unit and an overhead view state of the traffic object indicated by the image data. State estimator.

13. 13. The state estimation device according to claim 12, The diagnosing unit compares the posture of the traffic object indicated by the image data with the posture of the traffic object calculated based on the installation state parameters estimated by the estimating unit, and diagnoses the installation state of the imaging device when the degree of coincidence is higher than a judgment threshold. State estimator.

14. The computer estimating the installation state parameters of the imaging device that captured the input image data using a state estimation model that has been machine-learned to estimate the installation state parameters of the imaging device that captured the input image data, using first teacher data that includes image data of an image of a traffic environment and first correct answer value data of installation state parameters of the imaging device that captured the image data; diagnosing the installation state of the imaging device based on the estimated installation state parameters; Including, The state estimation method includes diagnosing whether the installation state of the imaging device is appropriate or not.

15. On the computer, estimating the installation state parameters of the imaging device that captured the input image data using a state estimation model that has been machine-learned to estimate the installation state parameters of the imaging device that captured the input image data, using first teacher data that includes image data of an image of a traffic environment and first correct answer value data of installation state parameters of the imaging device that captured the image data; diagnosing the installation state of the imaging device based on the estimated installation state parameters; Execute The diagnosis of the installation state of the imaging device is a state estimation program that diagnoses whether the installation state parameters are appropriate.

Citation Information

Patent Citations

  • Method for distress and road rage detection

    CN111741884A

  • Building stair pedestrian volume estimation system based on multi-dimensional MEMS inertial sensor

    CN112418649A

  • Monitoring camera parameter calibration method and device

    CN112950725A

  • Monitoring system, device of vehicle, device of roadside unit, traffic infrastructure system and method thereof

    CN114627638A

  • Visibility estimation device and method, and recording medium

    CN117291865A