System, program, apparatus, and method for estimating regional risk values for hazards based on images captured by moving body
The system uses mobile object cameras and learning models to estimate and aggregate risk values, addressing the limitations of existing technologies by providing real-time danger assessments for each region.
Patent Information
- Application Number
- JP2024039781
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-09-29
AI Technical Summary
Existing technologies are limited in detecting traffic risks to the area around the vehicle and do not account for changing traffic volumes over time, making it difficult to estimate the ever-changing degree of danger such as congestion, accidents, and incidents for each area of a region.
A system and method that utilizes a mobile object with a camera and positioning unit to capture images, estimate abnormality levels, aggregate these levels by area, and calculate risk values using learning models and population data to provide real-time risk assessments.
Enables precise estimation of changing dangers like traffic congestion and accidents across regions, aiding in route planning and risk avoidance.
Smart Images

Figure 2025140401000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for estimating a risk value of danger for each area (range) of a region from an image captured by a moving body. [Background technology]
[0002] Estimating the risk value of danger for moving vehicles and pedestrians in each area is important for ensuring public safety. There is a conventional behavior monitoring technology that uses images captured by an in-vehicle camera to improve the accuracy of detecting traffic risks based on the line of sight and relationships between a target person and surrounding people (see, for example, Patent Document 1). This technology detects the behaviors of multiple people captured in images captured by the in-vehicle camera, and detects attribute information for each person and correlation information that indicates the relationships between the people. Then, the behavior of each person is evaluated based on the attribute information and correlation information.
[0003] There is also a technology that, in the event of a disaster, updates a hazard map based on the traffic volume at the time of the disaster to support the decision-making of road administrators (see, for example, Patent Document 2). This technology acquires traffic-related state quantities and hazard map information, calculates the degree of influence of the state quantities and the hazard map information when a disaster occurs, and then calculates the possibility of further disasters occurring. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2020-095357 [Patent Document 2] Japanese Patent Application Publication No. 2018-092467 [Non-patent literature]
[0005] [Non-Patent Document 1] Xingchao Peng, Zijun Huang, Yizhe Zhu, Kate Saenko, Federated Adversarial Domain Adaptation, [online], [Retrieved March 12, 2024], Internet<URL:https: / / arxiv.org / abs / 1911.02054> [Non-patent document 2] Characteristics and History of Regional Mesh Statistics, [online], [searched February 14, 2024], Internet<URL:https: / / www.stat.go.jp / data / mesh / pdf / gaiyo1.pdf> [Non-patent document 3] KDDI Location Data, [online], [searched March 12, 2024], Internet<URL:https: / / k-locationdata.kddi.com / > [Non-patent document 4] Mu2ReST, Multi-resolution Recursive Spatio-Temporal Transformer for Long-Term Prediction, [online], [Retrieved March 12, 2024], Internet<URL:https: / / dl.acm.org / doi / abs / 10.1007 / 978-3-031-05933-9_6> Summary of the Invention [Problem to be solved by the invention]
[0006] According to the above-mentioned Patent Document 1, the area in which traffic risks can be detected is limited to the area around the vehicle equipped with the on-board camera. Furthermore, according to the above-mentioned Patent Document 2, since only the traffic volume at a single point in time is reflected, it is not possible to follow changes in traffic volume.
[0007] In response to this, the inventors of the present application wondered whether it would be possible to clearly indicate to a user or operator in advance the ever-changing degree of danger, such as congestion, accidents, and incidents, for each area of a region. For example, they wondered whether it would be possible to collect images taken by an in-vehicle camera and estimate a risk value for each area of a region. In this case, it is preferable to be able to estimate a risk value for danger with as high a degree of certainty as possible, even for areas where only a small number of images have been collected.
[0008] Therefore, the present invention aims to provide a system, program, device, and method that can estimate the risk value of dangers that change from moment to moment, such as traffic congestion, accidents, and incidents, for each area of a region from images captured by a camera mounted on a mobile object. [Means for solving the problem]
[0009] According to the present invention, there is provided a system including a mobile object and an estimation device that estimates a risk value of a danger for each area p, the system comprising: The moving body is a captured image database that stores images captured by a camera mounted on the vehicle in association with the image capturing position; an abnormality degree estimation means for estimating an abnormality degree for each captured image in the captured image database using a learning model previously trained with teacher data including captured images labeled normal and captured images labeled abnormal; an abnormality level transmitting means for transmitting the abnormality level and the photographing position for each photographed image to the estimation device; and The estimation device an area aggregation means for aggregating the abnormality level of each captured image received from a moving object for each area p including the image capturing position; a risk value calculation means for calculating a statistical value of the degree of abnormality for each area p as a risk value v(p) for the mobile object; The present invention is characterized by having the following.
[0010] According to another embodiment of the system of the present invention, For a moving object, the captured image of the camera is an image or a video image for a predetermined period of time. It is also preferable.
[0011] According to another embodiment of the system of the present invention, The mobile object is a vehicle traveling on a road. The images were taken by a camera mounted on the vehicle. The shooting position was measured by the positioning unit installed in the vehicle. Or, The mobile object is a mobile device held by a person, The photographed image is taken by a camera mounted on a mobile device. The shooting position was measured by the positioning unit installed in the mobile device. It is also preferable.
[0012] According to another embodiment of the system of the present invention, Regarding the estimation device, A map display means for dividing the map into areas p and displaying the risk value v(p) for each area p is provided. It is also preferable to have
[0013] According to another embodiment of the system of the present invention, The abnormality degree estimation means in the mobile object estimates the abnormality degree as normal (0) or abnormal (1), The risk value calculation means in the estimation device calculates a ratio of the number of images captured during abnormality to the number of all images captured in the area p as a statistical value of the degree of abnormality for each area p, Or, The abnormality degree estimation means in the mobile object estimates a value within a predetermined numerical range as the abnormality degree. It is also preferable.
[0014] According to another embodiment of the system of the present invention, The photographed image database for the moving object further associates each photographed image with a photographed time, The risk value calculation means in the estimation device calculates a risk value v(t, p) for each pair of area p and time period t. It is also preferable.
[0015] According to another embodiment of the system of the present invention, Regarding the estimation device, The system further includes a population database that stores the density of the population staying in each area p during each time period t, x(t, p), The risk value calculation means, when the number of photographed images aggregated in each area p is equal to or less than a predetermined number, For each area p, calculate a population density ratio r by dividing the visitor population density x(t, p) in time period t by the visitor population density x(t, Ω(p)) in time period t in a predetermined surrounding area Ω(p) centered on the area p; r=x(t,p) / x(t,Ω(p)) For each area p, calculate the spatially corrected risk value v'(t,p) by multiplying the predetermined peripheral risk value v(t,Ω(p)) for time period t in the predetermined peripheral area Ω(p) centered on the area p by the population density ratio r. v'(t,p)=r×v(t,Ω(p)) It is also preferable.
[0016] According to another embodiment of the system of the present invention, Regarding the estimation device, The risk value calculation means, when the number of photographed images aggregated in each area p is equal to or less than a predetermined number, For each area p, calculate the adjusted risk value v''(t,p) for the time period t based on the risk value v(t,p) and the risk value y'(t,p) predicted from the risk value y(t-Δt,p) for the time period t-Δt Δt before using an autoregressive prediction model. y'(t,p)=c+a×y(t-Δt,p)+ε c and a are coefficients, and ε is the error term v''(t,p)=y'(t,p)+K(v(t,p)-y'(t,p)) K(): Kalman gain It is also preferable.
[0017] According to another embodiment of the system of the present invention, Regarding the estimation device, For each area p, a spatiotemporal regression prediction model is used to make predictions from the risk value v(t,p) for time period t and the risk value y'(t,p) predicted from the risk value y(t-Δt,p) for the time period Δt before, t-Δt. The spatiotemporal corrected risk value v'''(t,p) is calculated by multiplying the specified surrounding risk value z(t,Ω(p)) for time period t in the specified surrounding area Ω(p) centered on the area p by the population density ratio r. y'(t,p)=c+a×y(t-Δt,p)+Σ j b j ×y(t-Δt,Ω j (p))+ε c, a, b j is the coefficient, and ε is the error term j is the index that classifies the surrounding area Ω(p) into multiple blocks v'''(t,p)=y'(t,p)+K(v(t,p)-y'(t,p)) K(): Kalman gain It is also preferable.
[0018] According to another embodiment of the system of the present invention, Regarding mobile devices, further comprising an image feature extraction means for extracting image feature values from a plurality of photographed images for each area p; The abnormality level transmitting means further transmits the image feature amount to the estimation device; Regarding the estimation device, further comprising a clustering means for classifying image features of each captured image into a plurality of clusters c for each area p; The risk value calculation means calculates a risk value v in each cluster c for each area p. c Calculate (p) It is also preferable.
[0019] According to another embodiment of the system of the present invention, Regarding the estimation device, The method further comprises a cluster labeling means for assigning, to each cluster c, a label that makes the plurality of image features included in the cluster c relatively similar to each other, using training data in which labels have been assigned in advance to the image features; The risk value calculation means calculates a risk value v for each area p in the label of each cluster c. c Calculate (p) It is also preferable.
[0020] According to the present invention, there is provided a program for causing a computer to function to estimate a risk value of a danger to a moving object, the program comprising: a captured image database that stores images captured by a camera mounted on a moving object in association with the image capturing position; an abnormality degree estimation means for estimating an abnormality degree for each captured image in the captured image database using a learning model previously trained with teacher data including captured images labeled normal and captured images labeled abnormal; an area aggregation means for aggregating the abnormality level of each photographed image in the photographed image database for each area p including the photographing position; a risk value calculation means for calculating a statistical value of the degree of abnormality for each area p as a risk value v(p) for the mobile object; The present invention is characterized by the fact that it makes a computer function by using the above-mentioned method.
[0021] According to the present invention, there is provided an estimation device for estimating a risk value of a danger to a moving object while it is moving, comprising: a captured image database that stores images captured by a camera mounted on a moving object in association with the image capturing position; an abnormality degree estimation means for estimating an abnormality degree for each captured image in the captured image database using a learning model previously trained with teacher data including captured images labeled normal and captured images labeled abnormal; an area aggregation means for aggregating the abnormality level of each photographed image in the photographed image database for each area p including the photographing position; a risk value calculation means for calculating a statistical value of the degree of abnormality for each area p as a risk value v(p) for the mobile object; The present invention is characterized by having the following.
[0022] According to the present invention, there is provided an estimation method for an apparatus for estimating a risk value of a danger to a moving object during movement, the method comprising: The device is a photographed image database that stores photographed images of a camera mounted on a moving body in association with the photographed position; a first step of estimating an anomaly level for each captured image in the captured image database using a learning model previously trained with training data including captured images labeled normal and captured images labeled abnormal; a second step of aggregating the degree of abnormality for each captured image in the captured image database for each area p that includes the capture position; A third step is to calculate the statistical value of the degree of anomaly for each area p as a risk value v(p) for the mobile object. The present invention is characterized by carrying out the following. [Effects of the Invention]
[0023] The system, program, device, and method of the present invention can estimate the risk value of dangers that change from moment to moment, such as traffic congestion, accidents, and incidents, for each area of a region from images captured by a camera mounted on a mobile object. [Brief explanation of the drawings]
[0024] [Figure 1] FIG. 2 is a first functional configuration diagram of a moving body in the system of the present invention. [Figure 2] FIG. 10 is an explanatory diagram of a learning model in an abnormality degree estimation unit of a moving object. [Figure 3] FIG. 2 is a first functional configuration diagram of an estimation device in the system of the present invention. [Figure 4] FIG. 2 is an explanatory diagram of an area aggregation unit in the present invention. [Figure 5] 10A and 10B are explanatory diagrams of spatial direction corrections made by the risk value calculation unit in the present invention. [Figure 6] FIG. 2 is a second functional configuration diagram of a moving body in the system of the present invention. [Figure 7] FIG. 2 is a second functional configuration diagram of the estimation device in the system of the present invention. [Figure 8] FIG. 10 is a third functional configuration diagram of the estimation device according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0025] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0026] According to the system of the present invention, a plurality of moving objects 1 and an estimation device 2 communicate with each other via a network. The moving object 1 is equipped with a camera for capturing images. The moving object 1 may be a vehicle traveling on a road, or may be a portable terminal held by a person. The estimation device 2 estimates a risk value of danger for each area of a region based on images captured by the moving object 1.
[0027] FIG. 1 is a first functional configuration diagram of a mobile object in the system of the present invention.
[0028] The images captured by the camera of the moving object 1 may be moving images captured over a predetermined period of time or may be images (still images). When the moving object is a vehicle, the captured image is taken by a camera mounted on the vehicle, and the captured image is associated with the time of capture. The captured position is also measured by a positioning unit mounted on the vehicle. Alternatively, if the moving object is a mobile terminal, the captured image is taken by a camera mounted on the mobile terminal, and the captured image is associated with the time of capture. The captured position is also measured by a positioning unit mounted on the mobile terminal.
[0029] The shooting position is positioning information obtained by a positioning unit such as a GPS (Global Positioning System) installed in a vehicle or a mobile terminal. Positioning information is generally expressed in latitude and longitude, but may be expressed as map coordinates or converted into map mesh numbers.
[0030] 1, a moving object 1 has a captured image database 10, an abnormality degree estimation unit 11, and an abnormality degree transmission unit 12. These functional components are realized as programs that cause a computer installed in the device to function. The processing flow of these functional components can also be understood as an estimation method. It should be noted that the system may further include a learning parameter transmitter 13 and an integrated parameter receiver 14 for associative learning, which will be described later.
[0031] [Photo database 10] The photographed image database 10 stores images photographed by a camera mounted thereon, in association with at least the photographing position. Each photographed image may also be associated with the photographing time.
[0032] [Abnormality degree estimation unit 11] The abnormality degree estimation unit 11 estimates the abnormality degree for each captured image in the captured image database 10 using a "learning model" that has been trained in advance using training data including captured images labeled normal and captured images labeled abnormal. The learning model of the abnormality degree estimation unit 11 may estimate the abnormality degree as normal (0) or abnormal (1) for each captured image, or may estimate a value within a predetermined numerical range (probability value from 0 to 1).
[0033] [Anomaly Level Transmitting Unit 12] The abnormality level transmission unit 12 associates at least the "degree of abnormality" and the "photographing position" with each captured image estimated by the abnormality level estimation unit 11, and transmits the images to the estimation device 2. It is also preferable to associate the images with the "photographing time." It is not necessary to transmit the captured image itself to the estimation device 2.
[0034] FIG. 2 is an explanatory diagram of a learning model in the abnormality degree estimation unit of the mobile object.
[0035] According to FIG. 2, the learning model in the anomaly degree estimation unit 11 is an example based on domain adaptation using adversarial learning (see, for example, Non-Patent Document 1). The abnormality probability estimation unit 11 of each moving body 1 has a "Source model" and a "Target model" that shares initial parameters with the trained Source model. For a moving object 1 having an "abnormal label," a Source model is trained so as to estimate the "abnormal label" from each captured image in the captured image database 10. For a moving object 1 that does not have an "anomaly label," the Target model is used to input each captured image in the captured image database 10 and output features. The discrimination unit then discriminates between the output of the Source model and the output of the Target model, and by training using adversarial learning, the output distribution of the Source model and the output distribution of the Target model are brought closer together. In other words, the learning parameters of the Target model are updated. The images input to the Source model and the Target model may be generated within the moving object 1 by, for example, an image generation model. Alternatively, the Source model and the Target model may be neural networks based on a convolutional neural network (CNN) or a Transformer. The Target model may be a Source model with an additional layer, and only a portion of the Source model and the additional portion (the portion close to the output layer) may be fine-tuned using its own captured images.
[0036] According to FIG. 2, the system includes a learning parameter transmitting unit 13 and an integrated parameter receiving unit 14. When a learning model based on domain adaptation is applied, the learning parameter transmission unit 13 transmits the learning parameters of the Source model to the estimation device 2. This allows the estimation device 2 to perform federated learning of the learning models of multiple moving objects 1. If there is an "anomaly label" corresponding to an image captured by the moving object 1, but the moving object 1 does not store it and the estimation device 2 has a correspondence between the captured image and the "anomaly label," the estimation device 2 may perform fine tuning using the "anomaly label" instead of the moving object 1. Furthermore, the integrated parameter receiving unit 14 receives the learning parameters integrated by the associative learning from the estimation device 2, and incorporates them into the Source model.
[0037] "Federated learning" is a method of training machine learning models in an environment where data sets are distributed. In conventional machine learning models, distributed data sets must first be aggregated into one large data set, and then a single learning model must be trained. In contrast, in federated learning, distributed data sets remain distributed and are trained individually using a common learning model. Then, only the parameters of each learning model can be aggregated at a central location. By combining multiple parameters received from each learning model, the center can create learning parameters as if the distributed data sets were aggregated into one large data set and trained.
[0038] The source model in the anomaly degree estimation unit 11 may be based on "weakly supervised learning." This involves training using teacher data to which incomplete labels are assigned for moving images or still images. For example, the source model may be labeled not for the entire image but for only a portion of the image. In the case of supervised learning, for example, for photographed images of traffic accidents, it is necessary to provide an "anomaly label" indicating the presence or absence of an abnormality as a correct label as training data. In contrast, in the case of weakly supervised learning using still images, the teacher data only needs to be the presence or absence of a "person" as an anomalous object in the captured image, which is taught as an "anomaly label" indicating the presence or absence of an anomaly in the correct answer label.In the case of weakly supervised learning using moving images, the teacher data only needs to be the presence or absence of a "traffic accident" as an anomalous object in a part (snippet) of a video clip, which is taught as an "anomaly label" indicating the presence or absence of an anomaly in the correct answer label. This significantly reduces the cost of creating training data and enables efficient learning even for large datasets.
[0039] FIG. 3 is a first functional configuration diagram of an estimation device in the system of the present invention.
[0040] 3, the estimation device 2 corresponds to the mobile object 1 in FIG. 1 and includes an area aggregation unit 21, a risk value calculation unit 22, a map display unit 23, and a population database 24. These functional components are realized as programs that cause a computer installed in the device to function. The processing flow of these functional components can also be understood as an estimation method. For associative learning, a learning parameter receiving unit 25 and a learning parameter integrating unit 26 may be further included.
[0041] [Area Aggregation Section 21] The area aggregation unit 21 aggregates the abnormality degree for each captured image received from a moving object for each area p that includes the image capturing position. The area may be, for example, a regional division (divided into, for example, 100 m meshes), or may be a regional unit such as a city, ward, town, village, or block. In the case of a regional division, divisions are made at different granularities (see, for example, Non-Patent Document 2).
[0042] FIG. 4 is an explanatory diagram of the area aggregation unit in the present invention. 4, the area on the map is divided into a mesh pattern. A plurality of anomaly degrees are aggregated for each area p of the target range.
[0043] [Risk Value Calculation Unit 22] The risk value calculation unit 22 calculates the statistical value of the degree of abnormality for each area p as a risk value v(p) for the mobile object. Furthermore, based on the time period element, it calculates a risk value v(t, p) for each pair of area p and time period t.
[0044] The risk value calculation unit 22 calculates the risk value v(p) as a statistical value of the degree of abnormality for each area p, most simply as follows: (1) The ratio of the number of images captured when an abnormality occurs to the number of all images captured in area p (number of images captured when an abnormality occurs / total number of images captured) is calculated. (2) Calculate the mean, median, or mode of the degree of abnormality in area p.
[0045] [Map display section 23] The map display unit 23 divides the map into areas p and displays the risk value v(p) for each area p. Specifically, the map may be displayed by color coding to indicate whether the risk value is high or low for each area p.
[0046] [Population Database 24] The population database 24 stores the visitor population density x(t, p) for each time period t for each area p. A specific example of the database may be population data that has been expanded and estimated with reference to official demographic statistics based on GPS location information obtained from a user's smartphone (see, for example, Non-Patent Document 1). In another embodiment, the density of visiting population x(t, p) may be calculated simply by dividing a large number of moving objects whose photographed images are collected in the photographed image database 10 into groups of areas and time periods.
[0047] <Embodiment for risk value correction> As another embodiment of the risk value calculation unit 22 described above, when "the number of photographed images aggregated in each area p is equal to or less than a predetermined number," it is preferable to correct the risk value as follows. -Spatial direction correction Time direction correction - Correction of space-time direction
[0048] Spatial Direction Correction The risk value calculation unit 22 uses a population database 24 that stores the visitor population density x(t, p) for each time period t for each area p. The risk value calculation unit 22 calculates a population density ratio r for each area p by dividing the visitor population density x(t, p) at time period t by the visitor population density x(t, Ω(p)) at time period t in a specified surrounding area Ω(p) centered on the area p. r=x(t,p) / x(t,Ω(p)) Then, for each area p, a spatially corrected risk value v'(t,p) is calculated by multiplying the predetermined surrounding risk value v(t,Ω(p)) for time period t in a predetermined surrounding area Ω(p) centered on the area p by the population density ratio r. v'(t,p)=r×v(t,Ω(p))
[0049] FIG. 5 is an explanatory diagram of spatial direction correction by the risk value calculation unit in the present invention.
[0050] 5, the area p is divided into a mesh, as in FIG. 4, and the area adjacent to the area p to be judged is set as the predetermined surrounding area Ω(p). The risk value v(p) is then expanded and corrected in the spatial direction. Of course, the predetermined surrounding area is not limited to an area adjacent to it, but may be an area included in a predetermined radius, or an area with mesh-like granularity that is one level coarser.
[0051] [Time direction correction] The risk value calculation unit 22 calculates a corrected risk value v''(t,p) in the time direction for each area p based on the risk value v(t,p) for time period t and the risk value y'(t,p) predicted from the risk value y(t-Δt,p) for time period t-Δt Δt ago using an autoregressive prediction model. y'(t,p)=c+a×y(t-Δt,p)+ε c and a are coefficients, and ε is the error term v''(t,p)=y'(t,p)+K(v(t,p)-y'(t,p)) K(): Kalman gain The Kalman gain is used to minimize the error of the estimated value relative to the true value. Specifically, it is evaluated using the minimum mean square error (MMSE). As a result, the risk value v(p) for one area p to be determined is corrected by extending it in the time direction.
[0052] [Space-time direction correction] The risk value calculation unit 22 calculates a corrected risk value v'''(t,p) in the spatiotemporal direction by multiplying the population density ratio r by a specified surrounding risk value z(t,Ω(p)) for time period t in a specified surrounding area Ω(p) centered on the area p, using a spatiotemporal regression prediction model that makes predictions from the risk value v(t,p) for time period t and the risk value y'(t,p) predicted from the risk value y(t-Δt,p) for time period t - Δt Δt ago. y'(t,p)=c+a×y(t-Δt,p)+Σ j b j ×y(t-Δt,Ω j (p))+ε c, a, b j is the coefficient, and ε is the error term j is the index that classifies the surrounding area Ω(p) into multiple blocks v'''(t,p)=y'(t,p)+K(v(t,p)-y'(t,p)) K(): Kalman gain
[0053] Furthermore, the spatio-temporal regression prediction model may be, for example, a model using a multi-resolution recursive spatio-temporal transformer for long-term prediction (Mu2ReST) (see, for example, Non-Patent Document 4). Conventionally, multi-resolution long-term spatio-temporal prediction has been used for smart city applications such as smart transportation, but it is difficult to apply both multi-resolution knowledge and spatio-temporal dependencies. In contrast, Mu2ReST predicts recursively from coarse to fine resolution. Transformers are good at capturing long-term dependencies, so they can also be applied to spatiotemporal prediction. This can simultaneously learn spatial and temporal dependencies by implementing spatiotemporal attention.
[0054] <Implementation of risk value classification> As another embodiment of the risk value calculation unit 22 described above, it is preferable to not only calculate the risk value for the captured image to be judged in area p, but also classify it into risk types (clustering). To this end, by classifying the captured images in area p according to the image feature level, it is possible to identify which abnormal state the risk value of the captured image to be judged corresponds to.
[0055] FIG. 6 is a second functional configuration diagram of a mobile object in the system of the present invention. The moving object 1 in FIG. 6 further includes an image feature extraction unit 15 compared to FIG.
[0056] [Image feature extraction unit 15] The image feature extraction unit 15 extracts image features from the multiple captured images for each area p. The image feature extraction unit 15 may extract features of the image itself, or may extract features for object detection. For example, the image feature extraction unit 15 may extract features that differ significantly between when a person is captured in the captured image and when a vehicle is captured in the captured image. Then, the abnormality degree transmitting unit 13 transmits the image feature amount for each captured image to the estimation device 2 together with the image feature amount.
[0057] FIG. 7 is a second functional configuration diagram of the estimation device in the system of the present invention. 7 further includes a clustering unit 27 and a cluster label assignment unit 28, as compared with FIG.
[0058] [Clustering Department 27] The clustering unit 27 classifies the image features of each captured image into multiple clusters c for each area p. The cluster algorithm may be, for example, k-means, which is unsupervised learning. As a result, in the case of object detection features, for example, when a captured image includes a person and when a captured image includes a vehicle, the features are classified into different clusters.
[0059] [Cluster Label Assignment Unit 28] The cluster labeling unit 28 uses training data in which labels have been assigned to image features in advance to assign, to each cluster c, a label that makes the image features included in the cluster c relatively similar. As a result, the risk value calculation unit 22 calculates the risk value v c Calculate (p)
[0060] FIG. 8 is a diagram showing a third functional configuration of the estimation device according to the present invention.
[0061] According to FIG. 8, the functional units of the moving object 1 in FIG. 6 and the functional units of the estimation device 2 in FIG. 7 are configured into a single estimation device 2. Each functional unit is the same as those described above. These functional units are realized as a program that causes a computer installed in the device to function. The processing flow of these functional units can also be understood as an estimation method.
[0062] As explained in detail above, the system, program, device, and method of the present invention can estimate the risk value of dangers that change from moment to moment, such as traffic congestion, accidents, and incidents, for each area of a region from images taken by a camera mounted on a mobile object. By calculating precise regional risks and providing this information, local residents will be able to avoid high-risk areas, and vehicle navigation systems will be able to select routes that avoid high-risk areas.
[0063] Furthermore, this will make it possible to, for example, "detect abnormalities in ever-changing dangers such as traffic jams, accidents, and incidents from video footage captured by the camera," which will contribute to Goal 3 of the United Nations-led Sustainable Development Goals (SDGs), which is to "Ensure healthy lives and promote well-being for all at all ages."
[0064] With respect to the various embodiments of the present invention described above, various changes, modifications, and omissions that fall within the scope of the technical spirit and aspects of the present invention may be easily made by those skilled in the art. The above description is merely illustrative and is not intended to be limiting in any way. The present invention is limited only by the claims and their equivalents. [Explanation of symbols]
[0065] 1. Mobile 10 Photo database 11 Abnormality degree estimation part 12. Anomaly level transmitter 13 Learning parameter transmission unit 14 Integrated parameter receiver 15 Image feature extraction unit 2 Estimation device 21 Area Aggregation Section 22 Risk Value Calculation Unit 23 Map display section 24 Population Database 25 Learning parameter receiver 26 Learning parameter integration unit 27 Clustering Department 28 Cluster label assignment
Claims
1. A system having a mobile object and an estimation device that estimates a risk value of a danger for each area p, The moving body is a captured image database that stores images captured by a camera mounted on the vehicle in association with the image capturing position; an abnormality degree estimation means for estimating an abnormality degree for each captured image in the captured image database using a learning model previously trained with teacher data including captured images labeled normal and captured images labeled abnormal; an abnormality level transmitting means for transmitting the abnormality level and the photographing position for each photographed image to the estimation device; and The estimation device an area aggregation means for aggregating the abnormality level of each captured image received from a moving object for each area p including the image capturing position; a risk value calculation means for calculating a statistical value of the degree of abnormality for each area p as a risk value v(p) for the mobile object; A system comprising:
2. For a moving object, the captured image of the camera is an image or a video image for a predetermined period of time.
2. The system of claim 1.
3. The mobile object is a vehicle traveling on a road. The images were taken by a camera mounted on the vehicle. The shooting position was measured by the positioning unit installed in the vehicle. Or, The mobile object is a mobile device held by a person, The photographed image is taken by a camera mounted on a mobile device. The shooting position was measured by the positioning unit installed in the mobile device.
2. The system of claim 1.
4. Regarding the estimation device, A map display means for dividing the map into areas p and displaying the risk value v(p) for each area p is provided. The system of claim 1 further comprising:
5. The abnormality degree estimation means in the mobile object estimates the abnormality degree as normal (0) or abnormal (1), The risk value calculation means in the estimation device calculates a ratio of the number of images captured during abnormality to the number of all images captured in the area p as a statistical value of the degree of abnormality for each area p, Or, The abnormality degree estimation means in the mobile object estimates a value within a predetermined numerical range as the abnormality degree.
2. The system of claim 1.
6. The photographed image database for the moving object further associates each photographed image with a photographed time, The risk value calculation means in the estimation device calculates a risk value v(t, p) for each pair of area p and time period t.
2. The system of claim 1.
7. Regarding the estimation device, The system further includes a population database that stores the density of the population staying in each area p during each time period t, x(t, p), The risk value calculation means, when the number of photographed images aggregated in each area p is equal to or less than a predetermined number, For each area p, calculate a population density ratio r by dividing the visitor population density x(t,p) in time period t by the visitor population density x(t,Ω(p)) in time period t in a predetermined surrounding area Ω(p) centered on the area p, r = x(t,p) / x(t,Ω(p)) For each area p, calculate the spatially corrected risk value v'(t,p) by multiplying the predetermined peripheral risk value v(t,Ω(p)) for time period t in the predetermined peripheral area Ω(p) centered on the area p by the population density ratio r. v'(t,p) = r × v(t,Ω(p)) The system of claim 6 .
8. Regarding the estimation device, The risk value calculation means, when the number of photographed images aggregated in each area p is equal to or less than a predetermined number, For each area p, calculate the adjusted risk value v''(t,p) for the time direction based on the risk value v(t,p) for time period t and the risk value y'(t,p) predicted from the risk value y(t-Δt,p) for the time period t-Δt Δt before using the autoregressive prediction model. y'(t,p)=c+a×y(t-Δt,p)+ε c and a are coefficients, and ε is the error term v''(t,p)=y'(t,p)+K(v(t,p)-y'(t,p)) K(): Kalman gain The system of claim 6 .
9. Regarding the estimation device, For each area p, a spatiotemporal regression prediction model is used to make predictions from the risk value v(t,p) for time period t and the risk value y'(t,p) predicted from the risk value y(t-Δt,p) for the time period Δt before, t-Δt. The spatiotemporal corrected risk value v'''(t,p) is calculated by multiplying the specified surrounding risk value z(t,Ω(p)) for time period t in the specified surrounding area Ω(p) centered on the area p by the population density ratio r. y'(t,p)=c+a×y(t-Δt,p)+Σ j b j ×y(t-Δt,Ω j (p))+ε c, a, b j is the coefficient, and ε is the error term j is the index that classifies the surrounding area Ω(p) into multiple blocks v'''(t,p)=y'(t,p)+K(v(t,p)-y'(t,p)) K(): Kalman gain The system of claim 6 .
10. Regarding mobile devices, further comprising an image feature extraction means for extracting image feature values from a plurality of photographed images for each area p; The abnormality level transmitting means further transmits the image feature amount to the estimation device; Regarding the estimation device, further comprising a clustering means for classifying image features of each captured image into a plurality of clusters c for each area p; The risk value calculation means calculates a risk value v in each cluster c for each area p. c Calculate (p) 2. The system of claim 1.
11. Regarding the estimation device, The method further comprises a cluster labeling means for assigning, to each cluster c, a label that makes the plurality of image features included in the cluster c relatively similar to each other, using training data in which labels have been assigned in advance to the image features; The risk value calculation means calculates a risk value v for each area p in the label of each cluster c. c Calculate (p) The system of claim 10.
12. A program that causes a computer to function to estimate a risk value of a danger to a moving object during movement, a captured image database that stores images captured by a camera mounted on a moving object in association with the image capturing position; an abnormality degree estimation means for estimating an abnormality degree for each captured image in the captured image database using a learning model previously trained with teacher data including captured images labeled normal and captured images labeled abnormal; an area aggregation means for aggregating the abnormality level of each photographed image in the photographed image database for each area p including the photographing position; a risk value calculation means for calculating a statistical value of the degree of abnormality for each area p as a risk value v(p) for the mobile object; A program that causes a computer to function.
13. An estimation device that estimates a risk value of a danger to a moving object while it is moving, a captured image database that stores images captured by a camera mounted on a moving object in association with the image capturing position; an abnormality degree estimation means for estimating an abnormality degree for each captured image in the captured image database using a learning model previously trained with teacher data including captured images labeled normal and captured images labeled abnormal; an area aggregation means for aggregating the abnormality level of each photographed image in the photographed image database for each area p including the photographing position; a risk value calculation means for calculating a statistical value of the degree of abnormality for each area p as a risk value v(p) for the mobile object; An estimation device comprising:
14. An estimation method for a device that estimates a risk value of a danger to a moving object during movement, comprising: The device is a photographed image database that stores photographed images of a camera mounted on a moving body in association with the photographed position; a first step of estimating an anomaly level for each captured image in the captured image database using a learning model previously trained with training data including captured images labeled normal and captured images labeled abnormal; a second step of aggregating the degree of abnormality for each photographed image in the photographed image database for each area p that includes the photographing position; a third step of calculating the statistical value of the degree of abnormality for each area p as a risk value v(p) for the mobile unit; A method for estimating an apparatus, comprising:
Citation Information
Patent Citations
Disaster occurrence probability calculating system and method, operation control apparatus, program, and recording medium
JP2018092467A
Behavior monitoring device, behavior monitoring system, and behavior monitoring program
JP2020095357A