Information provision system, terminal device, and program
The information provision system accurately tracks and transmits positional data for stationary objects captured by moving bodies, addressing the challenge of precise object positioning on Earth.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-08
AI Technical Summary
Existing technologies fail to accurately determine the position of stationary objects on Earth from images captured by moving bodies.
An information provision system comprising a terminal device mounted on a mobile body with image capture, tracking, and location information acquisition capabilities, which transmits positional data to a server when the object moves in or out of the frame, ensuring accurate positioning.
Enables high-accuracy acquisition of positional information for stationary objects by preventing significant deviation from their actual Earth location.
Smart Images

Figure 2026059957000001_ABST
Abstract
Description
Technical Field
[0001] The disclosed technology relates to an information providing system, a terminal device, and a program.
Background Art
[0002] Conventionally, as shown in Patent Document 1 below, for example, the display position of a target stationary object is detected from an image frame captured by a moving body, and the fact that the stationary object moves radially around the vanishing point within the image frame is utilized to predict the display position of the stationary object in the next captured image frame and narrow down the detection range of the stationary object. According to this technology, it is possible to reduce the processing time required for detecting the display position of the stationary object.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, with the technology described in Patent Document 1 above, although it is possible to efficiently detect the display position of a stationary object in an image frame captured by a moving body, it was not possible to accurately obtain the position of the stationary object on the earth.
[0005] The problem of the disclosed technology is to provide a technology for accurately acquiring information on the position of a specific object included in an image captured by a moving body.
Means for Solving the Problems
[0006] One aspect of the disclosed technology is an information provision system comprising a terminal device mounted on a mobile body and a server capable of communicating with the terminal device, wherein the terminal device comprises an image capture unit capable of capturing images, tracking means for tracking changes in the position of a specific object in an image captured by the image capture unit over time when the image includes the specific object, location information acquisition means for acquiring location information of the terminal device, and location information transmission means for transmitting the location information of the terminal device to the server when the specific object tracked by the tracking means moves out of the frame from the image captured by the image capture unit or moves into the frame from the image captured by the image capture unit, and the server is capable of transmitting the location information to other information processing devices different from the server and the terminal device.
[0007] In the above embodiment, it is desirable that the imaging unit is capable of capturing images in the direction forward of the moving body's direction of travel, and that the position information transmission means transmits the position information of the terminal device to the server when the specific object being tracked by the tracking means moves out of frame from the image captured by the imaging unit.
[0008] According to this type of information provision system, the terminal device tracks the temporal changes in the position of a specific object in an image (captured image) taken in front of the direction of movement of a moving object, and transmits the terminal device's position information when the specific object moves out of frame from the captured image to the server as the position information of the specific object. Therefore, it is prevented that the position of the specific object transmitted to the server will not deviate significantly from the actual location of the specific object on Earth. Thus, it is possible to obtain the position information of the specific object with high accuracy.
[0009] In the above embodiment, the imaging unit may be capable of capturing images toward the rear in the direction of travel of the moving body, and the position information transmission means may transmit the position information of the terminal device to the server when the specific object being tracked by the tracking means enters the frame of the image captured by the imaging unit.
[0010] According to this type of information provision system, the terminal device tracks the temporal changes in the position of a specific object in an image (captured image) taken toward the rear in the direction of movement of a moving object, and transmits the terminal device's position information when the specific object enters the frame of the captured image to the server as the position information of the specific object. Therefore, it is prevented that the position of the specific object transmitted to the server will not deviate significantly from the actual location of the specific object on Earth. Thus, it is possible to obtain the position information of the specific object with high accuracy.
[0011] The terminal devices included in the above-mentioned information provision system, as well as the programs executed by those terminal devices, are novel and useful. [Effects of the Invention]
[0012] The disclosed technology makes it possible to provide a technique for accurately acquiring positional information of a specific object included in an image captured from a moving object. [Brief explanation of the drawing]
[0013] [Figure 1] This diagram shows the overall configuration of the information provision system according to the first embodiment. [Figure 2] This diagram illustrates the target of image recognition performed by an in-vehicle terminal device. [Figure 3] This figure shows the configuration of the in-vehicle terminal device and the user terminal device in the information provision system according to the first embodiment. [Figure 4] This figure shows the configuration of the intermediate server and the GIS server in the information provision system according to the first embodiment. [Figure 5]It is a sequence diagram showing the procedure of processing during the operation of an image recognition application in an in-vehicle terminal device. [Figure 6] It is a sequence diagram showing the procedure of processing when a browsing request is made from a user terminal device. [Figure 7] It is an explanatory diagram showing an example of a captured image captured by an in-vehicle terminal device. [Figure 8] It is an explanatory diagram showing an example of a blurred image obtained by applying a blurring process to the captured image shown in FIG. 7. [Figure 9] It is an explanatory diagram showing an example of a map screen of a GIS application in a user terminal device. [Figure 10] It is an enlarged view of a small window screen in the map screen of the GIS application shown in FIG. 9. [Figure 11] It is an explanatory diagram showing the positional relationship between the object to be recognized by the image recognition application of the in-vehicle terminal device and the vehicle (mobile body) equipped with the in-vehicle terminal device. [Figure 12] It is an explanatory diagram showing the change in the position of the recognition object over time in an image captured in the forward direction of travel of the vehicle (mobile body). [Figure 13] It is a diagram for explaining a vector showing the change in the position of the recognition object over time in an image captured by an in-vehicle terminal device mounted in the forward direction of travel of the vehicle (mobile body). [Figure 14] It is a flowchart showing the procedure of image analysis processing (front image analysis processing) by an in-vehicle terminal device when the in-vehicle terminal device is capturing an image in the forward direction of travel of the vehicle (mobile body). [Figure 15] (A) is an explanatory diagram of object information, and (B) is an explanatory diagram of a method for specifying the existence position of an object when it is framed out. [Figure 16] It is an explanatory diagram of the second embodiment, showing the positional relationship between the object to be recognized by the image recognition application of the in-vehicle terminal device and the vehicle (mobile body) equipped with the in-vehicle terminal device. [Figure 17] It is an explanatory diagram of the second embodiment, showing the change in the position of the recognition object over time in an image captured in the backward direction of travel of the vehicle (mobile body). [Figure 18]This is a diagram for explaining a vector indicating the change in the position of a recognition target object over time in an image captured by an in-vehicle terminal device mounted rearward in the traveling direction of a vehicle (mobile object). [Figure 19] This is a flowchart showing the procedure of image analysis processing (rear image analysis processing) by an in-vehicle terminal device when the in-vehicle terminal device captures an image rearward in the traveling direction of a vehicle (mobile object). [Figure 20] This is an explanatory diagram of a method for specifying the existence position of an object when framing in.
Embodiments for Carrying Out the Invention
[0014] 1. First Embodiment Hereinafter, embodiments embodying the disclosed technology will be described in detail with reference to the accompanying drawings. FIG. 1 shows an information providing system 1 according to the first embodiment. The information providing system 1 includes an in-vehicle terminal device A mounted on a mobile object M which is a vehicle, an intermediate server X, a GIS server Y, and a user terminal device B, and each of these devices is connected to a network 2 such as the Internet. The in-vehicle terminal device A and the intermediate server X are communicably connected to each other via the network 2. Also, the intermediate server X and the GIS server Y are communicably connected to each other via the network 2. Also, the GIS server Y and the user terminal device B are communicably connected to each other via the network 2. Note that the connection mode of the in-vehicle terminal device A is a wireless connection. The connection modes of the intermediate server X, the GIS server Y, and the user terminal device B to the network 2 may be either wired or wireless.
[0015] A plurality of in-vehicle terminal devices A for acquiring information about a specific object from the destination of the movement of the mobile object M can be registered in the information providing system 1, and a plurality of users (users) of the information collected by the in-vehicle terminal device A can also be registered The information providing system 1 is a system for collecting various information about a specific object by the in-vehicle terminal device A at the destination of the movement of the mobile object M and providing this information to necessary users.
[0016] In-vehicle terminal device A is a smartphone equipped with image recognition AI capable of detecting specific objects from captured images. In-vehicle terminal device A may be a terminal device other than a smartphone, such as a dashcam equipped with image recognition AI. When in-vehicle terminal device A detects a specific object, it transmits information about that object to the intermediate server X.
[0017] Intermediate server X stores and stores information about specific objects received from in-vehicle terminal device A, and transmits it to GIS server Y.
[0018] GIS server Y transmits information about a specific object to user terminal device B in response to a request from user terminal device B.
[0019] User terminal device B is a personal computer (hereinafter referred to as "PC") on which an application capable of overlaying information about a specific object onto a map screen is installed. User terminal device B may be a terminal device other than a PC, such as a smartphone or tablet. User terminal device B sends a viewing request to GIS server Y and receives information (including information about a specific object) from GIS server Y according to the content of the request.
[0020] Specific objects targeted for image recognition by the in-vehicle terminal device A include power poles, power transmission equipment (power lines, transmission towers), and electrical manholes, as well as power transmission and distribution equipment experiencing abnormalities (power transmission and distribution equipment in an emergency state). Power transmission and distribution equipment is required to be monitored regularly to ensure a stable power supply, and if an abnormality is detected, continuous monitoring and elimination of the abnormality are required. For this reason, power supply operators such as power companies carry out regular patrols and monitoring by their employees.
[0021] Here, as shown in Figure 2, examples of abnormalities in utility poles include bird nests being built on them or vines becoming entangled around them. Examples of abnormalities in power transmission equipment include situations where a construction site is located beneath the power lines, posing a risk of cranes or other vehicles getting caught on them. Examples of abnormalities in electrical manholes include situations where there are cracks in the road surrounding the manhole. Utility poles with bird nests, utility poles entangled with vines, power lines near construction site installations, and electrical manholes near cracked roads are all examples of power transmission and distribution equipment in an emergency state. The specific objects targeted for image recognition by the in-vehicle terminal device A may be various types of power transmission and distribution equipment (utility poles, power lines, transmission towers, electrical manholes, etc.), regardless of whether they are abnormal or not.
[0022] In addition to the routine tasks of inspecting and monitoring electrical equipment (transmission and distribution equipment) as described above, power supply companies also have tasks such as employee visits to customer sites. If information on the current state of electrical equipment can be obtained incidentally while performing these other tasks, it would be possible to eliminate the need for separate inspections solely for collecting information on the current state of electrical equipment, thereby improving the efficiency of electrical equipment inspection and monitoring operations.
[0023] Therefore, in this form of information provision system 1, an in-vehicle terminal device A with the image recognition application 16 (described later) installed is mounted on a mobile vehicle. When the running image recognition application 16 detects a specific object (for example, a utility pole with a bird's nest on it (hereinafter also referred to as a "utility pole with a nest")), the in-vehicle terminal device A takes a photograph and transmits the information related to the captured image to the intermediate server X. The intermediate server X then works in conjunction with the GIS server Y to make the information collected by the in-vehicle terminal device A available for use by a user terminal device B with the GIS application 36 (described later) installed. In other words, information provision system 1 can efficiently share information collected from the city.
[0024] Furthermore, the specific objects targeted for image recognition by the in-vehicle terminal device A are not limited to electrical equipment (power transmission and distribution equipment), but may also include other facilities such as vacant houses. If vacant houses are targeted for image recognition and information collection, it can contribute to improving the efficiency of the vacant house patrol and monitoring work that local governments regularly conduct. In this configuration, the users of the information would be local governments. Also, the mobile entity is not limited to company cars of power supply companies, but may include other vehicles such as buses, taxis, and garbage trucks. In addition, the mobile entity may be a non-vehicle mobile entity such as a drone.
[0025] Next, the configurations of the in-vehicle terminal device A, user terminal device B, intermediate server X, and GIS server Y will be explained based on Figures 3 and 4. Figure 3(A) shows the configuration of the in-vehicle terminal device A. The in-vehicle terminal device A is a smartphone with camera and GPS functions. In other words, as shown in Figure 3(A), the in-vehicle terminal device A includes a controller 10 containing a CPU 12 and memory 14, a user IF (user interface) 20, a communication IF (communication interface) 21, an imaging unit 22 (an example of an imaging unit), and a GPS receiver 23. The user IF 20, communication IF 21, imaging unit 22, and GPS receiver 23 are electrically connected to the controller 10. Note that the controller 10 is a general term for the hardware and software used to control the in-vehicle terminal device A, and does not necessarily represent a single piece of hardware actually present in the in-vehicle terminal device A.
[0026] The CPU 12 executes various processes according to the program read from memory 14 and based on the operator's input. Memory 14 stores various programs, including various application programs (hereinafter referred to as "apps"), and various data. Memory 14 is also used as a workspace when various processes are executed. The buffer provided by the CPU 12 is also an example of memory. Note that memory 14 is not limited to ROM, RAM, HDD, etc., built into the in-vehicle terminal device A, but may also be a storage medium that the CPU 12 can read and write to.
[0027] The user interface 20 includes hardware that displays a screen to inform the user of information and hardware (input device) that accepts user operations. Specifically, the user interface 20 of in-vehicle terminal device A includes a touch panel equipped with information display and input acceptance functions. The user interface 20 of in-vehicle terminal device A also includes a speaker, which is hardware capable of outputting various sounds. The various sounds output from the speaker include warning sounds and voices (announcement voices, text-to-speech voices, etc.).
[0028] Communication IF21 includes hardware for communicating with external devices such as intermediate server X via network 2. The communication standards for communication IF21 include Wi-Fi®, 4G, 5G, etc. The in-vehicle terminal device A may have multiple communication IF21s that support multiple communication standards.
[0029] The imaging unit 22 includes a camera with a lens and other components, and is capable of capturing still images and videos. In other words, the in-vehicle terminal device A has the function of capturing images in the direction that the lens of the imaging unit 22 is pointed. In this embodiment, it is assumed that the in-vehicle terminal device A is mounted on the back of the vehicle's rearview mirror or the like so that the lens of the imaging unit 22 faces forward in the direction of travel of the vehicle (more specifically, so that the lens faces approximately 0 degrees when the direction of travel is considered 0 degrees). Note that the direction of the lens of the imaging unit 22 does not matter as long as it is within 30 degrees to the left and right in the direction of travel. If the direction of the lens is within this range, it is considered to be within the range of facing forward in the direction of travel of the vehicle. Furthermore, it is most desirable for the angle of view of the lens to be 180 degrees. This is because, as will be described later, it is possible to improve the accuracy of position identification when acquiring position information of a specific object by tracking that specific object in the captured image. While a lens angle of view between 100 and 180 degrees is acceptable, it is preferable that it be between 130 and 180 degrees, and even more preferable that it be between 150 and 180 degrees.
[0030] The GPS receiver 23 receives GPS signals. The in-vehicle terminal device A has the function of acquiring its location information (latitude and longitude information) on Earth based on the analysis results of the GPS signals received by the GPS receiver.
[0031] The memory 14 of the in-vehicle terminal device A stores an operating system (hereinafter referred to as "OS") 15 and an image recognition application 16. The memory 14 is also provided with an object information storage unit 18 for storing information about specific objects, which will be described later. The OS 15 is a multitasking OS that can process multiple tasks in parallel by managing and switching between multiple tasks, such as iOS® or Android®.
[0032] The image recognition application 16 installed on the in-vehicle terminal device A is an application program that includes an image recognition AI program (image recognition AI program 17) capable of recognizing specific objects. This image recognition AI program 17 employs a pre-trained model that has been trained using supervised deep learning to recognize specific objects.
[0033] Specific targets include, as mentioned above, utility poles with bird nests or vines wrapped around them, power transmission equipment near construction site installations (such as construction guards), and electrical manholes near cracks in roads. The image recognition application 16 in this form includes a pre-trained image recognition AI program 17 capable of recognizing all of the following: utility poles, power transmission equipment, electrical manholes, bird nests, vines, construction site installations (such as construction guards), and cracked roads. Furthermore, users of the image recognition application 16 can selectively input the target objects for image recognition (recognition targets) according to the purpose of information collection by operating the user IF 20. Alternatively, the target objects for image recognition by the image recognition application 16 may be fixed to predetermined ones and cannot be selected. In addition, the types of targets to be recognized by the image recognition application 16 can be changed as appropriate.
[0034] When the in-vehicle terminal device A recognizes a selectively input object to be recognized (in this embodiment, a "utility pole with a nest") using the image recognition application 16, it captures an image of the object using the imaging unit 22 and transmits various information related to that image (captured image) to the intermediate server X via the communication IF 21. The in-vehicle terminal device A also acquires its location information (latitude and longitude) based on the signal from the GPS receiver 23 when it captures the image of the object to be recognized (in this embodiment, a utility pole with a nest), and transmits this information to the intermediate server X via the communication IF 21. Information related to the captured image (specific captured image) including the object to be recognized (in this embodiment, a utility pole with a nest), and location information related to that specific captured image (location information on Earth based on the GPS receiver 23) are stored in the object information storage unit 18 of the memory 14. This information may be stored separately for each object to be recognized (captured object) or not.
[0035] Figure 3(B) shows the configuration of user terminal device B. User terminal device B is a desktop or laptop PC. As shown in Figure 3(B), user terminal device B comprises a controller 30 including a CPU 32 and memory 34, a user interface (IF) 37, and a communication interface (IF) 38. The user interface 37 and communication interface 38 are electrically connected to the controller 30. Note that the controller 30 is a general term for the hardware and software used to control user terminal device B, and does not necessarily represent a single piece of hardware actually present in user terminal device B.
[0036] The CPU 32 executes various processes according to the program read from memory 34 and based on the operator's input. Memory 34 stores various programs, including various applications, and various data. Memory 34 is also used as a workspace when various processes are executed. The buffer provided by the CPU 32 is also an example of memory. Note that memory 34 is not limited to ROM, RAM, HDD, etc., built into the user terminal device B, but may also be a storage medium that the CPU 32 can read and write to.
[0037] User IF37 includes input and output devices. Specifically, User IF37 of user terminal device B includes input devices such as a keyboard and mouse, and a liquid crystal display as a display device capable of displaying images. User IF37 may also be a touch panel equipped with display and input reception functions.
[0038] The communication interface 38 includes hardware for communicating with external devices such as the GIS server Y via network 2. The communication standards for the communication interface 38 include Ethernet® and Wi-Fi®. The user terminal device B may have multiple communication interface 38s supporting multiple communication standards.
[0039] The memory 34 of user terminal device B stores the OS (operating system) 35 and the GIS application 36. The OS 35 is a multitasking OS that can process multiple tasks in parallel by managing and switching between multiple tasks, such as Windows®, macOS®, Linux®, iOS®, and Android®.
[0040] The GIS application 36 installed on user terminal device B is a map application program that utilizes GIS (Geographic Information System). The GIS application 36 works in conjunction with the image recognition application 16 installed on in-vehicle terminal device A to allow users to view information about specific objects (in this embodiment, "utility poles with nests") along with their location information on the map screen. In other words, the GIS application 36 is an application that can overlay information about various facilities, such as specific objects, onto the map screen. The screen of the GIS application 36 will be described in detail later. Alternatively, the GIS application 36 may be a web application accessed via a browser through a web page provided by the GIS server Y.
[0041] Figure 4(A) shows the configuration of the intermediate server X. As shown in Figure 4(A), the intermediate server X includes a CPU 40, memory 42, and a communication interface 48. Note that the intermediate server X may be constructed using multiple servers.
[0042] The CPU 40 executes various processes according to the program read from memory 42. Memory 42 stores various programs and various data. Memory 42 is also used as a work area when various processes are executed. The buffer provided by the CPU 40 is also an example of memory. Note that memory 42 is not limited to ROM, RAM, HDD, etc. built into the intermediate server X, but may also be a storage medium that the CPU 40 can read and write to.
[0043] The communication IF48 includes hardware for communicating with external devices such as the in-vehicle terminal device A and the GIS server Y via network 2. The communication standard for the communication IF48 is Ethernet®, Wi-Fi®, etc. The intermediate server X may have multiple communication IF48s that support multiple communication standards.
[0044] The memory 42 of the intermediate server X stores the in-vehicle terminal DB 43, which is a database (DB) that stores data for identifying the in-vehicle terminal device A. The in-vehicle terminal DB 43 stores information such as the terminal ID that is assigned to the in-vehicle terminal device A after the image recognition application 16 has been installed and registration for use has been completed.
[0045] Furthermore, the memory 42 of the intermediate server X is equipped with an object DB 44, which is a DB (database) for storing object information (information about a specific object). Object information is stored in the object DB 44 separately for each in-vehicle terminal device A to which a terminal ID has been assigned. In other words, information about a specific object received from one in-vehicle terminal device A, namely in-vehicle terminal device A1, is stored in the storage area for in-vehicle terminal device A1, and information about a specific object received from another in-vehicle terminal device A, namely in-vehicle terminal device A2, is stored in the storage area for in-vehicle terminal device A2. Note that the object DB 44 may store object information without separating it by in-vehicle terminal device A. In addition, the in-vehicle terminal DB 43 and object DB 44 may be provided on an external device (e.g., a database server) that can communicate with the intermediate server X.
[0046] The information regarding a specific object includes text data indicating what the object is (name of the object), the object's position in the photograph (captured image) containing the object, and the object's position on Earth (position based on the GPS receiver 23), as well as image data of the captured image. The information regarding a specific object will be described in more detail later.
[0047] Furthermore, when the intermediate server X satisfies predetermined transmission conditions, it sends text data from the information about a specific object stored in the object DB44 to the GIS server Y. In this configuration, the transmission condition is the arrival of a periodic transmission timing. The transmission condition can be changed as needed; for example, it could be set to when a request is received from the GIS server Y.
[0048] Figure 4(B) shows the configuration of GIS server Y. As shown in Figure 4(B), GIS server Y includes a CPU 50, memory 52, and a communication interface 58. GIS server Y may be constructed using multiple servers.
[0049] The CPU 50 executes various processes according to the program read from memory 52. Memory 52 stores various programs and various data. Memory 52 is also used as a workspace when various processes are executed. The buffer provided by the CPU 50 is also an example of memory. Note that memory 52 is not limited to ROM, RAM, HDD, etc., built into the GIS server Y, but may also be a storage medium that the CPU 50 can read and write to.
[0050] Communication IF58 includes hardware for communicating with external devices such as user terminal device B and intermediate server X via network 2. The communication standards for communication IF58 include Ethernet® and Wi-Fi®. GIS server Y may have multiple communication IF58s supporting multiple communication standards.
[0051] The memory 52 of the GIS server Y stores a GIS database 53, which is a database (DB) that stores GIS data. The GIS database 53 stores data representing map elements (for example, cities, rivers, mountains, roads, buildings, etc.), data of map images containing those elements, and data of equipment registered by the user (utility poles, power transmission equipment, electrical manholes, etc.). The GIS server Y uses the data stored in the GIS database 53 to implement functions such as displaying a map screen by the GIS application 36 on the user terminal device B.
[0052] Furthermore, the memory 52 of the GIS server Y stores a user account database 54, which is a DB (database) that stores user account data. The user account database 54 stores information such as user account information assigned to users who have installed the GIS application 36 and completed user registration (for example, a company that has a usage agreement for the GIS application 36). In other words, the memory 52 of the GIS server Y stores information about users who have the authority to use the information provision system 1. Users of the GIS application 36 can log in to the GIS application 36 on their user terminal device B and use the functions of the GIS application 36.
[0053] Furthermore, the memory 52 of the GIS server Y is equipped with an object DB 55, which is a DB (database) that stores object information (information about a specific object). The object DB 55 stores information about a specific object received from the intermediate server X. Note that the information stored in the object DB 55 of the GIS server Y and the information stored in the object DB 44 of the intermediate server X do not have to be exactly the same.
[0054] Furthermore, when predetermined transmission conditions are met (in this configuration, in response to a request from user terminal device B), the GIS server Y transmits information about a specific object stored in the object DB 55 to user terminal device B. This allows the user to check the type of a specific object (specifically, "utility pole with a nest") and its location (location on Earth based on the GPS receiver 23) on the map screen of the GIS application 36 launched on user terminal device B. In addition, the user can view a modified image (see blurry image 120, Figure 8, described later) taken by the in-vehicle terminal device A on the map screen of the GIS application 36, which has been partially modified, allowing them to check the condition of a specific object (such as where and what kind of bird nest is built on the utility pole).
[0055] Next, the procedure for collecting information about a specific object in the information provision system 1 of this embodiment will be explained with reference to the sequence diagram in Figure 5. As shown in Figure 5, when the image recognition application 16 is launched on the in-vehicle terminal device A, the in-vehicle terminal device A sends a connection request to the intermediate server X (A), and the intermediate server X responds to it. This establishes a connection between the in-vehicle terminal device A and the intermediate server X. Once the connection with the intermediate server X is established, the in-vehicle terminal device A starts taking images with the imaging unit 22 (B). Note that the imaging unit 22 takes images continuously while the image recognition application 16 is running.
[0056] Next, the in-vehicle terminal device A analyzes whether a specific object is included in the image (captured image) being captured by the imaging unit 22 (C). As described above, in this embodiment, the specific object is a utility pole with a nest inside. The specific object can be changed as appropriate depending on the purpose.
[0057] If the in-vehicle terminal device A determines that a specific object (a utility pole with a nest) is included in the captured image (opt: [object detected]), the imaging unit 22 takes a photograph (still image) (D). The image file of the captured photograph is stored in the object information storage unit 18. In other words, the in-vehicle terminal device A stores the image data of the image containing the object (a utility pole with a nest) in the object information storage unit 18. The image file format is, for example, jpeg. Alternatively, the in-vehicle terminal device A may be configured to capture and store a video when a specific object is included in the captured image.
[0058] Next, the in-vehicle terminal device A transmits the inference results (AI inference results) of the image recognition AI program 17 and location information to the intermediate server X as text data (E). The AI inference results include the names of the objects included in the captured image (specifically, "utility pole with a nest") and the positions of the objects in the captured image (specifically, "left and top in the captured image"). The location information includes the location information of the in-vehicle terminal device A (latitude and longitude information) and the location information of the specific object (information on whether it is to the left, right, top, or bottom in relation to the direction of travel of the moving object).
[0059] Furthermore, the location information of the in-vehicle terminal device A transmitted to the intermediate server X is location information on Earth based on the GPS receiver 23. In this configuration, in order to obtain the location information of the discovered object as accurately as possible, the in-vehicle terminal device A acquires the location information when the object goes out of frame from the captured image (in other words, when it disappears from the captured image). This point will be explained later.
[0060] When the intermediate server X receives AI inference results and location information text data from the in-vehicle terminal device A, it stores this data in the object DB 44 of the memory 42 (F). In this configuration, the storage area of the object DB 44 of the memory 42 is divided according to the terminal ID of the in-vehicle terminal device A, so the data received by the intermediate server X from the in-vehicle terminal device A is stored in the storage area corresponding to the terminal ID of the source in-vehicle terminal device A.
[0061] Furthermore, following the transmission process of AI inference results and location information as text data (E), the in-vehicle terminal device A transmits information about the captured photograph (hereinafter referred to as "image information") as text data to the intermediate server X (G). The image information referred to here is information indicating the storage location (path information) and file name of the image data of the captured image to be transmitted later.
[0062] When the intermediate server X receives text data of image information from the in-vehicle terminal device A, it stores this data in a memory area in the object DB44 of memory 42 corresponding to the terminal ID of the source in-vehicle terminal device A (H).
[0063] When the intermediate server X determines that a predetermined transmission timing has arrived (opt:[transmission timing arrived]), it transmits the AI inference results, location information, and image information text data received from the in-vehicle terminal device A to the GIS server Y (I). In this configuration, the transmission timing is the time elapsed since the previous transmission. In other words, data transmission from the intermediate server X to the GIS server Y is performed periodically. In this transmission process (I), if data related to multiple photographs has not yet been transmitted to the GIS server Y, all untransmitted data is transmitted to the GIS server Y. Note that AI inference results and image information are examples of specific information.
[0064] When GIS server Y receives AI inference results, location information, and image information (information about specific objects excluding image data) from intermediate server X, it distinguishes this information for each photograph and stores it in object DB55 (J).
[0065] Furthermore, when the in-vehicle terminal device A determines that a predetermined transmission timing has arrived (opt:[transmission timing has arrived]), it transmits the image data of the captured photograph to the intermediate server X (K). In this configuration, the transmission timing is when tracking of the object (tracking until the object leaves the frame of the captured image), as described later, is not being performed. In other words, in this configuration, the conditions for transmitting image data are different from the conditions for transmitting text data. Since transmitting image data has a greater transmission load than transmitting text data, the image data is transmitted at a time when the processing load on the in-vehicle terminal device A is low. Note that the image data transmission timing can be changed as appropriate, such as by setting it to a fixed period. Alternatively, the image data transmission timing may be set to when the user IF20 of the in-vehicle terminal device A receives an operation to instruct the transmission of image data. In other words, the in-vehicle terminal device A may be configured to transmit image data to the intermediate server X based on human operation.
[0066] When the intermediate server X receives image data from the in-vehicle terminal device A, it stores this data in the object DB44 of memory 42 (L). As described above, the in-vehicle terminal device A has already sent information about the storage location and file name of the image data as image information to the intermediate server X. In other words, the storage location and file name of the image data have been specified. Therefore, the intermediate server X saves the image data in the specified storage location (storage area of object DB44) with the specified file name.
[0067] In the information provision system 1, while the image recognition application 16 is running on the in-vehicle terminal device A, the series of processes described above, from the analysis of captured images (C) to the transmission of image data (K), are repeatedly performed.
[0068] In this configuration, when the image recognition application 16 detects an object (in this configuration, a "utility pole with a nest"), the in-vehicle terminal device A transmits the AI inference results (such as the name of the object included in the captured image and the position of the object in the captured image), location information, image information (such as the storage location and file name of the image data of the captured image), as well as the image data of the captured photograph, to the intermediate server X. Therefore, information about a specific object is efficiently accumulated in the intermediate server X in accordance with the movement of the vehicle equipped with the in-vehicle terminal device A.
[0069] Furthermore, the intermediate server X transmits only the text data from the data received from the in-vehicle terminal device A to the GIS server Y. Therefore, information about specific objects, excluding image data, can be efficiently collected on the GIS server Y. In addition, compared to a configuration in which the intermediate server X transmits the image data of the captured images to the GIS server, it is possible to prevent the image data from leaking to servers other than the GIS server Y, and to prevent personal information that may be contained in the images based on the image data from being leaked to third parties.
[0070] Furthermore, in this configuration, when the image recognition application 16 detects an object, the in-vehicle terminal device A first sends only text data to the intermediate server X. Therefore, compared to a configuration that immediately sends image data of a photograph containing the object upon detection, it is possible to reduce the processing load when an object is detected.
[0071] Next, the procedure for a user to view information about a specific object in this information provision system 1 will be explained with reference to the sequence diagram in Figure 6. When the GIS application 36 is launched on the user terminal device B and the user inputs an information viewing instruction using the user IF 37, the user terminal device B makes an information viewing request to the GIS server Y, as shown in Figure 6 (M).
[0072] If the GIS server Y receives a request from user terminal device B to view information other than an image (alt:[Request for non-image viewing]), it sends text data related to the object stored in the object DB 55 in memory 52 to user terminal device B (N). This allows the user to view information related to the object (such as the object name and location) on the map screen of the GIS application 36. Here, since the information that can be viewed is limited to text data, the possibility of personal information leakage is lower compared to a configuration that allows viewing of the image itself. The map screen of the GIS application 36 will be described later.
[0073] Furthermore, if the GIS server Y receives a request from user terminal device B to view information, and that request is for viewing an image (alt:[Image viewing request]), it will send an image viewing request to intermediate server X (O).
[0074] When intermediate server X receives an image viewing request from GIS server Y, it performs blurring processing (P). The blurring processing includes image processing that blurs objects other than the specific object (non-specific objects) included in the image being viewed. This is because the photographs taken by the in-vehicle terminal device A may contain various objects other than the target object. Note that the blurring processing may be performed by another device that can communicate with intermediate server X. Specifically, for example, a server for blurring processing may be set up separately from intermediate server X, and this server may receive image data of the captured image from intermediate server X, perform blurring processing, and pass the blurred image data to GIS server Y, either via intermediate server X or without it.
[0075] Figure 7 shows an example of an image taken by the in-vehicle terminal device A. The image 100 shown in Figure 7 includes a specific object, "a utility pole 101 on which a bird's nest 102 has been built," as well as people (pedestrians, etc.) 110 and other vehicles (oncoming vehicles, etc.) 112. When a photograph containing a specific object (a utility pole 101 on which a bird's nest 102 has been built) also includes people 110 and other vehicles 112, it is undesirable from the standpoint of protecting personal information if such a photograph becomes accessible to a wide range of people. Therefore, in this embodiment, blurring processing is used to obscure the people 110 and other vehicles 112 that appear in the photograph taken by the in-vehicle terminal device A.
[0076] Figure 8 shows a blurred image 120 obtained by applying blurring processing to the captured image 100 shown in Figure 7. In the blurred image 120 shown in Figure 8, the part of the utility pole 101 on which the bird's nest 102 was built remains clear, but the parts of the person 110 and other vehicles 112 are blurred to the extent that they cannot be identified. In other words, the image of person 110 is processed into a blurred person image 110A, and the image of other vehicles 112 is processed into a blurred vehicle image 112A. It should be noted that a trained image recognition AI is used to recognize people and vehicles in the photograph during the blurring process. Furthermore, the blurring process may be configured to blur everything in the photograph taken by the in-vehicle terminal device A except for the specific object. In addition, other known methods besides blurring can be appropriately adopted as methods for blurring things other than the specific object.
[0077] Following the blurring process, the intermediate server X transmits the image data of the blurred image 120, in which people and other vehicles have been blurred by the blurring process, to the GIS server Y (Q). The GIS server Y then transmits the image data of the blurred image 120 received from the intermediate server X to the user terminal device B (R). This allows the user to view the blurred image 120 on the screen of the GIS application 36, clearly showing the specific object, "the utility pole 101 on which the bird's nest 102 was made." In other words, by viewing the blurred image 120, the user can clearly confirm the current state of the object (the utility pole 101 on which the bird's nest 102 was made) (size of the bird's nest 102, location where it was made, etc.). On the other hand, since people and vehicles are blurred in the blurred image 120, the user cannot identify them, thus protecting personal information. In addition, since the blurred image 120 can be an image with less information than the original captured image 100, it is possible to reduce the transmission load of image data. As a result, the information provision system 1 makes it possible to increase the degree of freedom in using images captured by the in-vehicle terminal device A and improve convenience. In particular, in this configuration, since the images sent from the intermediate server X to another information processing device such as the GIS server Y are blurry images 120, it is possible to share the information (images) that have been widely collected and stored on the intermediate server X with many users, which easily leads to benefits for society as a whole.
[0078] Next, we will describe the map screen displayed by the GIS application 36 installed on the user terminal device B. Figure 9 is an example of a map screen displayed by the GIS application 36. The map screen (information viewing screen) 60 in the example shown in Figure 9 includes the username 62 of the user using the GIS application 36 and a map 64 of the area selected by the user. The user can display a map 64 of any area on the map screen 60 by operating the user IF 37. The map 64 displays an arrow 65 indicating the location of an object (in this embodiment, a utility pole with a nest) discovered by the vehicle-mounted terminal device A. The direction of the arrow 65 indicates which side of the road the object was on, as viewed from the vehicle-mounted terminal device A mounted on the mobile body M.
[0079] In the map screen 60 shown in Figure 9, six arrows 65 are displayed on the map 64. Each arrow 65 also functions as a button that accepts instructions to display detailed information about an object. In other words, the user can select any one of the arrows 65 on the map screen 60 by operating the user IF37, and the detailed information of the object corresponding to that arrow 65 will be displayed. Of the six arrows 65 in the map 64, the arrow with a border 66 (hereinafter referred to as "thick arrow 65a") indicates the arrow that has been selected by the user. When the user selects any of the arrows 65, the GIS application 36 displays a small window screen 70 overlaid on top of the map screen 60.
[0080] The small window screen 70 includes an enlarged map 72 that shows the location of the thick arrow 65a and its surrounding area, and a detailed information section 74 that provides detailed information about the object corresponding to the thick arrow 65a. Figure 10 is an enlarged view of the small window screen 70. As shown in Figure 10, the enlarged map 72 in the small window screen 70 displays arrows 67 at each location pointed to by arrow 65 in map 64, and text 73 indicating the name of the object at the location of arrow 67 is displayed near each arrow 67. In this example, the text 73 reads "Nested Utility Pole". By looking at the enlarged map 72, the user can recognize that there are three utility poles along Street A. In the enlarged map 72, the arrow 67 corresponding to the thick arrow 65a in map 64 is displayed as a thick arrow 67a with a border. Also, in the enlarged map 72, the direction of the arrow 67 indicates the direction of movement of the moving object M when a photograph of the object at the location of arrow 67 was taken. The detection of the direction of movement of the moving object will be described later.
[0081] The detailed information section 74 displays detailed information about the object corresponding to the thick arrow 65a (thick arrow 67a). The detailed information section 74 includes the following items: "Key code," "Class name (name of the object)," "Photo file name," "Photo file," "Display position (left / right)," "Display position (up / down)," "Longitude," "Latitude," "Date taken," "Direction (degrees)," and "Speed (m / s)." All items are displayed as text.
[0082] The "Key Code" field displays the code assigned to access information about the object corresponding to the thick arrow 67a. The "Class Name (Name of Item)" field displays the name of the object corresponding to the thick arrow 67a. In this example, the "Class Name (Name of Item)" field displays the text "Nested Utility Pole".
[0083] The "Photo File Name" field displays the file name of the photo taken by the in-vehicle terminal device A that photographs the object corresponding to the thick arrow 67a. The word "Display" is underlined in the "Photo File" field. This "Display" word functions as a button that accepts the instruction to display the photo corresponding to the photo file. In other words, by clicking the "Display" word in the small window screen 70 through the operation of user IF37, the user can display the photo of the object corresponding to the thick arrow 67a (the photo with the file name listed in the "Photo File Name" field, hereinafter also referred to as the "Target Photo") on user IF37's display. In short, user terminal device B makes a request to GIS server Y to view the image based on the click of the "Display" word in the small window screen 70.
[0084] The image displayed in response to the user's actions is a blurred image 120, as shown in Figure 8. As mentioned above, in the blurred image 120, the object, "the utility pole 101 on which the bird's nest 102 was built," is clearly displayed, but passersby (people 110) and oncoming vehicles (other vehicles 112) are blurred to the extent that they cannot be identified. Therefore, the user can accurately grasp the situation of the object, but cannot grasp the situation of non-objects that are not the target of surveillance (the target of information collection). In other words, personal information cannot be obtained from the blurred image 120.
[0085] In addition, the "Display Position (Left / Right)" item in the detailed information section 74 of the small window screen 70 displays information on whether the object (utility pole with nest) is positioned to the left or right of the center in the photograph. In this example, "Left" is displayed. The "Display Position (Up / Down)" item displays information on whether the object (utility pole with nest) is positioned above or below the center in the photograph. In this example, "Up" is displayed.
[0086] The "Longitude" and "Latitude" fields display the longitude and latitude of the location information transmitted by the in-vehicle terminal device A to the intermediate server X when it detects an object. In other words, it displays the longitude and latitude indicating the Earth's location of the object detected by the in-vehicle terminal device A.
[0087] The "Date Taken" field displays the date the photograph was taken, in Western calendar year. The "Direction (degrees)" field displays the direction of movement of the object M at the time the photograph was taken. The direction of movement is displayed in degrees, with true north being 0 degrees, true east 90 degrees, true south 180 degrees, and true west 270 degrees. The "Speed (m / s)" field displays the speed of movement of the object M at the time the photograph was taken, in meters per second. In this example, the "Direction (degrees)" field displays "84.70" and the "Speed (m / s)" field displays "9.90," indicating that the photograph was taken when the object M was moving eastward at a speed of 9.9 m / s. Additionally, the "Display Position (left / right)" and "Display Position (up / down)" fields indicate that the object is located in the upper left corner of the photograph. Based on this information, the GIS application 36 identifies the display location (on either the left or right side of the road) of the object (utility pole with a nest) corresponding to the thick arrow 65a in the enlarged map 72 on the small window screen 70.
[0088] The in-vehicle terminal device A mounted on the mobile body M is equipped with a known acceleration sensor, an angular velocity sensor (gyro sensor), and a compass sensor (digital compass), and these sensors are capable of detecting the direction and speed of movement of the mobile body M at the time of photography. Furthermore, these sensors enable the detection of the position of the in-vehicle terminal device A even in locations where GPS signal reception is difficult, such as inside tunnels. In addition, the various text data (information about a specific object excluding image data) concerning the object sent from the in-vehicle terminal device A to the intermediate server X and from the intermediate server X to the GIS server Y shall include, in addition to AI inference results, location information, and image information, text data on the date the target photograph was taken, as well as the direction and speed of movement of the mobile body M at the time the target photograph was taken.
[0089] Next, we will explain in detail the timing at which the in-vehicle terminal device A, which has discovered the object, acquires location information (in other words, the narrowing down of the object's location) in the information provision system 1. Figure 11 is a diagram showing the positional relationship between the moving body M and the object O when the moving body M is moving. In Figure 11, it is assumed that the object O (a utility pole with a nest in this embodiment) is to the left in the direction of movement of the moving body M. As mentioned above, in this embodiment, the in-vehicle terminal device A is mounted so that the lens of the imaging unit 22 is facing forward in the direction of movement of the moving body M. Therefore, in the situation shown in Figure 11(A), the image recognition application 16 of the in-vehicle terminal device A discovers (recognizes) the object O which is ahead in the direction of movement of the moving body M. Subsequently, as the moving body M moves, the distance between the object O and the moving body M decreases, and as shown in Figure 11(B), the moving body M attempts to pass by the side of the object O. The position of the moving body M when it attempts to cross over object O is closer to the actual location of object O on Earth than the position of the moving body M (position of the in-vehicle terminal device A) when the image recognition application 16 detects object O. Therefore, in this embodiment, the position information of the in-vehicle terminal device A when the moving body M attempts to cross over object O is adopted as the position information of object O.
[0090] Figure 12 is a schematic diagram showing the change in the position of object O in the captured image over time as a moving object M approaches the object O. In Figure 12, the first captured image 130 is the image displayed by the imaging unit 22 at the first timing t1a (the image acquired for image analysis at the first timing t1a), the second captured image 132 is the image displayed by the imaging unit 22 at the second timing t2a (which is the next image analysis timing after the first timing t1a), which is later than the first timing t1a, and the third captured image 134 is the image displayed by the imaging unit 22 at the third timing t3a (which is the next image analysis timing after the second timing t2a), which is later than the second timing t2a.
[0091] In the first captured image 130, the object O is located near the center in the left-right direction and is captured at a first size. This first captured image 130 is, for example, an image taken when the moving object M and the object O are in the positional relationship shown in Figure 11(A). In the second captured image 132, the moving object M is closer to the object O than when the first captured image 130 was taken, and the object O is located to the left and below the position in the first captured image 130, slightly away from the edge of the image, and is captured at a second size, which is larger than the first size. In the third captured image 134, the moving object M is even closer to the object O than when the second captured image 132 was taken, and the object O is located to the left and below the position in the second captured image 132, on the lower left edge of the image, and is captured at a third size, which is larger than the second size. This third captured image 134 is, for example, an image taken when the moving object M and the object O are in the positional relationship shown in Figure 11(B). After this, object O moves out of frame from the captured image and disappears. In other words, the third captured image 134 is a photograph taken when object O is about to move out of frame from the captured image. In this embodiment, when object O is positioned close to the edge of the captured image, as in the third captured image 134, the in-vehicle terminal device A acquires location information based on the GPS receiver 23. More specifically, the third captured image 134 is the photograph taken just before object O moves out of frame from the captured image, that is, the last photograph taken before it moves out of frame. In this embodiment, as will be described later, the image recognition application 16 repeatedly performs image inference processing (processing to infer whether an object is included in the captured image), but the in-vehicle terminal device A acquires location information at the timing (time) of the inference processing one step prior to the inference processing in which the object disappeared from the captured image (the last inference processing in which it was inferred that an object was included in the captured image). The last captured image before the object moves out of frame is called the pre-loss frame.
[0092] Next, we will explain the operation of the CPU 12 of the in-vehicle terminal device A for acquiring the position information of the object. Before explaining the operation of the CPU 12, we will explain the parameters used for determining the identity of the object in the captured image. In this embodiment, in order to determine the identity of the object in the captured image, the movement range of the object is confirmed using the magnitude of the movement direction vector connecting the position of the object in the previous analysis image to the position of the object in the current analysis image, and the angle of the movement direction vector calculated this time with respect to the movement direction vector calculated last time.
[0093] First, the direction of movement vectors will be explained based on Figures 13(A) and 13(B). Figure 13(A) is a schematic diagram showing the direction of movement vector P representing the change in position of object O from the first timing t1a to the second timing t2a shown in Figure 12. The direction of movement vector P shown in Figure 13(A) (an example of the first vector) is a vector connecting the center of gravity position a of object O at the first timing t1a to the center of gravity position b of object O at the second timing t2a. Figure 13(B) is a schematic diagram showing the direction of movement vector Q representing the change in position of object O from the second timing t2a to the third timing t3a shown in Figure 12. The direction of movement vector Q shown in Figure 13(B) (an example of the second vector) is a vector connecting the center of gravity position b of object O at the second timing t2a to the center of gravity position c of object O at the third timing t3a. The centroid position is represented by xy coordinates with a predetermined point in the captured image as the origin. The magnitude of the direction of movement vector can be determined by calculating the distance between two centroid positions. Note that any known calculation method can be appropriately used to calculate the centroid position of the object O. Alternatively, the geometric center of the shape of the object O (in this embodiment, a utility pole with a nest) in the photograph can be calculated, and this geometric center can be assumed to be the centroid position to determine the direction of movement vector. In other words, either the centroid position or the center position can be used as the position of the object in the captured image. Any known method can also be appropriately used to calculate the center position.
[0094] After the third timing t3a shown in Figure 12, the object O is no longer within the field of view of the imaging unit 22 of the in-vehicle terminal device A mounted on the mobile body M, and therefore the object O is not captured in the image (see the dashed line shown in Figure 13(B)). In other words, the object O goes out of frame from the captured image. For this reason, when the lens of the imaging unit 22 is pointed forward in the direction of travel of the mobile body M, as in this embodiment, it is desirable for the in-vehicle terminal device A to acquire position information based on the GPS receiver 23 at the timing when the object O is about to go out of frame from the captured image. This is because the position of the in-vehicle terminal device A is close to the actual position of the object. For this reason, in this embodiment, as shown in the third captured image 134 in Figure 12, the in-vehicle terminal device A acquires position information (latitude and longitude) at the timing when the position of the object O in the captured image reaches the edge of the captured image.
[0095] In tracking an object O until it moves out of frame in the captured image, the in-vehicle terminal device A confirms the identity of the object O using movement direction vectors (movement direction vector P, movement direction vector Q) that indicate the change in the position of the object O. In other words, the in-vehicle terminal device A tracks the same object O while confirming whether the movement range of the object O is within a realistically possible range by determining whether the movement direction vectors (movement direction vector P, movement direction vector Q) exceed a predetermined distance threshold.
[0096] Next, the angle between the movement direction vector calculated this time and the movement direction vector calculated last time will be explained based on Figure 13(C). As shown in Figure 13(C), the angle between the movement direction vector Q (a vector showing the change in the position of the object from the second timing t2a to the third timing t3a) and the movement direction vector P (a vector showing the change in the position of the object from the first timing t1a to the second timing t2a) is represented by θ1. Hereafter, angle θ1 will represent the angle between the movement direction vector calculated this time and the movement direction vector calculated last time. In this embodiment, when tracking the object O until it moves out of frame from the captured image, the in-vehicle terminal device A also uses angle θ1 to confirm the identity of the object O. That is, the in-vehicle terminal device A tracks the same object O while confirming whether the movement range of the object O is within a realistically possible range by determining whether angle θ1 exceeds a predetermined angle threshold.
[0097] Next, we will explain in detail the operation of the CPU 12 of the in-vehicle terminal device A for acquiring the location information of the target object. When the image recognition application 16 is launched, the CPU 12 of the in-vehicle terminal device A executes the "forward image analysis processing" shown in Figure 14. In other words, the "forward image analysis processing" is included in the operation of the image recognition application 16. The "forward image analysis processing" is a detailed explanation of the processes (C), (D), (E), and (G) in the sequence diagram of Figure 5.
[0098] As shown in Figure 14, in the forward image analysis process, the CPU 12 reads the image captured by the imaging unit 22 (S001) and performs inference (analysis) to determine whether the read image contains an object (in this embodiment, a utility pole with a nest) (S002). Next, the CPU 12 performs blur and shake detection processing (S003). In blur and shake detection processing (S003), it is determined whether the image inferred in step S002 is an image with defects such as being out of focus. If the image is free of defects, the inference result from step S002 is kept for determination in step S004. If the image has defects, the inference result from step S002 is discarded (deleted).
[0099] Next, the CPU 12 determines whether the captured image contains an object based on the inference results for the captured image that was deemed free of defects in the blur / shake detection process (S004). If the result of the determination in step S004 is NO, the CPU 12 returns to step S001. In other words, the CPU 12 repeatedly performs inference (analysis) of the captured image until an object is found. Note that if the CPU 12 determined that there is a defect in the captured image in the blur / shake detection process in step S003, it will determine NO in step S004. Thus, in this embodiment, the CPU 12 does not determine whether an object is included based on the inference results for captured images that have defects such as out of focus. This improves the accuracy of object detection.
[0100] Meanwhile, if the result of step S004 is YES, the CPU 12 determines whether or not an object ID has already been assigned to the discovered object (S005). An object ID is unique identification information for an object, and one is assigned to each object. In other words, even if objects are of the same type, if they are different individuals, each individual will be assigned a different object ID.
[0101] If the result of the determination in step S005 is NO, that is, if an object has been newly discovered in the captured image, the CPU 12 assigns an object ID to that object (S006).
[0102] Then, following step S006, the CPU 12 performs the imaging process (S007). In the imaging process (S007), the imaging unit 22 takes a photograph, and the image file of that photograph is stored in the memory 14. In other words, the CPU 12 stores the image data of the image (specifically captured image) that includes the target object in the memory 14.
[0103] Next, the CPU 12 stores various information (object information other than image data) related to the object included in the specific captured image in the memory 14 as text data. The object information stored in the memory 14 includes various types of information as shown in Figure 15(A). Specifically, as shown in Figure 15(A), it includes the object ID assigned to the object, the name of the object (for example, "utility pole with nest"), the date and time of shooting, the position of the object in the captured image (x and y coordinates in the specific captured image), the file name of the image data, the storage location (path) of the image data, and information such as whether the object was on the left or right side in relation to the direction of travel of the moving body M. In other words, the object information stored in the memory 14 includes the inference results (AI inference results) and image information from the image recognition AI program 17 described above. Note that the object information may also include information on the orientation of the lens of the imaging unit 22 (forward in the direction of travel of the vehicle, etc.) and information on the position of the object in the captured image (left / right information, up / down information).
[0104] Next, the CPU 12 determines whether or not the object is about to move out of frame in the captured image, based on the positional information (coordinates in the captured image) of the object in the captured image that was analyzed (S009). If the result of the determination in step S009 is NO, that is, if the CPU 12 determines that it is not the time for the object to move out of frame, it returns to step S001.
[0105] Once an object is detected in the captured image and assigned an object ID, the captured image acquired in step S001 will include the object until the object with that object ID moves out of the frame of the captured image. Therefore, after determining NO in step S009, CPU 12 determines YES in step S004 and executes step S005. In step S005, if CPU 12 determines that the object included in the captured image is an object to which an object ID has already been assigned (YES in S005), it calculates a movement direction vector (S010) based on the information of the position of the object in the captured image (previous analysis image) from the previous analysis result (previous position) and the position of the object in the captured image (current analysis image) from the current analysis result (current position). As described above, the movement direction vector is a vector connecting the previous position to the current position, and is a vector that shows the change in the position of the object from the time of the previous image analysis to the time of the current image analysis (see movement direction vector P shown in Figure 13(A) and movement direction vector Q shown in Figure 13(B)).
[0106] Following step S010, the CPU 12 determines whether the magnitude of the movement direction vector calculated in step S010 is less than or equal to a predetermined distance threshold (S011). The distance threshold is determined through verification and other means to be a reasonable value as a realistically possible movement distance, taking into consideration the movement speed of the moving object M, the type of object, the interval of image analysis by the image recognition application 16, etc. The distance threshold may also be calculated by the CPU 12 based on measured values of the movement speed of the moving object M, etc. For example, the distance threshold can be set to about 10 pixels in the case of a captured image with horizontal and vertical pixels of 1280 x 720 pixels.
[0107] If the CPU 12 determines in step S011 that the magnitude of the movement direction vector is less than or equal to a predetermined distance threshold (YES in S011), it then determines whether the angle of the movement direction vector calculated this time relative to the movement direction vector calculated last time (see angle θ1 shown in Figure 13(C)) is less than or equal to a predetermined angle threshold (S012). The angle threshold, like the distance threshold, is determined through verification and other means to be a reasonable value as an angle that is realistically possible, taking into account the movement speed of the moving object M, the type of object, the interval of image analysis of the image recognition application 16, etc. The angle threshold may also be calculated by the CPU 12 based on measured values of the movement speed of the moving object M, etc. The angle threshold can be set to, for example, about 5°.
[0108] If the CPU 12 determines in step S012 that the angle θ1 is less than or equal to a predetermined angle threshold (YES in S012), it proceeds to step S009. In other words, if both the magnitude of the movement direction vector and the angle θ1 are less than or equal to their respective thresholds, the CPU 12 treats the object included in the current analysis image as the same object as the object associated with the object ID already assigned. Therefore, it proceeds to step S009 without performing the processing in steps S006 to S008 for the object included in the current analysis image.
[0109] In response, if the CPU 12 determines NO in step S011 (i.e., if it determines that the magnitude of the movement direction vector exceeds the distance threshold), or if it determines NO in step S012 (i.e., if it determines that the angle θ1 exceeds the angle threshold), it proceeds to step S006. In other words, if either the magnitude of the movement direction vector or the angle θ1 is below the threshold, the CPU 12 considers the object included in the current analysis image as a separate object from the object to which an object ID has already been assigned, and assigns a new object ID to that object (S006). Then, it takes a photograph including the object to which the new object ID has been assigned (S007) and stores the object information (S008). Thus, in this embodiment, if the position of an object in the analysis image has moved to a position that is not realistically possible (i.e., if the magnitude of the movement direction vector exceeds the distance threshold, or if the angle θ1 exceeds the angle threshold), it is treated as a separate object from the object to which an object ID has already been assigned (i.e., the object being tracked). In this way, the accuracy of object tracking is improved.
[0110] In step S009, if the CPU 12 determines that the object being tracked is about to move out of frame from the captured image, it acquires location information (latitude and longitude information) based on the GPS receiver 23 (S013). The CPU 12 then transmits this location information and object information relating to the object moving out of frame to the intermediate server X (S014). Note that both the object information and location information transmitted in step S014 are text data. Therefore, the load of this transmission is light.
[0111] Here, the object information includes information on whether the object was on the left or right side in relation to the direction of movement of the moving body M, as described above. The CPU 12 determines whether the object was on the left or right side in relation to the direction of movement of the moving body M based on whether the object goes out of frame to the left edge or the right edge of the captured image. In this embodiment, since the image is taken facing forward in the direction of movement of the moving body M, as shown in Figure 15(B), if the object goes out of frame from the left edge of the captured image, the CPU 12 determines that the object is on the left side in relation to the direction of movement of the moving body M, and if the object goes out of frame from the right edge of the captured image, the CPU 12 determines that the object is on the right side in relation to the direction of movement of the moving body M.
[0112] Next, CPU12 determines whether tracking of all objects assigned object IDs has been completed (i.e., whether all objects assigned object IDs have gone out of frame) (S015). If the result of this determination is NO, CPU12 returns to step S001. In other words, CPU12 repeats the image analysis (continues tracking of objects) until all objects assigned object IDs have gone out of frame.
[0113] On the other hand, if the result of step S015 is YES, the CPU 12 then determines whether or not an operation to terminate the image recognition application 16 has been performed (S016). This termination operation is performed using the user IF 20 of the in-vehicle terminal device A. If the CPU 12 determines NO in step S016, it returns to step S001. In other words, the CPU 12 continues image analysis by the image recognition application 16 until a termination operation is performed. On the other hand, if the CPU 12 determines YES in step S016, that is, if a termination operation using the user IF 20 has been performed, it terminates the forward image analysis process. In other words, it terminates the image recognition application 16. The determination process in step S009 in the forward image analysis process is performed for all object IDs. In this embodiment, the termination operation in step S016 is performed using the user IF 20 of the in-vehicle terminal device A, but it may also be an operation to turn off the electrical system using the ignition switch of the mobile body M (a configuration in which the forward image analysis process is terminated in conjunction with the OFF operation of the ignition switch).
[0114] As described in detail above, the information provision system 1 of the first embodiment comprises an in-vehicle terminal device A mounted on a mobile body M and an intermediate server X capable of communicating with the in-vehicle terminal device A. The in-vehicle terminal device A comprises an imaging unit 22, a determination means for determining whether a specific object (in this embodiment, a utility pole with a nest) is included in the image captured by the imaging unit 22, a text data transmission means for transmitting information relating to the specific object (such as the name of the object and its position in the captured image) as text data to the intermediate server X when the determination means determines that the specific object is included, a GPS receiver 23 (an example of a location information acquisition means) for acquiring the location information (latitude and longitude) of the in-vehicle terminal device A, and a location information transmission means for transmitting the location information acquired by the GPS receiver 23 to the intermediate server X when text data is transmitted by the text data transmission means (see Figure 5). The intermediate server X can transmit the text data and location information received from the in-vehicle terminal device A to a GIS server Y (an example of another information processing device). The CPU 12 that performs the processes shown in Figures 5 and 14 comprises a determination means, a text data transmission means, and a location information transmission means.
[0115] According to this information provision system 1 configuration, if an image captured by the imaging unit 22 of the in-vehicle terminal device A mounted on the mobile body M includes a target object (a utility pole with a nest), text data of information related to that object and the location information of the in-vehicle terminal device A are transmitted to the intermediate server X (see Figure 5). The intermediate server X then transmits this text data and location information to the GIS server Y (see Figure 5). Therefore, by attaching the in-vehicle terminal device A to a mobile body M such as a car that is moving for various purposes, it is possible to efficiently collect information about the target object incidentally while carrying out those purposes.
[0116] Furthermore, with this configuration of information provision system 1, users can access the GIS server Y from user terminal device B to check information about objects stored on the GIS server Y (see Figure 6). The information about objects stored on the GIS server Y and transmitted to user terminal device B is in text data format (see processing (J) in Figure 5 and processing (M) and (N) in Figure 6). Therefore, compared to a configuration where the image data of the captured image including the object is stored on the GIS server Y, there is no risk of providing images containing personal information (e.g., person 110 or other vehicle 112, see Figure 7) to a third party, thus reducing the possibility of personal information leakage.
[0117] Furthermore, in this embodiment of the information provision system 1, the in-vehicle terminal device A includes an image data transmission means that transmits image data of a specific image (captured image 100 shown in Figure 7) containing a specific object (a utility pole with a nest) captured by the imaging unit 22, and related to the transmitted object information, to the intermediate server X (processing (K) shown in Figure 5). The CPU 12 constitutes the image data transmission means.
[0118] According to this configuration of information provision system 1, not only text data about the object contained in the image captured by the in-vehicle terminal device A, but also the image data of the captured image itself can be stored on the intermediate server X. Therefore, by allowing users to view the image of the object upon request, it becomes possible to clearly show the condition of the object.
[0119] Furthermore, in this form of information provision system 1, the intermediate server X is equipped with a correction means that modifies the specific image (captured image 100 shown in Figure 7) so that non-specific objects (people 110, other vehicles 112) that are different from the specific object (utility pole with a nest) are unrecognizable (processing shown in (P) in Figure 6), and can transmit the image data of the corrected specific image (unclear image 120 shown in Figure 8) corrected by the correction means to the GIS server Y (processing shown in (Q) in Figure 6). The CPU 40 of the intermediate server X constitutes the correction means.
[0120] According to this information provision system 1 configuration, users of the GIS application 36 can access the GIS server Y from the user terminal device B to view images including the target object (a utility pole with a nest). However, these images are blurred images 120 in which people 110 and other vehicles 112 are obscured. Therefore, it is possible to prevent the leakage of personal information due to the provision of images containing personal information.
[0121] In this form of information provision system 1, the intermediate server X (an example of a specific server) performs corrections using correction means (the blurring process shown in (P) in Figure 6) and transmits the image data of the corrected specific image (the blurred image 120 shown in Figure 8) to the GIS server Y (the process shown in (Q) in Figure 6) in response to a request from the GIS server Y (an example of another server). The user can then view the blurred image 120 as a function of the GIS application 36 provided by the GIS server Y.
[0122] This configuration of information provision system 1 allows for the distribution of processing load between the intermediate server X and the GIS server, compared to a configuration where all processing, from the collection of information about the target object by the in-vehicle terminal device A to the provision of that information to the user terminal device B, is performed on a single server. Furthermore, it improves the maintainability and security of the servers.
[0123] Furthermore, in this embodiment of the information provision system 1, the image recognition application 16 of the in-vehicle terminal device A includes an image recognition AI program 17. In other words, the determination means according to the present invention determines whether a specific object is included in the image captured by the imaging unit 22 based on image recognition by AI. This makes it possible to suitably detect the object from the captured image.
[0124] Furthermore, in this configuration of information provision system 1, the information concerning the target object (utility pole with a nest) sent from the in-vehicle terminal device A to the intermediate server X includes information on the position of the target object in the image captured by the imaging unit 22 (for example, left and top, see Figure 15). Therefore, even devices other than the in-vehicle terminal device A can identify whether the target object is on the left or right side of the road based on this information.
[0125] Furthermore, in this form of information provision system 1, the object recognized by the image recognition application 16 of the in-vehicle terminal device A is a utility pole with a bird's nest on it. Therefore, it becomes possible to efficiently detect abnormalities in electrical equipment, such as bird nests being built on utility poles.
[0126] This embodiment discloses an in-vehicle terminal device A that can communicate with an intermediate server X (an example of a predetermined server) and can be mounted on a mobile body M, comprising: an imaging unit 22; a determination means for determining whether a specific object (in this embodiment, a utility pole with a nest) is included in an image taken by the imaging unit 22; a text data transmission means for transmitting information relating to the specific object (such as the name of the object and its position in the captured image) as text data to the intermediate server X when the determination means determines that the specific object is included; a GPS receiver 23 for acquiring location information (latitude and longitude) of the in-vehicle terminal device A; and a location information transmission means for transmitting the location information acquired by the GPS receiver 23 to the intermediate server X when text data is transmitted by the text data transmission means.
[0127] Furthermore, this embodiment discloses a program (image recognition application 16) that is executed by an in-vehicle terminal device A that can communicate with an intermediate server X and can be mounted on a mobile body M, and which causes the in-vehicle terminal device A to execute the following: a determination process (for example, the process in step S004 shown in Figure 14) that determines whether a specific object (in this embodiment, a utility pole with a nest) is included in an image taken by the imaging unit 22; an object information transmission process (for example, the process in step S014 shown in Figure 14) that transmits information relating to the specific object as text data to the intermediate server X if the determination process determines that the specific object is included; and a location information transmission process (for example, the process in step S014 shown in Figure 14) that transmits location information (latitude and longitude) acquired by a GPS receiver 23 capable of acquiring the location information of the in-vehicle terminal device A to the intermediate server X when text data is transmitted in the object information transmission process.
[0128] Furthermore, the information provision system 1 in this embodiment includes an in-vehicle terminal device A mounted on a mobile body M and an intermediate server X capable of communicating with the in-vehicle terminal device A. The in-vehicle terminal device A includes an imaging unit 22 capable of capturing images in the direction forward of the mobile body M's movement, a tracking means (processing such as steps S001 to S012 shown in Figure 14) that tracks the change in the position of a specific object (in this embodiment, a utility pole with a nest) in the image captured by the imaging unit 22 over time when the image includes such a specific object, a GPS receiver 23 (an example of a position information acquisition means) that acquires the position information (latitude and longitude) of the in-vehicle terminal device A, and a position information transmission means (processing such as steps S009, S013, S014 shown in Figure 14) that transmits the position information of the in-vehicle terminal device A to the intermediate server X when the specific object being tracked by the tracking means goes out of frame from the image captured by the imaging unit 22. The intermediate server X can then transmit the location information received from the in-vehicle terminal device A to the GIS server Y (an example of another information processing device) (see process (I) in Figure 5). The CPU 12 that executes the processes shown in Figures 5 and 14 constitutes the tracking means and the location information transmission means.
[0129] According to this configuration of information provision system 1, the in-vehicle terminal device A tracks the change in the position of an object (in this embodiment, a utility pole with a nest) over time in the captured image taken in front of the direction of travel of the moving body M, and transmits the position information (latitude and longitude) of the in-vehicle terminal device A when the object moves out of the frame of the captured image as the position information of the object to the intermediate server X. Therefore, compared to a configuration in which the position information of the in-vehicle terminal device A is sent to the intermediate server X immediately after detecting that the object is included in the captured image, it is prevented that the position of the object transmitted to the intermediate server X will deviate significantly from the actual location of the object on Earth. Thus, it is possible to provide a technology that can accurately acquire the position information of an object.
[0130] In this embodiment of the information provision system 1, the CPU 12 of the in-vehicle terminal device A tracks the object O captured at the first timing t1a and the object O captured at the second timing t2a as the same object O if the distance (magnitude of the movement direction vector P) between the first position a (center of gravity position a, see Figure 13(A)), which is the position of the object O (in this embodiment, a utility pole with a nest) in the first captured image 130 captured by the imaging unit 22 at the first timing t1a (see Figure 12), and the second position b (center of gravity position b, see Figure 13(A)), which is the position of the object O in the second captured image 132 captured by the imaging unit 22 at the second timing t2a (see Figure 12), which is less than or equal to a predetermined distance threshold (see steps S010 and S011 shown in Figure 14). The distance threshold is set to a realistic range of movement. Therefore, even if the position of the object changes in the captured images that are analyzed multiple times over time, it is possible to accurately track the object as if it were a single object.
[0131] Furthermore, in this embodiment of the information provision system 1, if the distance (magnitude of the movement direction vector P) between the first position a and the second position b (see Figure 13(A)) is greater than a predetermined distance threshold, the CPU 12 of the in-vehicle terminal device A starts tracking the object captured at the second timing t2a as a separate object from the object captured at the first timing t1a (step S006 is executed if the answer to step S011 shown in Figure 14 is NO). Therefore, compared to a configuration that does not track objects as separate objects, it is possible to prevent objects from being missed in the captured image.
[0132] Furthermore, in this information provision system 1, the CPU 12 of the in-vehicle terminal device A tracks the object O photographed previously (for example, at the second timing t2a) and the object O photographed this time (for example, at the third timing t3a) as the same object if the angle θ1 (see Figure 13(C)) of the currently calculated movement direction vector (for example, the movement direction vector Q calculated at the third timing t3a, see Figure 13(B)) relative to the previously calculated movement direction vector (for example, the movement direction vector P calculated at the second timing t2a, see Figure 13(A)) is less than or equal to a predetermined angle threshold (see step S012 shown in Figure 14). The angle threshold is set to a realistic range of movement. Therefore, even if the position in which the object is captured changes in the captured images which are analyzed many times over time, it is possible to accurately track the object as a single object.
[0133] Furthermore, in this embodiment of the information provision system 1, the CPU 12 of the in-vehicle terminal device A starts tracking the object captured this time (for example, at the third timing t3a, see Figure 13(B)) as a separate object from the object captured last time (at the second timing t2a) if the angle θ1 (see Figure 13(C)) of the movement direction vector calculated this time (for example, at the third timing t3a, see Figure 13(B)) relative to the movement direction vector calculated last time (for example, the movement direction vector P calculated at the second timing t2a, see Figure 13(A)) is greater than a predetermined angle threshold (step S006 is executed if the answer to step S012 shown in Figure 14 is NO). Therefore, compared to a configuration that does not track objects as separate objects, it is possible to prevent objects from being missed in the captured image.
[0134] Furthermore, in this embodiment of the information provision system 1, the in-vehicle terminal device A is equipped with object information transmission means that, when an image captured by the imaging unit 22 includes a specific object (in this embodiment, a utility pole with a nest), transmits specific information (such as AI inference results and image information) relating to that object to the intermediate server X. The intermediate server X can transmit the specific information and location information to the GIS server Y (see Figure 5). The CPU 12 constitutes the object information transmission means. With this configuration of the information provision system 1, it is possible to provide information relating to a specific object (in this embodiment, a utility pole with a nest) discovered while the mobile body M is moving, along with the object's precise location information (latitude and longitude), to the GIS server Y via the intermediate server X.
[0135] This embodiment discloses an in-vehicle terminal device A that can communicate with an intermediate server X (an example of a predetermined server) and can be mounted on a mobile body M, comprising: an imaging unit 22 capable of taking images toward the front in the direction of travel of the mobile body M; tracking means (processing such as steps S001 to S012 shown in Figure 14) for tracking the change in position of a specific object (in this embodiment, a utility pole with a nest) in the image taken by the imaging unit 22 when the image taken by the imaging unit 22 includes such a specific object; a GPS receiver 23 for acquiring the position information of the in-vehicle terminal device A; and position information transmission means (processing such as steps S009, S013, S014 shown in Figure 14) for transmitting the position information of the in-vehicle terminal device A to the intermediate server X when the specific object being tracked by the tracking means goes out of frame from the image taken by the imaging unit 22.
[0136] Furthermore, this embodiment discloses a program (image recognition application 16) that is executed by an in-vehicle terminal device A that can communicate with an intermediate server X and can be mounted on a mobile body M. The program causes the in-vehicle terminal device A to execute the following when an image captured by an imaging unit 22, which takes an image toward the front of the direction of travel of the mobile body M, includes an image capture unit 22 that
[0137] 2. Second Embodiment The information provision system of the second embodiment will be described below. In the information provision system of the second embodiment (hereinafter referred to as "information provision system 1A"), the in-vehicle terminal device A is mounted on the mobile body M with the lens of the imaging unit 22 pointed towards the rear in the direction of travel of the mobile body M (specifically, so that the lens is pointed approximately 180 degrees when the direction of travel is considered 0 degrees). In other words, the in-vehicle terminal device A takes images towards the rear in the direction of travel of the mobile body M. For this reason, the CPU 12 of the in-vehicle terminal device A in information provision system 1A performs rear image analysis processing (see Figure 19) instead of the forward image analysis processing (see Figure 14) described above. In the description of the second embodiment, components similar to those in information provision system 1 of the first embodiment are denoted by the same reference numerals and their explanation is omitted. In the second embodiment, the direction of the lens of the imaging unit 22 is acceptable as long as it is within the range of 150 to 210 degrees, with the direction of travel being 0 degrees. If the direction of the lens is within this range, it is considered to be in the category of being towards the rear in the direction of travel of the mobile body M.
[0138] When the in-vehicle terminal device A takes a picture facing backward in the direction of travel of the moving object M, as shown in Figure 16(A), the image recognition application 16 of the in-vehicle terminal device A detects (recognizes) the object O (a utility pole with a nest in this embodiment) located to the left of the direction of travel of the moving object M. As shown in Figure 16(B), the distance to the object O increases as the moving object M continues to move. Therefore, the position of the moving object M (the position of the in-vehicle terminal device A) when the image recognition application 16 detects the object O, that is, the position of the in-vehicle terminal device A (the position shown in Figure 16(A)) when the moving object M has just passed the object O, is the closest to the actual position of the object O on Earth. Thus, in this embodiment, the position information of the in-vehicle terminal device A when the moving object M has just passed the object O is used as the position information of the object O.
[0139] Figure 17 schematically shows the change in the position of object O in the captured image over time, when the in-vehicle terminal device A is taking pictures in the direction of backward movement of the moving object M, and the moving object M is moving away from the object O. Assume that object O is on the left side of the road R in the direction of movement of the moving object M, as shown in Figure 16. Because the image is taken in the direction of backward movement, object O appears to the right of the center in the horizontal direction of the captured image.
[0140] In Figure 17, the first captured image 150 is the image displayed by the imaging unit 22 at the first timing t1b (the image acquired for image analysis at the first timing t1b), the second captured image 152 is the image displayed by the imaging unit 22 at the second timing t2b (which is the next image analysis timing after the first timing t1b), which is later than the first timing t1b, and the third captured image 154 is the image displayed by the imaging unit 22 at the third timing t3b (which is the next image analysis timing after the second timing t2b), which is later than the second timing t2b.
[0141] In the first captured image 150, object O is positioned on the lower right edge of the image and is captured at a size of size 1. Assume that this first captured image 150 is the image taken when the moving object M and object O are in the positional relationship shown in Figure 16(A). That is, the first captured image 150 is the image taken when object O enters the frame of the captured image (in other words, when it appears in the captured image). In the second captured image 152, the moving object M is further away from object O than when the first captured image 150 was taken, and object O is positioned closer to the center and higher in the horizontal direction than in the first captured image 150, and is positioned slightly away from the edge of the image, and is captured at a size of size 2, which is smaller than size 1. In the third captured image 154, the moving object M is even further away from object O than when the second captured image 152 was taken, and object O is positioned closer to the center and higher in the horizontal direction than in the second captured image 152, and is positioned at the top of the image, close to the horizontal center, and is captured at a size 3, which is smaller than size 2. This third captured image 154 is assumed to be an image taken when the moving object M and the object O are in the positional relationship shown in Figure 16(B).
[0142] When photographing in the direction of backward movement of the moving object M, the position of the object O in the captured image changes over time as shown in Figure 17. Therefore, in this embodiment, when the object O is captured near the edge of the image (i.e., when it enters the frame of the image), as in the first captured image 150, the in-vehicle terminal device A acquires position information based on the GPS receiver 23. This makes the position information acquired by the in-vehicle terminal device A closer to the actual position of the object. More specifically, the first captured image 150 is the photograph taken immediately after the object O enters the frame of the image, i.e., the first photograph taken after it enters the frame. In this embodiment, the image recognition application 16 repeatedly performs image inference processing (processing to infer whether the image contains an object), and the in-vehicle terminal device A acquires position information at the timing (time) when the inference processing that caused the object to appear in the image (the first inference processing that inferred that the image contains an object) is performed.
[0143] Next, the operation of the CPU 12 for acquiring the position information of the object in the second embodiment will be described. Prior to describing the operation of the CPU 12, the movement direction vector in the second embodiment will be described based on Figures 18(A) and 18(B). Figure 18(A) is a schematic diagram showing the movement direction vector U that represents the change in the position of the object O from the first timing t1b to the second timing t2b shown in Figure 17. The movement direction vector U shown in Figure 18(A) (an example of the first vector) is a vector that connects the center of gravity position d of the object O at the first timing t1b to the center of gravity position e of the object O at the second timing t2b.
[0144] Prior to the first timing t1b shown in Figure 17, the object O is not within the field of view of the imaging unit 22 of the in-vehicle terminal device A mounted on the mobile body M, and therefore the object O is not captured in the image (see the dashed line shown in Figure 18(A)). In other words, the object O enters the frame of the captured image at the first timing t1b shown in Figure 17. For this reason, when the lens of the imaging unit 22 is pointed backward in the direction of travel of the mobile body M, as in the second embodiment, it is desirable for the in-vehicle terminal device A to acquire position information based on the GPS receiver 23 at the timing when the object O enters the frame of the captured image. This is because the position of the in-vehicle terminal device A is close to the actual position of the object. Therefore, in this embodiment, as shown in the first captured image 150 in Figure 17, the in-vehicle terminal device A acquires position information (latitude and longitude) at the timing when the object O enters the edge of the captured image.
[0145] Figure 18(B) is a schematic diagram showing the movement direction vector V representing the change in the position of object O from the second timing t2b to the third timing t3b shown in Figure 17. The movement direction vector V shown in Figure 18(B) (an example of the second vector) is a vector connecting the center of gravity position e of object O at the second timing t2b to the center of gravity position f of object O at the third timing t3b.
[0146] In the second embodiment, during tracking after the object O enters the frame of the captured image, the in-vehicle terminal device A confirms the identity of the object O using movement direction vectors (movement direction vector U, movement direction vector V) that indicate the change in the position of the object O. In other words, the in-vehicle terminal device A tracks the same object O while confirming whether the movement range of the object O is within a realistically possible range by determining whether the movement direction vectors (movement direction vector U, movement direction vector V) exceed a predetermined distance threshold.
[0147] Next, the angle of the currently calculated direction of movement vector relative to the previously calculated direction of movement vector in the second embodiment will be explained based on Figure 18(C). As shown in Figure 18(C), the angle of the direction of movement vector V (a vector indicating the change in the position of the object from the second timing t2b to the third timing t3b) relative to the direction of movement vector U (a vector indicating the change in the position of the object from the first timing t1b to the second timing t2b) is represented by θ2. Hereinafter, angle θ2 will represent the angle of the currently calculated direction of movement vector relative to the previously calculated direction of movement vector when the in-vehicle terminal device A is taking pictures facing backward in the direction of travel. In the second embodiment, when tracking an object O after it has entered the frame of the captured image, the in-vehicle terminal device A also uses angle θ2 to confirm the identity of the object O. That is, the in-vehicle terminal device A tracks the same object O while confirming whether the range of movement of the object O is within a realistically possible range by determining whether angle θ2 exceeds a predetermined angle threshold.
[0148] Next, we will describe the "rear image analysis processing" performed by the CPU 12 of the in-vehicle terminal device A in the second embodiment. In the second embodiment, when the image recognition application 16 is launched, the CPU 12 of the in-vehicle terminal device A performs the "rear image analysis processing" shown in Figure 19. In other words, the "rear image analysis processing" is included in the operation of the image recognition application 16 in the second embodiment.
[0149] The processes in steps S101 to S108 of the back-image analysis process shown in Figure 19 are the same as the processes in steps S001 to S008 of the forward-image analysis process (Figure 14) described above, so a detailed explanation is omitted. In the imaging process (S107) of the back-image analysis process (Figure 19), an image file of a photograph showing the object that has just entered the frame (in this embodiment, a utility pole with a nest), such as the first captured image 150 shown in Figure 17, is stored in memory 14. In addition, in the object information storage process (S108) of the back-image analysis process (Figure 19), the CPU 12 determines whether the object was to the left or right in the direction of travel of the moving body M based on whether the object entered the frame from the left or right edge of the captured image, and stores this information in memory 14. In the second embodiment, since the image is taken facing backward in the direction of travel of the moving body M, as shown in Figure 20, if the object enters the frame from the left edge of the captured image, the CPU 12 determines that the object is on the right side in the direction of travel of the moving body M, and if the object enters the frame from the right edge of the captured image, the CPU 12 determines that the object is on the left side in the direction of travel of the moving body M.
[0150] In the back-image analysis process (Figure 19), following step S108, the CPU 12 acquires position information (latitude and longitude information) based on the GPS receiver 23 (S109). In other words, in the second embodiment, the timing of the discovery of a new object is the timing when the object enters the frame of the captured image, so at this timing the CPU 12 acquires the position information and stores it in the memory 14.
[0151] Next, the CPU 12 determines whether or not to terminate tracking of the object (S110). Object tracking ends, for example, when the object that entered the frame of the captured image leaves the frame. In the example shown in Figure 17, the object enters the frame from the left edge of the captured image and leaves the frame from the top edge. The position of the object in the captured image is determined by the x and y coordinates within the captured image. Alternatively, the tracking of the object may be terminated when the in-vehicle terminal device A has moved a predetermined distance from the object O (for example, the second timing t2b shown in Figure 17), without tracking the object until the object leaves the frame of the captured image. In this case, it is sufficient to determine that the in-vehicle terminal device A has moved a predetermined distance from the object O by the position of the object in the captured image.
[0152] If the result of step S110 is NO, CPU 12 returns to step S101. In other words, CPU 12 continues tracking the object to which the object ID has been assigned.
[0153] In the back-image analysis process (Figure 19), if the CPU 12 determines YES in step S105, it calculates the direction of movement vector (S111), determines whether it is within the distance threshold (S112), and determines whether it is within the angle threshold (S113). Since each of these processes is the same as the processes in steps S010 to S012 in the forward-image analysis process (Figure 14) described above, a detailed explanation is omitted. If the result of the determination in step S112 is NO, or if the result of the determination in step S113 is NO, the CPU 12 proceeds to step S106. In other words, the CPU 12 treats the object included in the current analysis image as separate from the object to which an object ID has already been assigned.
[0154] On the other hand, if the result of step S113 is YES, the CPU 12 proceeds to step S110. That is, the CPU 12 treats the object included in the current analysis image as the same as the object to which an object ID has already been assigned. The processing in steps S111 to S113 improves the accuracy of object tracking. In other words, if the position of the object in the analysis image has moved to a position that is not realistically possible, the tracking accuracy of the object is improved by treating it as a different object from the object to which an object ID has already been assigned (i.e., the object being tracked).
[0155] Furthermore, if CPU 12 determines YES in step S110, it proceeds to step S114. The processing in steps S114 to S116 is the same as the processing in steps S014 to S016 in the forward image analysis process (Figure 14) described above, so a detailed explanation is omitted.
[0156] The movement direction vector calculated in step S111 is, as described above, a vector connecting the position of the object in the captured image (previous analysis image) from the previous analysis result (previous position) to the position of the object in the captured image (current analysis image) from the current analysis result (current position) (see movement direction vector U shown in Figure 18(A) and movement direction vector V shown in Figure 18(B)). In the determination in step S113, the CPU 12 determines whether the angle θ2 (see Figure 18(C)) of the movement direction vector calculated this time with respect to the movement direction vector calculated last time is less than or equal to a predetermined angle threshold. The determination process in step S010 in the back image analysis process is performed for all object IDs.
[0157] As described in detail above, the information provision system 1A of the second embodiment comprises an in-vehicle terminal device A mounted on a mobile body M and an intermediate server X capable of communicating with the in-vehicle terminal device A. The in-vehicle terminal device A comprises an imaging unit 22 capable of capturing images toward the rear in the direction of travel of the mobile body M, a tracking means (processing such as steps S101 to S113 shown in Figure 19) that tracks the change in the position of a specific object (in this embodiment, a utility pole with a nest) in the image captured by the imaging unit 22 over time when the image captured by the imaging unit 22 includes such an object, a GPS receiver 23 (an example of a position information acquisition means) that acquires the position information (latitude and longitude) of the in-vehicle terminal device A, and a position information transmission means (processing such as steps S104 to S109, S114 shown in Figure 19) that transmits the position information of the in-vehicle terminal device A to the intermediate server X when the specific object being tracked by the tracking means enters the frame of the image captured by the imaging unit 22. The intermediate server X can then transmit the location information received from the in-vehicle terminal device A to the GIS server Y (an example of another information processing device) (see process (I) shown in Figure 5). The CPU 12 that executes the processes shown in Figures 5 and 19 constitutes the tracking means and the location information transmission means.
[0158] According to this configuration of information provision system 1A, the in-vehicle terminal device A tracks the temporal changes in the position of a specific object (in this embodiment, a utility pole with a nest) in the captured image taken toward the rear in the direction of travel of the moving body M, and transmits the position information of the in-vehicle terminal device A when the object enters the frame of the captured image as the position information of the object to the intermediate server X. Therefore, compared to a configuration in which the position information of the in-vehicle terminal device A is sent to the intermediate server X some time after the detection of the specific object being included in the captured image, it is prevented that the position of the object transmitted to the intermediate server X will deviate significantly from the actual location of the object on Earth. Thus, it is possible to provide a technology that can accurately acquire the position information of an object.
[0159] In the information provision system 1A of the second embodiment, even if the image captured by the imaging unit 22 at one timing (for example, the second timing t2b shown in Figure 17) includes an object O, the CPU 12 of the in-vehicle terminal device A will not transmit the position information of the in-vehicle terminal device A at one timing to the intermediate server X if the object O entered the frame before one timing (in the example shown in Figure 17, it entered the frame at the first timing t1b). (If steps S104, S105, S112, and S113 shown in Figure 19 are all determined to be YES, S109 will not be executed.) This makes it possible for the CPU 12 to transmit only the position information closest to the actual location of the object on Earth to the intermediate server X.
[0160] Although the information provision system 1 of the first embodiment and the information provision system 1A of the second embodiment have been described above, the present invention is not limited to the embodiments described above and can be modified as appropriate without departing from the spirit of the invention.
[0161] In each of the embodiments described above, the intermediate server X is configured to perform blurring processing (processing (P) shown in Figure 6). Alternatively, the in-vehicle terminal device A may perform blurring processing and transmit the image data of the blurred image 120 to the intermediate server X. In other words, the in-vehicle terminal device A may be equipped with correction means to correct non-specific objects (e.g., people or other vehicles) that are different from specific objects (e.g., utility poles with nests) included in the image captured by the imaging unit 22 to make them unrecognizable, and the image data transmitted from the in-vehicle terminal device A to the intermediate server X may be the image data of the blurred image 120 (corrected specific image). In this case, the image data of the blurred image 120 is transmitted to the GIS server Y by the intermediate server X.
[0162] Furthermore, in each of the embodiments described above, when tracking the same object in the captured image (for example, a utility pole with a nest), the system is configured to perform both comparison with a distance threshold and comparison with an angle threshold. However, it is also possible to configure the system to perform only one of these processes. Even performing only one of these processes can improve the accuracy of object tracking compared to a configuration that does not perform either process. Moreover, if improving the accuracy of tracking the same object is not a consideration, the system may be configured to not perform either the comparison with the distance threshold or the comparison with the angle threshold.
[0163] Furthermore, in the first embodiment, the system is configured to acquire location information when the object is located at the edge of the captured image, but it is also possible to configure the system to acquire location information when the object has completely disappeared from the captured image. In other words, the in-vehicle terminal device A may acquire location information at the timing (time) when the image recognition application 16 performs an inference process that assumes the object has disappeared from the captured image. The first photograph taken after the object has gone out of frame from the captured image (the photograph taken immediately after it has gone out of frame) is called the lost frame. Both the location information when the object is about to go out of frame from the captured image (location information just before it goes out of frame) and the location information when the object has completely gone out of frame from the captured image (location information immediately after it has gone out of frame) are included in the "location information of the terminal device when a specific object goes out of frame from an image captured by the shooting unit" according to the present invention. Furthermore, "when it goes out of frame" according to the present invention is most preferably at the time of the last inference that the captured image includes the object, but it also includes the range of several inferences (for example, 2 to 3 times) before and after that time. Within this range, it is unlikely that the moving object has moved significantly, so there is no particular problem with the accuracy of the acquired location information. Furthermore, "when the image goes out of frame" according to the present invention may be within a predetermined time before or after the last inference that the image contains an object (for example, within 1 second before or after, or within 500ms before or after). Furthermore, "when the image goes in of frame" according to the present invention is most preferably when the image contains an object, but it also includes the time when inferences are performed several times before or after (for example, 2 to 3 times). Furthermore, "when the image goes in of frame" according to the present invention may be within a predetermined time before or after the first inference that the image contains an object (for example, within 1 second before or after, or within 500ms before or after).
[0164] Furthermore, in the first embodiment, if improvement of the accuracy of the object's position information is not considered, the in-vehicle terminal device A may acquire the position information at a time other than when the object moves out of frame from the captured image, such as when the in-vehicle terminal device A acquires the position information at the same time as the object is detected.
[0165] Furthermore, in the embodiments described above, the intermediate server X and the GIS server Y were configured separately, but it is also possible to provide a single server that handles the functions of both the intermediate server X and the GIS server Y. In this case, the user terminal device B would be an example of another information processing device.
[0166] Furthermore, while the above configurations involve sending the image data itself from the intermediate server X to the GIS server Y, it is also possible to send information indicating the location where the image data is stored (a so-called link). This information indicating the location of the image data is also considered an example of image data. Such a configuration can reduce the load on image data transmission.
[0167] Furthermore, in each of the embodiments described above, the type of information relating to the object transmitted from the in-vehicle terminal device A to the intermediate server X can be changed as appropriate. Specifically, for example, the system may transmit only the name of the object and location information based on the GPS receiver 23. Also, for example, if there is only one object to be discovered and tracked by the image recognition application 16 (the object being monitored), the information transmitted from the in-vehicle terminal device A to the intermediate server X may consist only of location information based on the GPS receiver 23. Also, for example, the system may not transmit information about the position of the object in the captured image (e.g., left and top). Also, for example, the system may not transmit image data of the captured image that includes the object. Furthermore, if the protection of personal information and reduction of transmission load are not considered, the system may not transmit the AI inference results, which are text data, but instead transmit image data of the captured image and location information based on the GPS receiver 23. In this case, the image data is an example of specific information.
[0168] Furthermore, in each of the embodiments described above, blurred images 120 were created by blurring people and other vehicles included in the captured image. However, the object to be blurred (non-specific object) may be something other than people or vehicles, such as a residence. Blurring the residence in the captured image is preferable from a security standpoint, as it may be possible to predict whether someone is home or out based on the exterior of the residence.
[0169] Furthermore, while the embodiments described above show examples where an object goes out of or into the frame from the left or right edge of the captured image, the CPU 12 may also be configured to determine that there is an object (e.g., a power line) above the road on which the moving body M is traveling (above the moving body M) based on the object going out of the frame from the upper edge of the captured image, or that there is an object (e.g., a manhole) on the road on which the moving body M is traveling (below the moving body M) based on the object going out of the frame from the lower edge of the captured image. In this case, the position from which the object enters the frame of the captured image may be any of the top, bottom, left, or right edges of the captured image. Furthermore, if the lens of the imaging unit 22 is oriented backward in the direction of travel of the moving body M, the CPU 12 may determine that there is an object (e.g., a power line) above the road on which the moving body M is traveling (above the moving body M) based on the object entering the frame from the upper edge of the captured image, or the CPU 12 may determine that there is an object (e.g., a manhole) on the road on which the moving body M is traveling (below the moving body M) based on the object entering the frame from the lower edge of the captured image. In this case, the position from which the object exits the frame of the captured image can be any of the top, bottom, left, or right edges of the captured image.
[0170] Furthermore, in the first embodiment described above, the position information of the object may be calculated by taking into account not only the position information of the in-vehicle terminal device A when the object goes out of frame from the captured image, but also information on which edge (top, bottom, left, or right) of the captured image the object entered and exited from. This configuration makes it possible to further improve the accuracy of the position information of the object. Furthermore, in the second embodiment described above, the position information of the object may be calculated by taking into account not only the position information of the in-vehicle terminal device A when the object enters frame from the captured image, but also information on which edge (top, bottom, left, or right) of the captured image the object entered and exited from.
[0171] Furthermore, in each of the embodiments described above, images were captured with the lens of the imaging unit 22 facing forward or backward in the direction of movement of the moving object M. However, the lens of the imaging unit 22 may also be directed to the left (approximately -90 degrees with the direction of movement being 0 degrees) in the direction of movement of the moving object M, or directed to the right (approximately 90 degrees with the direction of movement being 0 degrees) in the direction of movement of the moving object M. In this case, the timing for acquiring the location information of the tracked object may be when the object goes out of frame or when it enters frame. In this configuration, the field of view of the lens of the imaging unit 22 may be narrow. Alternatively, the imaging unit may be configured to use a 360-degree camera of the full-spherical or hemispherical type to capture images of the area around the moving object M. In this case as well, the timing for acquiring the location information of the tracked object may be when the object goes out of frame or when it enters frame.
[0172] Furthermore, in each of the embodiments described above, the specific object recognized (tracked) by the image recognition application 16 was a utility pole with a bird's nest, but the specific object can be changed as appropriate depending on the purpose. Specifically, for example, the specific object may be various types of power transmission and distribution equipment (utility poles, power lines, transmission towers, electrical manholes, etc.), or it may be other types of power transmission and distribution equipment in emergency situations other than utility poles with bird nests (utility poles covered with vines, power lines near construction site installations, electrical manholes around cracked roads, etc.). The image recognition AI program should be appropriately trained according to the type of specific object.
[0173] Furthermore, in the embodiments described above, the object of recognition (tracking) by the image recognition application 16 was an emergency power transmission and distribution facility such as a utility pole with a bird's nest built on it, but it could also be a communication facility such as an antenna or a traffic facility such as a road sign. With such a configuration, it would be possible to provide information about the object to communication carriers, road management companies, etc. [Explanation of Symbols]
[0174] 1… Information provision system 12…CPU 14…Memory 16…Image recognition app 17…Image Recognition AI Program 21...Communication IF 22…IMG Department 23…GPS receiver 100... Captured images 120... Unclear image 130...First image 132...Second image 134...Third image M...Moving object A... In-vehicle terminal device B...User terminal device X...Intermediate server Y...GIS server P...Movement direction vector Q... Direction vector θ1…Angle
Claims
1. Terminal devices mounted on mobile vehicles, An information provision system comprising a server capable of communicating with the aforementioned terminal device, The aforementioned terminal device is A camera unit capable of taking images, When an image captured by the aforementioned imaging unit includes a specific object, a tracking means for tracking the change in the position of the specific object in the image captured by the aforementioned imaging unit over time, A location information acquisition means for acquiring the location information of the terminal device, The system includes a location information transmission means that transmits location information of the terminal device to the server when the specific object being tracked by the tracking means moves out of the frame of the image captured by the imaging unit, or when it moves into the frame of the image captured by the imaging unit. The information provision system is characterized in that the server is capable of transmitting the location information to other information processing devices different from the server and the terminal device.
2. The information provision system according to claim 1, The aforementioned imaging unit is capable of capturing images in the direction forward of the moving body's direction of travel. The location information transmission means is characterized by transmitting the location information of the terminal device to the server when the specific object being tracked by the tracking means goes out of frame from the image captured by the imaging unit.
3. The information provision system according to claim 2, The tracking means is characterized in that, when the distance between a first position, which is the position of the specific object in an image captured by the imaging unit at a first timing, and a second position, which is the position of the specific object in an image captured by the imaging unit at a second timing following the first timing, is less than or equal to a predetermined distance threshold, the tracking means tracks the specific object captured at the first timing and the specific object captured at the second timing as the same object.
4. The information provision system according to claim 3, The tracking means is characterized in that, when the distance between the first position and the second position is greater than the distance threshold, it starts tracking the specific object photographed at the second timing as a different object from the specific object photographed at the first timing.
5. An information provision system according to any one of claims 1 to 4, The tracking means is characterized in that, when a vector connecting the position of the specific object in an image captured by the imaging unit at a first timing and the position of the specific object in an image captured by the imaging unit at a second timing following the first timing is defined as the first vector, and a vector connecting the position of the specific object in an image captured by the imaging unit at a third timing following the second timing is defined as the second vector, the tracking means tracks the specific object captured at the second timing and the specific object captured at the third timing as the same object when the angle of the second vector with respect to the first vector is less than or equal to a predetermined angle threshold.
6. The information provision system according to claim 5, The tracking means is characterized in that, if the angle of the second vector with respect to the first vector is greater than the angle threshold, it starts tracking the specific object photographed at the third timing as a different object from the specific object photographed at the second timing.
7. An information provision system according to any one of claims 1 to 4, The terminal device includes object information transmission means that, when the image captured by the imaging unit includes the specific object, transmits specific information relating to the specific object to the server. The information provision system is characterized in that the server is capable of transmitting the specific information and the location information to the other information processing device.
8. The information provision system according to claim 7, The terminal device is an information provision system characterized by comprising determination means for determining whether the image captured by the imaging unit contains the specific object based on AI-based image recognition.
9. The information provision system according to claim 1, The aforementioned imaging unit is capable of capturing images toward the rear in the direction of travel of the moving body. The location information transmission means is characterized by transmitting the location information of the terminal device to the server when the specific object being tracked by the tracking means enters the frame of an image captured by the camera unit.
10. The information provision system according to claim 9, The location information transmission means is characterized in that, even if the image captured by the imaging unit at a given time includes the specific object, if the specific object entered the frame before the given time, the location information of the terminal device at that time is not transmitted to the server.
11. A terminal device that can communicate with a designated server and can be mounted on a mobile device, A camera unit capable of taking images, When an image captured by the aforementioned imaging unit includes a specific object, a tracking means for tracking the change in the position of the specific object in the image captured by the aforementioned imaging unit over time, A location information acquisition means for acquiring the location information of the terminal device, A terminal device characterized by comprising: location information transmission means for transmitting location information of the terminal device to the server when the specific object being tracked by the tracking means goes out of frame from the image captured by the imaging unit, or when it enters the frame of the image captured by the imaging unit.
12. A program that can communicate with a designated server and is executed by a terminal device that can be mounted on a mobile device, When an image captured by an image capture unit includes a specific object, a tracking process is performed to track the change in the position of the specific object in the image captured by the image capture unit over time. A program for causing a terminal device to perform a location information transmission process, which transmits the location information of the terminal device to the server when the specific object being tracked by the tracking process moves out of the frame of the image captured by the imaging unit, or when it moves into the frame of the image captured by the imaging unit.
Citation Information
Patent Citations
Image processor and image processing method
JP2013196387A