Image processing device and explanatory text generation system
The image processing device addresses the lack of human-like descriptive text in conventional systems by assigning positional standards to vehicle images, enhancing search accuracy through natural-sounding descriptions that reflect driving conditions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- HITACHI LTD
- Filing Date
- 2024-11-20
- Publication Date
- 2026-06-01
AI Technical Summary
Conventional image caption generation techniques fail to generate natural descriptions of surrounding situations in vehicle images that reflect the presence or absence of objects based on traffic scenes, as they lack the tacit knowledge of human drivers, leading to discrepancies in descriptive text.
An image processing device that assigns positional relationship standards to vehicle images, using large-scale language models to generate descriptive text that reflects the intended operator's perspective, incorporating vehicle position and driving conditions to improve search accuracy.
Enhances search accuracy by generating descriptive text that aligns with human-intended positional relationships, improving the relevance of search results based on vehicle image data.
Smart Images

Figure 2026089184000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus and a description generation system.
Background Art
[0002] Conventionally, a technique called image caption generation for recognizing images and videos and generating descriptions of those images has been disclosed (for example, Patent Document 1). In the technique described in Patent Document 1, the surrounding objects moving into the image of an in-vehicle camera are recognized, and a text including the positional relationship between the vehicle and the surrounding objects is output to provide it as a description of the surrounding situation.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in conventional techniques such as Patent Document 1, it has not been particularly considered to generate a natural description of the surrounding situation as written by a human, which can express the presence or absence of surrounding objects to be focused on according to the traffic scene in which the vehicle is placed, rather than a sentence listing the surrounding objects in the image. In particular, when using a vehicle image database that searches and displays vehicle images collected from a connected car, it is desirable that a description close to the search query input by a human is given.
[0005] When generating descriptive text based on images acquired from in-vehicle cameras, a person with driving experience can explain the positional relationship with surrounding vehicles based on analogies of the vehicle's position and direction of travel within the image. However, large-scale language models that perform image understanding rely on the positions of surrounding vehicles within the image, which presents a challenge in that they may provide descriptive text that differs from the traffic conditions described by a human. In other words, for example, when a worker creates an explanatory document, the worker creates the document based on the tacit knowledge that "the explanatory document must be in a positional relationship relative to the driver's position." Such a document can be called a driver-referenced explanatory document. However, since generative AI does not necessarily possess such tacit knowledge, the explanatory document created by the generative AI will not necessarily be the aforementioned driver-referenced explanatory document.
[0006] This invention was made in consideration of these circumstances, and one of its objectives is for the AI to generate explanatory text in a positional relationship intended by the operator. For example, in the field of automobiles, the objective is to provide an image processing device that generates natural-sounding explanatory text about the surrounding situation, which can express the presence or absence of surrounding objects of interest according to the traffic scene in which the vehicle is placed. [Means for solving the problem]
[0007] One feature of this invention is, for example, the assignment of positional relationship standards to an image. [Effects of the Invention]
[0008] According to the present invention, for example, it becomes possible to generate descriptive text in a positional relationship intended by the operator, thereby improving the operator's search accuracy. For example, in the field of automobiles, when generating descriptive text for a vehicle image, image description data including the presence or absence of surrounding objects of interest based on the vehicle's position information within the image is generated, and driving condition description data is generated based on the image description data, thereby improving the search accuracy from search terms entered by a human. [Brief explanation of the drawing]
[0009] [Figure 1] This is a diagram illustrating the system configuration for Example 1. [Figure 2] This is a diagram showing the configuration of the image processing apparatus according to Example 1. [Figure 3] This is a detailed configuration diagram of the image processing device according to Example 1. [Figure 4] This figure shows an example of vehicle driving image data extracted from vehicle driving video data related to Example 1. [Figure 5] This is a flowchart of the explanatory text generation unit for Example 1. [Figure 6] This figure shows an example of camera mounting position data for Example 1. [Figure 7] This figure shows an example of vehicle driving control data related to Example 1. [Figure 8] This is a detailed flowchart of the process performed by the vehicle position information superimposition unit in Example 1. [Figure 9] This figure shows the procedure for correcting the vehicle's position information based on the camera mounting position data related to Example 1. [Figure 10] This figure shows the procedure for correcting the vehicle's position information based on the vehicle driving control data related to Example 1. [Figure 11] This figure shows superimposed image data of the vehicle's position, which has been corrected based on camera mounting position data and vehicle driving control data for Example 1. [Figure 12] This figure shows the details of the image description data model that defines the descriptive information for the image related to Example 1. [Figure 13] This flowchart details the processes performed by the image description information generation unit for Example 1. [Figure 14] This figure shows the image description data used by the large-scale language model unit for Example 1 to describe the image data with the vehicle's position superimposed on it, based on the image description data model. [Figure 15] This flowchart details the process of generating driving condition description data performed by the driving condition description generation unit in Example 1. [Figure 16] It is a diagram showing an example of driving situation explanation image data related to Example 1. [Figure 17] It is a flowchart of the search unit related to Example 1. [Figure 18] It is a diagram showing the image search interface of the input / output unit related to Example 1. [Figure 19] It is a display screen diagram showing the details of the search result of FIG. 18 related to Example 1. [Figure 20] It is a display screen diagram showing the details of the search result of FIG. 18 related to Example 1. [Figure 21] It is a plan view of the vehicle related to Example 1. [Figure 22] It is a configuration diagram of the image explanation system related to Example 2. [Figure 23] It is a diagram showing the details of the driving log related to Example 2. [Figure 24] It is a diagram showing an example of still image data which is a part of the vehicle driving video data related to Example 2. [Figure 25] It is a configuration diagram of the vehicle image analysis apparatus related to Example 2. [Figure 26] It is a configuration diagram of the necessity table related to Example 2. [Figure 27] It is a configuration diagram of the recognition table related to Example 2. [Figure 28] It is a flowchart of the explanatory text generation unit related to Example 2. [Figure 29] It is a hardware configuration diagram of the vehicle image analysis apparatus related to Example 2. [Figure 30] It is a detailed flowchart of the generation process of the image explanatory text related to Example 2. [Figure 31] It is a flowchart showing an example of the specific operation of the generation process of the image explanatory text described in FIG. 30 related to Example 2. [Figure 32] It is a detailed flowchart of the generation process of the image explanatory text for general roads related to Example 2. [Figure 33]This table shows the intermediate data obtained when image caption generation processing was performed on the still image data of Figure 24 related to Example 2. [Figure 34] This table shows the output data after the explanatory text generation unit has removed unnecessary data from the intermediate data in Figure 33 relating to Example 2. [Figure 35] This figure shows the details of the process for generating GPS explanatory text for Example 2. [Figure 36] This figure shows the details of the process for generating the control explanation text for Example 2. [Figure 37] This table shows an example of the explanatory text generated in Figures 35 and 36 relating to Example 2. [Figure 38] This is a table showing the instructions given to the large-scale language model unit used in the process of generating the driving condition description text for Example 2. [Figure 39] This table shows an example generated from a different image than the image caption for Figure 34 relating to Example 2. [Figure 40] This is a flowchart showing the processing of the search unit in Example 2. [Figure 41] This figure shows the image search interface of the input / output unit for Example 2. [Figure 42] This is the playback screen that appears when the search result (Scene 1) in Figure 41 related to Example 2 is clicked. [Figure 43] This is the playback screen that appears when the search result (Scene 2) in Figure 41 related to Example 2 is clicked. [Modes for carrying out the invention]
[0010] Examples 1 and 2 will be described below with reference to the drawings. [Examples]
[0011] Figure 1 is a diagram showing the configuration of the image explanation system 100. The image description system 100 is connected via a communication line 8 between the image processing device 1P and each of the connected cars 91, 92, and 93. Vehicles 91-93 use mounted onboard cameras 99 to photograph the area around them. These onboard cameras 99 may be configured to photograph the forward view (indicated by arrows in Figure 1), as illustrated in Figure 1, or they may be configured to photograph any desired view.
[0012] Vehicles 91-93 then transmit the following various data (vehicle driving log data) acquired while driving or stopped via communication line 8. • Vehicle driving video data 11AP is video data captured by the in-vehicle camera 99. • Camera mounting position data 11BP (see Figure 6 for details) indicates the mounting position data for the in-vehicle camera 99. • Vehicle driving control data 11CP (details in Figure 7) consists of vehicle control data such as speed, acceleration / deceleration, and steering angle. Communication line 8 collects various data transmitted from vehicles 91, 92, and 93 via the communication network and transmits it to image processing device 1P. Image processing device 1P collects the various data received via communication line 8 and stores it in the driving log DB 11P.
[0013] Figure 2 is a diagram showing the configuration of the image processing device 1P. The driving log DB11P stores vehicle driving video data 11AP, camera mounting position data 11BP, vehicle driving control data 11CP, and image description data model 11DP (see Figure 3 for details). The explanatory text generation unit 12P analyzes the driving conditions of each vehicle using the large-scale language model unit 13P based on the vehicle driving video data 11AP, camera mounting position data 11BP, and vehicle driving control data 11CP, and generates natural language text that describes the driving conditions of the vehicles. Therefore, the explanatory text generation unit 12P includes a vehicle position information superimposition unit 12AP, an image explanatory information generation unit 12BP, and a driving condition explanatory text generation unit 12CP (see Figure 3 for details).
[0014] The large-scale language model unit 13P is a generative AI called by the descriptive text generation unit 12P. The large-scale language model unit 13P generates image description data 15AP in text for images and generates driving condition description data 15BP which integrates multiple sentences and pieces of information. The description DB15P stores image description data 15AP, driving condition description data 15BP, and driving condition description image data 15CP (see Figure 3 for details). The search unit 16P performs a search of the description database 15P according to instructions from the input / output unit 17P and returns the results to the input / output unit 17P. The input / output unit 17P is operated by the user, receives a search query from the user, issues a query instruction to the search unit 16P, and provides a response to the user based on the search results from the search unit 16P.
[0015] Figure 3 is a detailed configuration diagram of the image processing device 1P. The driving log DB11P stores vehicle driving video data 11AP, camera mounting position data 11BP, vehicle driving control data 11CP, and image description data model 11DP. The vehicle position information superimposition unit 12AP generates vehicle position superimposed image data 114 (a superimposed image on which the large-scale language model unit 13P superimposes positional relationship references for explanation) by superimposing vehicle position information onto vehicle driving image data 111 (images captured by the onboard camera 99) extracted from vehicle driving video data 11AP.
[0016] The description DB15P stores image description data 15AP, driving condition description data 15BP, and driving condition description image data 15CP. The image description information generation unit 12BP generates image description data 15AP, which is descriptive information for the vehicle position superimposed image data 114, based on the image description data model 11DP (see Figure 14 for details). Therefore, the image description data model 11DP defines descriptive information for the vehicle driving image data 111 (objects to be described in the driving situation description) (see Figure 12 for details).
[0017] The driving condition description generation unit 12CP generates driving condition description data 15BP, which is a driving condition description for the vehicle position superimposed image data 114, based on the image description data 15AP (see Figure 16 for details). Furthermore, the driving situation description generation unit 12CP generates driving situation description image data 15CP by adding driving situation description data 15BP to the vehicle position superimposed image data 114 (see Figure 16 for details), and stores the generation result in the description DB 15P.
[0018] Figure 4 shows an example of vehicle driving image data 111 extracted from vehicle driving video data 11AP. This is still image data extracted from video footage of an onboard camera 99 of a vehicle driving on a highway. The vehicle driving image data 111 can be extracted, for example, at regular time intervals such as every 10 seconds, or by extracting 10 images at equal time intervals from a single video file.
[0019] Figure 5 is a flowchart of the explanatory text generation unit 12P. Processing will begin from S21. In S22, the vehicle position information superimposition unit 12AP generates vehicle position superimposed image data 114 by superimposing the vehicle position information onto the vehicle driving image data 111 (see Figure 8 for details). In S23, the image description information generation unit 12BP generates image description data 15AP for the vehicle position superimposed image data 114 based on the image description data model 11DP (see Figure 13 for details). In S24, the driving condition description generation unit 12CP generates driving condition description data 15BP based on the image description data 15AP (see Figure 15 for details). The process will be completed in S25. As a result, the image processing device 1P generates image description data 15AP that explains surrounding objects to be noticed and checked according to the vehicle's driving conditions, based on the vehicle's position information in the vehicle driving image data 111, which is an image from the in-vehicle camera 99. This enables the image processing device 1P to generate natural driving condition description data 15BP according to the traffic scene, improving the search accuracy for user search queries.
[0020] Figure 6 shows an example of camera mounting position data 11BP. The camera mounting position data 11BP stores the horizontal deviation Δ of the mounting position of the in-vehicle camera 99 as -10% and the horizontal angular deviation θ as -5 degrees.
[0021] Figure 7 shows an example of vehicle driving control data 11CP. The vehicle driving control data 11CP stores the vehicle's speed v as 65 [km] and the vehicle's steering angle φ as 3 [deg].
[0022] Figure 21 is a plan view of vehicle 91. A vehicle-mounted camera 99 is attached to the front of vehicle 91, with the front wheels 97 pointing slightly to the right and the rear wheels 98 pointing in the direction of travel. The camera mounting position data 11BP shows the horizontal deviation Δ and the horizontal angular deviation θ of the vehicle-mounted camera 99. In a typical vehicle 91, a rearview mirror is installed on the central axis of the vehicle 91. Therefore, the drive recorder is often installed in a position that does not obstruct the view of the occupants, resulting in a horizontal deviation Δ. The horizontal deviation Δ represents the horizontal deviation of the mounting position of the onboard camera 99 from the central axis of the vehicle. Furthermore, when installing a dashcam, a horizontal angle deviation θ occurs due to the curvature of the windshield and misalignment during installation. The vehicle driving control data 11CP shows the vehicle's speed v and the steering angle φ of the front wheels 97. The vehicle's speed v indicates the speed of the vehicle 91 in the direction of travel. The steering angle φ indicates the angle between the direction of travel of the vehicle 91 and the steering wheels.
[0023] Figure 8 is a detailed flowchart of the process (S22) executed by the vehicle position information superimposition unit 12AP. From S221, the vehicle position information superimposition unit 12AP starts processing. In S222, the vehicle position information superimposition unit 12AP acquires vehicle driving image data 111. In S223, the vehicle position information superimposition unit 12AP corrects the vehicle position information based on the camera mounting position data 11BP (see Figure 9 for details). In S224, the vehicle position information superimposition unit 12AP corrects the vehicle position information based on the vehicle driving control data 11CP (see Figure 10 for details). In step S225, the vehicle position information superimposition unit 12AP generates vehicle position superimposed image data 114 (see Figure 11 for details). At S226, the vehicle position information superimposition unit 12AP completes processing.
[0024] Figure 9 shows the procedure for correcting the vehicle's position information based on the camera mounting position data 11BP. In the vehicle driving image data 112, the initial position of the vehicle 1121 is the predicted path (driving line) of the initial position of the vehicle in the vehicle driving image data 111, and is defined as a line extending from bottom to top from the center point in the left-right direction on the camera image. Then, as a horizontal correction 1122, a correction equivalent to the horizontal deviation Δ is applied, and the vehicle position after horizontal correction 1123 is obtained. Next, a correction equivalent to the horizontal angle deviation θ is performed as a horizontal angle correction 1124, and the vehicle's position after horizontal angle correction 1125 is obtained.
[0025] Figure 10 shows the procedure for correcting the vehicle's position information based on vehicle driving control data 11CP. In the vehicle driving image data 113, the vehicle position after camera mounting position correction 1131 is the vehicle position after the camera mounting position has been corrected, and in Embodiment 1, it is the same as the vehicle position after horizontal angle correction 1125 in Figure 9. In contrast, the vehicle position after speed correction 1132, which indicates the driving position after a certain period of time in the direction of travel of 1131, is determined based on the speed v. Next, the vehicle's position 1133 after steering angle correction is determined based on the steering angle φ.
[0026] Figure 11 shows superimposed vehicle position image data 114, which is obtained by superimposing the vehicle's position 1141, which has been corrected based on camera mounting position data 11BP and vehicle driving control data 11CP. The vehicle position 1141 in Figure 11 is the same as the vehicle position 1133 after steering angle correction in Figure 10. On the other hand, the vehicle position 1141 only needs to be a position based on at least one of the camera mounting position data 11BP and the vehicle driving control data 11CP. The vehicle position information superimposition unit 12AP generates vehicle position superimposed image data 114 by superimposing the vehicle position 1141 onto the vehicle driving image data 111. Note that the vehicle itself is not visible from the field of view of the camera mounted on the vehicle. Therefore, the vehicle position 1141 is the predicted path (driving line) that the vehicle will take in the future from the time the vehicle position superimposed image data 114 was captured. Furthermore, the vehicle position 1141 is data that indicates the positional relationship reference for the large-scale language model unit 13P to provide explanations for the vehicle driving image data 111 captured by the onboard camera 99.
[0027] As a result, the large-scale language model unit 13P can generate natural image description data 15AP and driving situation description data 15BP based on the vehicle position superimposed image data 114, according to the positional relationship between the vehicle and surrounding vehicles and pedestrians. Therefore, the search accuracy of the driving situation description data 15BP for user search queries is improved. For example, the superimposed image data 114 shows the left lane 1142 and the right lane 1143 in which truck 1144 is driving, and truck 1145 is changing lanes from the left lane 1142 to the right lane 1143. Here, in the superimposed image data 114 of the vehicle's position, the vehicle's position 1141 is located in the right lane 1143, so the large-scale language model unit 13P can recognize that "the truck 1145 is located ahead in the same lane as the vehicle."
[0028] Figure 12 shows the details of the image description data model 11DP, which defines explanatory information for an image. The image description data model 11DP defines the information to be described as driving condition description data 15BP and examples of its description, as follows: • "Image summary" is a natural-sounding sentence that provides a general overview of the image. "Road information" refers to information that indicates the characteristics of the road a vehicle is traveling on, such as "road type" and "road shape." "Surrounding vehicles" refers to information that describes the characteristics of surrounding vehicles, such as "vehicle type," "location," "color," and "behavior." "Nearby pedestrians" refers to information that describes the characteristics of nearby pedestrians, such as their location, clothing, and behavior. "Environment" refers to information that describes the characteristics of the surrounding environment, such as "time of day," "weather," and "road surface conditions," which change depending on the date and time. These road information, surrounding object information (surrounding vehicles, surrounding pedestrians), and surrounding environment information (environment) are included in the image description data model 11DP as camera surrounding information.
[0029] The image description data model 11DP is prepared in advance by an administrator or other person and registered in the image processing device 1P. The image description data model 11DP is prepared for each applicable task (driving test, safety instruction, etc.), and the items used in that task (for example, objects to be avoided in a driving test) are registered in it.
[0030] Figure 13 is a flowchart detailing the process (S23) executed by the image description information generation unit 12BP. According to this flowchart, the image description information generation unit 12BP generates image description data 15AP based on the vehicle position superimposed image data 114. From S231, the image description information generation unit 12BP starts processing. In S232, the image description information generation unit 12BP reads the vehicle position superimposed image data 114. In S233, the image description information generation unit 12BP reads the image description data model 11DP. In S234, the image description information generation unit 12BP instructs the large-scale language model unit 13P to describe the vehicle position superimposed image data 114 based on the image description data model 11DP, and obtains image description data 15AP from the large-scale language model unit 13P. At S235, the image description information generation unit 12BP terminates processing.
[0031] As a result, the image processing device 1P defines the information to be explained according to the use case of image search and the explanatory parts that experts focus on as an image explanation data model 11DP, and can generate image explanation data 15AP and driving condition explanation text data 15BP, thereby improving the search accuracy for user search queries.
[0032] Figure 14 shows the image description data 15AP, which is an image description data 104 superimposed on the vehicle's position, explained by the large-scale language model unit 13P based on the image description data model 11DP. The large-scale language model unit 13P recognizes the position of the small truck 1145 (#3-A-1 in Figure 14) as "forward" (#3-A-3 in Figure 14) as image description data 15AP, which is the result of recognizing various information described in the image description data model 11DP in Figure 12 from the self-position superimposed image data 114 in Figure 11. This recognition is possible because the truck 1145 is located on the extension of the self-position 1141 shown in the self-position superimposed image data 114. As a result, descriptions of "position" such as #3-A-2 and #3-B-2 in Figure 14 are described in terms of relative position to the self-position, improving the accuracy of the traffic situation description in the driving situation description data 15BP generated from the image description data 15AP. On the other hand, the large-scale language model unit 13P compares this to the case where it recognizes various information described in the image description data model 11DP in Figure 12 from the vehicle driving image data 111 in Figure 4. In this case, the position of the truck, which appears small, remains abstract information that cannot narrow down its relative position to the vehicle's position, such as "right side in the image".
[0033] Figure 15 is a flowchart detailing the process (S24) for generating driving condition description data 15BP, which is performed by the driving condition description generation unit 12CP. According to this flowchart, the driving condition description generation unit 12CP generates driving condition description data 15BP based on the image description data 15AP. From S241, the driving condition description generation unit 12CP starts processing. In S242, the driving condition description generation unit 12CP reads the image description data 15AP. In S243, the driving situation description generation unit 12CP instructs the large-scale language model unit 13P to generate driving situation description data 15BP based on the image description data 15AP, and retrieves the driving situation description data 15BP from the large-scale language model unit 13P. At S244, the driving condition description generation unit 12CP terminates processing.
[0034] As a result, the image processing device 1P can generate driving condition description data 15BP with different language, character count, and writing style based on the image description data 15AP, without having to perform image understanding processing using a computationally expensive large-scale language model each time.
[0035] Figure 16 shows an example of driving condition explanatory image data 15CP. The driving situation description image data 15CP is data obtained by superimposing information on surrounding vehicles obtained from the image description data 15AP onto the vehicle position superimposed image data 114. The input / output unit 17P may display the driving condition description image data 15CP on its own, or it may display the driving condition description image data 15CP and the driving condition description text data 15BP together. This enables the image processing device 1P to more intuitively help the user understand the surrounding traffic objects in the image data described in the image description data 15AP.
[0036] Figure 17 is a flowchart of the search unit 16P. Processing will begin at S51. In S52, the input / output unit 17P receives the search query entered by the user. In S53, the similarity between the search query entered by the user in S52 and the driving condition description data 15BP generated in S24 is evaluated, and video information containing descriptions with high similarity is obtained. For similarity evaluation, cosine similarity search based on document vectorization may be used. In S54, the system searches for vehicle driving video data 11AP (and the vehicle driving image data 111 or self-position superimposed image data 114 generated from it), driving situation description data 15BP, related information, etc., related to the video information searched in S53, and generates a response text that includes these search results. Related information includes, for example, information on surrounding vehicles and information on surrounding pedestrians from the self-position. At S55, the response is sent to input / output unit 17P. Processing will be completed in S56. As a result, the image processing device 1P can quickly present the user with the desired video by having it search for video data containing driving condition description data 15BP that is similar to the natural language search query entered by the user.
[0037] Figure 18 shows the image search interface 61 of the input / output unit 17P. The image search interface 61 consists of a search text input unit 62 and a search result display unit 63. The search input section 62 receives the search text received from the user via the input / output section 17P in S52. The search results display unit 63 shows multiple search results 631, 632, 633, and 634 as vehicle driving video data 11AP (video thumbnail images) from S54.
[0038] Figure 19 is a display screen showing details of the search result 631 from Figure 18. Result 631 also displays the following information: • Display area 631a for vehicle driving video data 11AP that was found to have a high similarity to the search query. In Figure 19, as an example of vehicle driving video data 11AP, the vehicle position superimposed image data 114, in which the vehicle position 631d is superimposed, is displayed. Display field 631b is the display field 631a which contains the driving situation description data 15BP that is associated with the video in display field 631a. Display field 631c displays related information (image, time, weather, location, vehicle type, and driving speed) to the content of display field 631b.
[0039] Figure 20 is a display screen showing details of the search result 632 from Figure 18. Result 632 also displays the following information: • Display area 632a for vehicle driving video data 11AP that was found to have a high similarity to the search query. In Figure 20, as an example of vehicle driving video data 11AP, vehicle driving image data 111 before the vehicle's position was superimposed is displayed. Display field 632b is for the driving condition description data 15BP, which is associated with the video in display field 632a. Display field 632c displays related information (image, time, weather, location, vehicle type, and driving speed) to the content of display field 632b.
[0040] As shown in Figures 19 and 20, the image processing device 1P prompts the user to input a search query in natural language into the search query input unit 62. The image processing device 1P searches for videos by referring to driving condition description data 15BP, vehicle driving video data 11AP, and related information that are similar to the input search query, and displays the search results on the search result display unit 63. This makes it possible to improve search efficiency.
[0041] In the first embodiment described above, the vehicle's position and driving trajectory were superimposed on the image from the in-vehicle camera 99, and the positional relationship with surrounding vehicles, pedestrians, etc., was explained based on the superimposed image of the vehicle's position. [Examples]
[0042] Next, we will describe Example 2. For example, when constructing an image database used in vehicle development, it is desirable to pre-associate a description of each image in the database, so that a description similar to the search query entered by the user is provided. Example 2 was developed taking these circumstances into consideration, and describes an image description system 100Q (Figure 22) that databases images associated with natural-sounding descriptions that match search terms entered by humans. This image description system 100Q allows for the creation of a database of images associated with natural-sounding descriptions that match search terms entered by humans.
[0043] Figure 22 is a diagram showing the configuration of the image explanation system 100Q. In the image description system 100Q, there are vehicles 91-93, each a connected car, that can communicate with the vehicle image analysis device 1Q via the communication line 8. Each vehicle 91-93 transmits various measurement data measured by its own vehicle to the vehicle image analysis device 1Q's driving log DB 11Q via the communication line 8. Each vehicle, 91-93, is equipped with an on-board camera that captures images. The vehicle image analysis device 1Q generates a descriptive text based on the images captured by the on-board camera. Furthermore, the communication line 8 can be a general public network, whether wired or wireless, such as the fifth-generation mobile communication system, or 5G (5th Generation), which enables "massive simultaneous connections" and "ultra-low latency." By taking advantage of the features of newer mobile phone systems beyond 5G, it is also possible to expect effects such as online (real-time while driving) generation of explanatory text.
[0044] Figure 23 shows the details of the driving log DB11Q. The driving log DB11Q stores measurement data from each vehicle (91-93) according to the following data types. • Vehicle driving video data 11AQ, which is video data captured by an onboard camera (not shown) while the vehicle is in motion or stopped. • Vehicle GPS (Global Positioning System) data, specifically vehicle driving GPS data 11BQ. • Vehicle driving log data obtained from the in-vehicle ECUQ (Electronic Control Unit), etc., includes vehicle driving control data such as speed, acceleration / deceleration, and steering angle (11CQ). GPS is one example of a satellite positioning system.
[0045] Figure 24 shows an example of still image data 111Q, which is part of the vehicle driving video data 11AQ. Still image data 111Q is an example of data extracted from vehicle driving video data 11AQ, which was captured by each vehicle 91-93 traveling on the highway.
[0046] Figure 25 is a diagram showing the configuration of the vehicle image analysis device 1Q. The vehicle image analysis device 1Q includes, in addition to the driving log DB 11Q described in Figure 23, an explanatory text generation unit 12Q, a large-scale language model unit 13Q, an explanatory target setting unit 14Q, an explanatory text DB (image database) 15Q, a search unit 16Q, and an input / output unit 17Q. The explanatory target setting unit 14Q includes a necessity table 14AQ and a recognition table 14BQ. The large-scale language model unit 13Q includes a VQA (Visual Question Answer) unit 13AQ and a summary generation unit 13BQ. Furthermore, the various data stored within the vehicle image analysis device 1Q (driving log DB11Q, necessity table 14AQ, and recognition table 14BQ) may be stored in a storage device (not shown) outside the vehicle image analysis device 1Q, and may be configured to be accessible from the vehicle image analysis device 1Q via a network from that storage device.
[0047] The explanatory text generation unit 12Q analyzes the driving conditions of each vehicle based on the data in the driving log DB 11Q, using the large-scale language model unit 13Q and the explanation target setting unit 14Q, and generates natural language text that explains the driving conditions of the vehicles according to the analysis results. Therefore, the explanatory text generation unit 12Q analyzes whether the driving conditions of the vehicles fall into one of the pre-classified traffic scenes. The large-scale language model unit 13Q is called by the explanatory text generation unit 12Q. The large-scale language model unit 13Q is implemented by LAVIS (LAnguage VISion), which uses a natural language conversation scheme that interacts between question sentences and answer sentences as an interface, and has the following processing units. The VQA unit 13AQ responds to natural language queries about images in natural language. Therefore, the VQA unit 13AQ prepares an image recognition model, such as a CNN (Convolutional Neural Network), in advance using training data, and obtains a description of the corresponding image by inputting a query to that image recognition model. The summary generation unit 13BQ responds with a summary sentence, which is a description of the traffic situation, created by summarizing (integrating multiple sentences) the content of the input natural language text (prompt). The summary generation unit 13BQ may use existing services such as GPT-4, CLIP, BLIP, and BLIP-2 as text generation AI services that perform summarization and translation processing.
[0048] The explanation target setting unit 14Q refers to the following table and sets the explanation target corresponding to the traffic scene in the explanation text generation unit 12Q, and stores it in the explanation text DB 15Q. The necessity table 14AQ (Figure 26) associates explanatory objects with each traffic scene and defines the degree of importance of whether or not to mention each explanatory object in the explanatory text. The recognition table 14BQ (Figure 27) defines the detailed information to be mentioned in the descriptive text for each individual object being described. The description DB15Q stores the driving condition descriptions generated by the description generation unit 12Q.
[0049] Thus, the explanatory text generation unit 12Q executes the following (process 1) to (process 3). (Process 1) Identify the scenes in the image received from the camera, and for each identified scene, read recognition requirements information for the objects in that image from the requirements table 14AQ. (Process 2) Based on the recognition necessity information read, the system recognizes objects that require recognition from within the image and generates a descriptive text for each object from the recognition results. (Process 3) Based on the identified scene and the descriptive text for each object, a situational description of the image is generated, and the situational description of the image is associated with the image and stored in the description DB15Q.
[0050] Figure 26 is a diagram showing the configuration of the necessity table 14AQ. The explanation target setting unit 14Q sets the explanation targets for each traffic scene, such as highways, general roads, and parking lots. For example, the combination of the traffic scene "highway" and the object to be explained "pedestrian" is marked "Required / Required". The recognition requirement information "Required" on the left side of this "Required / Required" notation indicates that the object to be explained needs to be recognized from the image. The explanation requirement information "Required" on the right side of the "Required / Required" notation indicates that the object to be explained, whether recognized or not recognized from the image, needs to be explained in the explanatory text. Therefore, in the "general road" traffic scenario, pedestrian recognition is "necessary," and even if pedestrians are not recognized, an explanation is "necessary." On the other hand, in the "general road" traffic scenario, pedestrian crossings are "necessary," but an explanation is not required.
[0051] In this manner, the explanatory text generation unit 12Q reads information on whether an object in the image needs to be explained from the necessity table 14AQ. If an object designated as requiring an explanation based on the read explanation necessity information cannot be recognized from within the image, it generates an explanatory text for that object stating that the object does not exist in the image. This allows the system to generate descriptive text only for objects that are not normally present in a given traffic scenario, and to generate descriptive text even when objects that are normally present are not present. As a result, more natural traffic situation descriptions can be generated, improving the accuracy of user searches.
[0052] Figure 27 is a diagram showing the configuration of the recognition table 14BQ. The recognition table 14BQ defines, as detailed items, the items to be analyzed using the VQA unit 13AQ for the recognized explanatory object, and the items to be included in the explanatory text based on the analysis results. For example, if a pedestrian is detected, the VQA unit 13AQ analyzes their location, clothing color, and actions. In this way, the descriptive text generation unit 12Q reads detailed items about objects in the image from the recognition table 14BQ and generates descriptive texts for each object based on the detailed items read from the image. This makes it possible to individually set the information to be added depending on the object being described, and thus natural traffic condition descriptions can be generated. As a result, the search accuracy for user search queries is improved.
[0053] Figure 28 is a flowchart of the explanatory text generation unit 12Q. The description generation unit 12Q generates an image description by having the VQA unit 13AQ perform image analysis on the still image data 111Q extracted from the vehicle driving video data 11AQ, as part of the image description generation process (S21Q). The extraction process of the still image data 111Q is, for example, to extract images taken at regular intervals such as every 10 seconds, or to extract 10 images at equal time intervals from a single video file.
[0054] The description generation unit 12Q generates a GPS description based on the vehicle driving GPS data 11BQ as part of the GPS description generation process (S22Q). In this process, the target GPS data (positioning data) is either GPS data from the same time as the still image data 111Q in S21Q, or the GPS data from the closest time. In other words, the explanatory text generation unit 12Q adds an explanatory text to the image situation description that relates to at least one of the following pieces of information: information about the time of day when the image was taken, and information about the driving position when the image was taken, based on GPS data read from the vehicle's onboard GPS.
[0055] The explanatory text generation unit 12Q generates a control explanatory text based on the vehicle driving control data 11CQ as part of the control explanatory text generation process (S23Q). In this process, the control data to be used is the control data from the same time as, or the closest time to, the still image data 111Q in S21Q. In other words, the descriptive text generation unit 12Q adds a descriptive text to the image's situation description, based on vehicle driving control data read from the vehicle's onboard ECU (Electronic Control Unit) of the vehicle 91, relating to at least one of the following pieces of information: speed information, acceleration / deceleration information, and steering angle information.
[0056] The description generation unit 12Q generates a driving situation description by having the summarization generation unit 13BQ summarize the image description generated in S21Q, the GPS description generated in S22Q, and the control description generated in S23Q, as part of the driving situation description generation process (S24Q). The description generation unit 12Q associates the generated driving situation description with the still image data 111Q that is the subject of the driving situation description and stores it in the description database 15Q. In other words, the summary generation unit 13BQ generates a summary text in accordance with the input prompt. Then, the descriptive text generation unit 12Q generates an image contextual text by inputting a prompt to the summary generation unit 13BQ that includes information to be included in the contextual text, information to be excluded from the contextual text, and a contextual text generation instruction based on example contextual texts, as well as descriptive texts for each object.
[0057] Figure 29 is a hardware configuration diagram of the vehicle image analysis device 1Q. The vehicle image analysis device 1Q is configured as a computer 900 having a CPU 901, RAM 902, ROM 903, HDD 904, communication I / F 905, input / output I / F 906, and media I / F 907. The communication interface 905 is connected to an external communication device 915. The input / output interface 906 is connected to the input / output device 916. The media interface 907 reads and writes data to the recording medium 917. Furthermore, the CPU 901 controls each processing unit by executing a program (also called an application or app) loaded into the RAM 902. This program can also be distributed via a communication line or by recording it on a recording medium 917 such as a CD-ROM.
[0058] Figure 30 is a detailed flowchart of the image caption generation process (S21Q in Figure 28). The description generation unit 12Q classifies the traffic scene captured in the still image data 111Q by querying the VQA unit 13AQ (S211Q). The VQA unit 13AQ receives the still image data 111Q and a traffic scene query (for example, "Where is this scene? For example, is it a road, highway, or parking lot?") and responds to the description generation unit 12Q with a reply (for example, highway) (line A01 in Figure 33). Note that the description generation unit 12Q reads the traffic scene query that has been set in the vehicle image analysis device 1Q in advance by the administrator, so users of the vehicle image analysis device 1Q do not need to generate the traffic scene query themselves. The explanatory text generation unit 12Q then executes the loop processing from S212Q to S217Q by sequentially selecting the items to be explained from the necessity table 14AQ (pedestrians, bicycles, cars, etc.). The items to be explained selected in this loop processing will be referred to as the selected items below.
[0059] The explanatory text generation unit 12Q determines whether recognition of the selected object is necessary in the traffic scene identified in S221Q by referring to the necessity table 14AQ (S212Q). If the answer in S212Q is Yes (necessary), the process proceeds to S213Q; otherwise, it proceeds to S217Q. The descriptive text generation unit 12Q queries the VQA unit 13AQ of the large-scale language model unit 13Q to determine whether the selected object exists in the still image data 111Q. This query is, for example, "Are there any pedestrians in this scene?" (row A02 in Figure 33). Based on the response to this query (the object recognition result), the descriptive text generation unit 12Q recognizes the object that needs to be recognized from within the image. As a result of this recognition, the descriptive text generation unit 12Q determines whether the selected object exists in the still image data 111Q (S213Q). If the answer in S213Q is Yes (it exists), the process proceeds to S214Q; otherwise, it proceeds to S215Q.
[0060] The description generation unit 12Q obtains detailed items about the selected object present in the still image data 111Q (S214Q). Therefore, the description generation unit 12Q refers to the recognition table 14BQ to obtain the detailed items corresponding to the selected object. Then, the description generation unit 12Q queries the VQA unit 13AQ of the large-scale language model unit 13Q for each of the detailed items about the selected object in the still image data 111Q. For example, the description generation unit 12Q determines that the selected object is a car, and the detailed items corresponding to a car in the recognition table 14BQ include "color," so it generates the query "What color is the car?".
[0061] The explanatory text generation unit 12Q determines whether or not an explanation (mention) of the selected object is necessary by referring to the necessity table 14AQ (S215Q). If the answer in S215Q is Yes (necessary), proceed to S216Q; otherwise, proceed to S217Q. The description generation unit 12Q generates a description of the selected object by one of the following means (S216Q): If the answer to S213Q is Yes, a description of the selected object is generated from the combination of the query statement and the answer statement for the detailed items of the selected object obtained in S214Q. If the answer to S213Q is No, a descriptive message is generated indicating that the selected object was not recognized in the still image data 111Q (e.g., line D07 in Figure 39). The explanatory text generation unit 12Q determines whether processing for all selected objects has been completed after finishing processing for the current selected object (S217Q). If the answer in S217Q is Yes (completed), processing ends; otherwise, it switches to the unprocessed selected object and returns to S212Q. This enables the generation of natural-sounding traffic situation descriptions that include explanations of surrounding objects that should be noticed and checked depending on the traffic scene in which the vehicle is traveling, thereby improving the accuracy of searches based on user queries.
[0062] Figure 31 is a flowchart showing an example of the specific operation of the image caption generation process described in Figure 30. The explanatory text generation unit 12Q branches into processing for each traffic scene (S301Q) according to the result of the processing that classifies the traffic scenes captured in the still image data 111Q (S211Q in Figure 30). • If the classification result is for a public road, generate an image description for public roads (S302Q). • If the classification result is for a highway, generate an image description for highways (S303Q). • If the classification result is for a parking lot, generate an image description for the parking lot (S304Q). • If the classification result is "Other," generate an image description specified for "Other" (S305Q).
[0063] Figure 32 is a detailed flowchart of the process for generating image captions for general roads (S302Q in Figure 31). The explanatory text generation unit 12Q generates a question and answer about traffic scenes (S211Q in Figure 30) and a question and answer about the presence or absence of traffic signals (S311Q). The explanatory text generation unit 12Q performs the determination process (S213Q) shown in Figure 30, asking whether there are pedestrians in the image (S312Q). If there are, it proceeds to S313Q; otherwise, it proceeds to S314Q. The explanatory text generation unit 12Q generates, as answers to questions about the detailed items of the pedestrian (S214Q in Figure 30), an answer indicating the presence of the pedestrian, a question and answer regarding the pedestrian's location, a question and answer regarding the pedestrian's color, and a question and answer regarding the pedestrian's actions (S313Q). The explanatory text generation unit 12Q generates a response indicating that there are no pedestrians (S314Q).
[0064] The explanatory text generation unit 12Q performs the determination process (S213Q) shown in Figure 30, asking whether there is a car in the image (S315Q). If there is a car, it proceeds to S316Q; otherwise, it proceeds to S317Q. The explanatory text generation unit 12Q generates, as answers to questions about the details of the automobile (S214Q in Figure 30), an answer indicating the presence of the automobile, a question and answer regarding the location of the automobile, a question and answer regarding the type of automobile, a question and answer regarding the color of the automobile, and a question and answer regarding the operation of the automobile (S316Q). The explanatory text generation unit 12Q performs the determination process (S213Q) shown in Figure 30, asking whether there is a bicycle in the image (S317Q). If there is, it proceeds to S318Q; otherwise, it proceeds to S319Q. The explanatory text generation unit 12Q generates, as answers to questions about the detailed items of the bicycle (S214Q in Figure 30), an answer indicating the presence of the bicycle, a question and answer regarding the bicycle's location, a question and answer regarding the bicycle's color, and a question and answer regarding the bicycle's operation (S318Q).
[0065] The explanatory text generation unit 12Q performs the determination process shown in Figure 30 (S213Q), asking whether there is a pedestrian crossing in the image (S319Q). If there is, it proceeds to S320Q; otherwise, it terminates the process. The explanatory text generation unit 12Q generates a response indicating the presence of a pedestrian crossing (S320Q). As explained above in Figure 31, the process for generating image captions for general roads is similar, but the process for generating image captions for other traffic scenes (S303Q~S305Q) is the same.
[0066] Figure 33 is a table showing the intermediate data obtained by performing the image description generation process (S21Q) on the still image data 111Q from Figure 24. In this table, each row consists of a question and its corresponding answer, which together form an image description. For example, line A01 is the result of the process (S211Q) that classifies traffic scenes. Lines A02 to A04 show the results of the process (S213Q) that determines whether or not the selected object exists within the still image data 111Q. Rows A05 to A13 are the results of the process (S214Q) for obtaining detailed information about the selected object.
[0067] Figure 34 is a table showing the output data after the explanatory text generation unit 12Q has removed unnecessary data from the intermediate data in Figure 33. The difference from Figure 33 is that the explanatory text for pedestrians (row A02), the explanatory text for bicycles (row A03), and the explanatory text for toll booths (row A13), which were deemed unnecessary when not recognized in the necessity table 14AQ of the explanatory target setting unit 14Q, have been removed in Figure 34.
[0068] Figure 35 shows the details of the GPS description generation process (S22Q). The explanatory text generation unit 12Q generates an explanation (S221Q) that classifies the time of shooting of the still image data 111Q into morning, noon, evening, night, etc., according to the time of day of the GPS information obtained from the vehicle driving GPS data 11BQ. The explanatory text generation unit 12Q generates an explanation (S222Q) that identifies the location where the still image data 111Q was taken, such as a major arterial road or a city name, according to the latitude and longitude of the GPS information obtained from the vehicle driving GPS data 11BQ. This allows for the generation of traffic condition descriptions that include information about the time of day and location of vehicle travel. As a result, the accuracy of user searches is improved.
[0069] Figure 36 shows the details of the control explanation text generation process (S23Q). The explanatory text generation unit 12Q generates explanatory text regarding driving speed control from the vehicle driving control data 11CQ (S231Q). The explanatory text may state, for example, that speeds below 20 km / h are low speed, 20 km / h to 60 km / h are medium speed, and 60 km / h and above are high speed. The explanatory text generation unit 12Q generates explanatory text regarding acceleration and deceleration control from the vehicle driving control data 11CQ (S232Q). The explanatory text may state, for example, that the vehicle is accelerating if the acceleration is above a certain level, decelerating if the deceleration is above a certain level, or constant speed otherwise. The explanatory text generation unit 12Q generates explanatory text related to steering control from the vehicle driving control data 11CQ (S233Q). The explanatory text may include, for example, left turn state, right turn state when the steering angle is above a certain level, and straight-ahead state in all other cases. This allows for the generation of more natural traffic situation descriptions that include explanations of the vehicle's control state and behavior. As a result, it has the effect of improving the accuracy of searches based on user queries.
[0070] Figure 37 is a table showing an example of the explanatory text generated in Figures 35 and 36. In line B01, the time zone description is generated in S221Q. In line B02, a description of the driving location is generated in S222Q. In line B03, the description of the driving speed is generated in S231Q. In line B04, a descriptive text about the acceleration and deceleration state is generated in S232Q. In line B05, a description of the steering state is generated in S233Q.
[0071] Figure 38 is a table showing the instructions given to the large-scale language model unit 13Q used in the process of generating driving condition descriptions (S24Q). Line C01 contains the instructions for generating a traffic condition description. Line C02 contains the information to be included in the traffic condition description. Line C03 contains information to be excluded from the traffic condition description. Line C04 contains an example of a traffic situation description.
[0072] Then, in the process of generating the driving condition description (S24Q), the description generation unit 12Q generates a prompt by sequentially combining the following texts from (item 1) to (item 4). Note that at least one of (item 2) and (item 3) may be omitted. (Item 1) Instructions for the large-scale language model unit 13Q (Figure 38) (Item 2) GPS explanation (lines B01 and B02 in Figure 37) (Item 3) Control description (lines B03, B04, and B05 in Figure 37) (Item 4) Image caption (Figure 34)
[0073] The descriptive text generation unit 12Q inputs the generated prompt to the large-scale language model unit 13Q to obtain a traffic situation description written in natural language. Below is an example of a traffic situation description generated by the large-scale language model unit 13Q from the prompt (corresponding to situation description 732Q in Figure 42 below). Traffic situation description: "It is noon, and you are driving straight on the Metropolitan Expressway Route 5 at high speed, slowing down considerably. In this scene, there are solid orange lane markings on the expressway. Additionally, a white truck is driving ahead of you on the road." This improves search accuracy for user searches by generating natural-sounding traffic condition descriptions tailored to usage and user preferences.
[0074] Figure 39 is a table showing an example generated from an image different from the image caption in Figure 34. The image caption for Figure 39 is for images taken from a vehicle traveling on a public road. Therefore, since the combination of public road and pedestrian in the necessity table 14AQ is "necessary" for explanation, the information "pedestrian was not recognized" is included in row D07. The traffic situation description generated by the large-scale language model unit 13Q through the driving situation description generation process (S24Q) from the prompt containing the image description in Figure 39 is as follows: Traffic situation description: "You are currently driving at a slow, steady speed on a city road in Mito City at night, preparing to turn left. There are no pedestrians in this scene."
[0075] The above explains the process of creating a database of driving condition descriptions (up to saving them to the description DB15Q) with reference to Figure 39. The following describes examples of how the database information can be used. Figure 40 is a flowchart showing the processing of the search unit 16Q. The search unit 16Q searches for images that match the input search query from the images stored in the description database 15Q. Specifically, the search unit 16Q outputs images as search results that have a high degree of similarity between the input search query and the situational descriptions stored in the description database 15Q. The search unit 16Q may, for example, list images in descending order of similarity in the search results and output the top 1 to X images (relative similarity determination), or it may output search results images whose similarity is higher than a pre-set threshold value (threshold Y) (absolute similarity determination). The following describes the processing details of the search unit 16Q.
[0076] The search unit 16Q receives a search query entered by the user from the input / output unit 17Q (S61Q). In this case, the user is, for example, a commentator at a traffic control center that manages highways, and the search query is, for example, a request to collect images from the database of situations similar to an accident that occurred at a specific location on a highway at a specific time. The commentator plans to edit the image materials obtained from the database to produce a news program about the accident that occurred. The search unit 16Q evaluates the similarity between the user's search query and the descriptions stored in the description database 15Q (S62Q), and retrieves video information (image information) associated with the descriptions with high similarity from the description database 15Q. For similarity evaluation, cosine similarity search based on document vectorization may be used.
[0077] The search unit 16Q retrieves not only the driving video and driving situation description acquired by S62Q, but also related information about the driving situation (such as weather information that is not in the description DB15Q but can be obtained from the weather database by specifying the location and date of the driving situation description), and generates a response based on these results (S63Q). In other words, the search unit 16Q may output related information obtained from a database other than the description DB 15Q as part of the search results, based on the information contained in the situational description corresponding to the image with a high similarity. The search unit 16Q sends the answer text from S63Q to the input / output unit 17Q (S64Q). This allows users to quickly find the desired video when searching for video data containing traffic condition descriptions similar to the natural language search query they entered.
[0078] Figure 41 shows the image search interface 71Q of the input / output unit 17Q. The image search interface 71Q consists of a search input section 72Q, which is the input field for the search text of S61Q, and a search result display section 73Q, which is the display field for the answer text of S64Q. Multiple search results (scenes 1 to 4) are displayed as icons or thumbnail images in the search result display section 73Q.
[0079] Figure 42 shows the playback screen when the search result (Scene 1) in Figure 41 is clicked. The following information is displayed on this playback screen from top to bottom. • Image 731Q from the search results. The description of the situation in image 731Q, 732Q, was extracted because of its high similarity to the search query. • Additional information such as GPS description (time, location), control description (vehicle type, driving speed), and weather (733Q).
[0080] Figure 43 shows the playback screen when the search result (Scene 2) in Figure 41 is clicked. Similar to Figure 42, this playback screen displays the image 741Q of the search result, its situation description 742Q, and additional information 743Q, just as in Figure 42. This allows users to improve search efficiency by entering search terms in natural language and then searching for videos by referencing traffic condition descriptions, videos, and related information that are similar to their search terms.
[0081] According to the embodiment 2 described above, when generating a description of a vehicle image, the description generation unit 12Q refers to the necessity table 14AQ and generates a natural description that includes the presence or absence of surrounding objects of interest based on the traffic scene in which the vehicle is placed. As a result, a natural description is created in the database that mentions necessary surrounding objects while omitting unnecessary ones, thereby improving the accuracy of database searches from search queries entered by humans.
[0082] The contents of Example 2 and Example 1 can be combined or partially substituted, as illustrated below. • As input data for the image description generation process (S21Q in Figure 28) by the description generation unit 12Q of Example 2, the still image data 111Q extracted from the vehicle driving video data 11AQ is replaced with vehicle position superimposed image data 114, which is obtained by superimposing vehicle position information onto the vehicle driving image data 111 of Example 1. In Example 2, the processes S22Q to S24Q in Figure 28 (such as the process in S24Q that causes the large-scale language model unit 13Q to generate a traffic situation description) also process the vehicle position superimposed image data 114 instead of the still image data 111Q. In other words, the image processing device 1P may perform S21 and S22 in Figure 5 to generate the vehicle position superimposed image data 114, which can then be treated as still image data 111Q for the vehicle image analysis device 1Q, and the vehicle image analysis device 1Q may perform S21Q to S24Q in Figure 28.
[0083] Furthermore, the hardware configuration of the image processing device 1P in Example 1 is the same as the hardware configuration of the vehicle image analysis device 1Q in Example 2 shown in Figure 29. Moreover, the image processing device 1P in Example 1 and the vehicle image analysis device 1Q in Example 2 may be configured as the same device housed in the same enclosure. For example, the large-scale language model unit 13P may have the same functions as the large-scale language model unit 13Q. Furthermore, each processing unit in the image processing device 1P of Example 1 and each processing unit in the vehicle image analysis device 1Q of Example 2 may operate on a single computer 900 (Figure 29), or they may be distributed and operated across multiple computers 900.
[0084] Furthermore, by combining the contents of Example 2 and Example 1, the following explanatory text generation system can be constructed. The description generation system comprises a processing unit (description generation unit 12P) of an image processing device 1P, and a database (description DB 15P, description DB 15Q) that can be searched by the operator using natural language. The explanatory text generation unit 12P performs the following processing. [First Processing] The generation model (large-scale language model unit 13Q) receives scene information indicating the scene recognized from the captured image (vehicle driving image data 111) or superimposed image (vehicle position superimposed image data 114) (S211Q in Figure 30, processing 1 in Figure 25). [Second Process] Based on whether or not an object needs to be described for each scene (necessity table 14AQ), prompts corresponding to the scene information are input to the generation model (large-scale language model unit 13P) (S24Q in Figure 28, process 2 in Figure 25). [Third Process] The generator receives a situational description (driving situation description data 15BP) related to the superimposed image generated by the generation model in response to the input prompt. [Fourth Processing] At least one image from the captured image and superimposed image is associated with the situation description text and stored in the database (driving situation description image data 15CP).
[0085] Furthermore, the present invention is not limited to the embodiments described above, and it goes without saying that various other applications and modifications can be taken as long as they do not depart from the gist of the invention as described in the claims. For example, the embodiments described above describe in detail and specifically the configuration of the image explanation systems 100P and 100Q in order to explain the present invention in an easy-to-understand manner, and are not necessarily limited to those that include all the components described. Also, it is possible to replace a part of the configuration of one embodiment with a component of another embodiment. It is also possible to add a component of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, replace, or delete other components for a part of the configuration of each embodiment. For example, superimposed images may be created by other image processing methods. In the examples, the reference for positional relationships was described as a driver, but this is illustrative, and the reference for positional relationships can be changed depending on the use and purpose of the explanatory text. Furthermore, reflecting tacit knowledge from a particular field in the operation of the generating AI is within the scope of disclosure in this specification. Examples 1 and 2 are examples relating to the automotive field, but the application to the automotive field is just one example, and the present invention is broadly applicable to other fields as well.
[0086] Furthermore, some or all of the above configurations, functions, and processing units may be implemented in hardware, for example, by designing them as integrated circuits. Broadly defined processor devices such as FPGAs (Field Programmable Gate Arrays) and ASICs (Application Specific Integrated Circuits) may be used as hardware. Furthermore, each component of the image explanation system 100P and 100Q according to the above-described embodiment may be implemented on any hardware, as long as the respective hardware can send and receive information from each other via a network. Also, the processing performed by a certain processing unit may be implemented by a single piece of hardware, or by distributed processing by multiple pieces of hardware. [Explanation of symbols]
[0087] 1P Image Processing Device (Description Generation System) 1Q Vehicle Image Analysis System (Descriptive Text Generation System) 8. Communication lines 11P Driving Log Database 12P Description Generation Unit 13P Large-Scale Language Model Section 15P Description Database 16P Search Section 17P input / output section 11AP Vehicle Driving Video Data 11BP Camera Mounting Position Data (Camera mounting position information) 11CP Vehicle Driving Control Data (Control Information for Moving Objects) 11DP Image Description Data Model 12AP Vehicle position information superimposition unit (processing unit) 12BP Image Description Information Generation Unit (Processing Unit) 12CP Driving Status Description Generation Unit (Processing Unit) 15AP Image Description Data 15BP Driving Condition Description Data 15CP Driving Condition Description Image Data 91 Vehicles (mobile objects) 99 In-car camera (camera) 100P, 100Q Image Description System 111 Vehicle driving image data (captured images) 114 Superimposed image data of the vehicle's position (superimposed image) 1141 Vehicle position (reference for positional relationship)
Claims
1. The system is characterized by having a processing unit that outputs a superimposed image in which a reference for the positional relationship used by the generative model to explain the captured image is superimposed on the captured image taken by the camera. Image processing device.
2. The processing unit is characterized by outputting the superimposed image based on the camera mounting position information or the control information of the mobile body on which the camera is mounted. The image processing apparatus according to claim 1.
3. The camera mounting position information is characterized by including information indicating the horizontal deviation of the mounting position, or information indicating the horizontal angular deviation of the mounting position. The image processing apparatus according to claim 2.
4. The moving body is a vehicle, and the control information for the moving body includes information indicating the vehicle's speed or information indicating the vehicle's steering angle. The image processing apparatus according to claim 2.
5. The image description data model is defined by the summary information of the captured image or the surrounding information of the camera. The processing unit is characterized by instructing the generation model to generate an explanatory text for the superimposed image based on the image explanatory data model. The image processing apparatus according to claim 4.
6. The surrounding information of the camera is characterized by including any of the following: road information, surrounding object information, and surrounding environment information. The image processing apparatus according to claim 5.
7. The aforementioned road information is characterized by including information on the shape of the road. The image processing apparatus according to claim 6.
8. The surrounding object information is characterized by including any of the following: information on the type of surrounding vehicle, information on the color of the surrounding vehicle, information on the behavior of the surrounding vehicle, information on the location of surrounding pedestrians, information on the clothing of surrounding pedestrians, and information on the behavior of surrounding pedestrians. The image processing apparatus according to claim 6.
9. The aforementioned processing unit, The system accepts the search query entered by the user. Among the descriptive sentences generated by the generation model based on the superimposed image, a descriptive sentence similar to the search sentence, The superimposed image or the captured image corresponding to the similar explanatory text, The search results are characterized by displaying related information, including peripheral information of the camera corresponding to the similar descriptive text, as search results. The image processing apparatus according to claim 1.
10. A descriptive text generation system comprising an image processing device as described in claim 1 and a database searchable by an operator using natural language, The aforementioned processing unit, The generation model receives scene information indicating the scene recognized from the captured image or the superimposed image, Based on whether or not an explanation of the objects set for each scene is necessary, prompts corresponding to the scene information are input to the generated model. The generation model receives a situational description regarding the superimposed image generated in response to the prompt, The database is characterized by associating at least one image from the captured image and the superimposed image with the situational description. Description generation system.