Image processing device and explanatory text generation system

The image processing device uses positional relationship standards and a large-scale language model to generate descriptive text that aligns with human driver intent, enhancing search accuracy by accurately depicting surrounding objects based on vehicle position and driving conditions.

WO2026110425A1PCT designated stage Publication Date: 2026-05-28HITACHI LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/027713
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-20
Filing Date
2025-08-05
Publication Date
2026-05-28

Smart Images

  • Figure JP2025027713_28052026_PF_FP_ABST
    Figure JP2025027713_28052026_PF_FP_ABST
Patent Text Reader

Abstract

This image processing device (1P) has a host vehicle position information superimposition unit (12AP) that outputs host vehicle position superimposition image data (114) in which a positional relationship reference for a large language model unit (13P) to perform explanation is superimposed on vehicle travel image data (111) captured by a vehicle-mounted camera (99). The host vehicle position information superimposition unit (12AP) outputs host vehicle position superimposition image data (114) on the basis of camera attachment position data (11BP) or vehicle travel control data (11CP). The camera attachment position data (11BP) includes information indicating a horizontal deviation of an attachment position or information indicating a horizontal angle deviation of the attachment position.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device and description text generation system

[0001] The present invention relates to an image processing device and a description text generation system.

[0002] Conventionally, a technique called image caption generation for recognizing images and videos and generating description texts for those images has been disclosed (for example, Patent Document 1). In the technique described in Patent Document 1, by recognizing peripheral objects moving into the image of an in-vehicle camera and outputting text including the positional relationship between the vehicle and the peripheral objects, it is provided as a peripheral situation description text.

[0003] Japanese Patent Application Laid-Open No. 2019-214320

[0004] However, in conventional techniques such as Patent Document 1, it has not been particularly considered to generate a natural peripheral situation description text as described by a human, which can express the presence or absence of peripheral objects to be focused on according to the traffic scene where the vehicle is placed, rather than a text listing the peripheral objects in the image. In particular, when using a vehicle image database that searches and displays vehicle images collected from a connected car, it is desirable that a description text similar to the search query input by a human is given.

[0005] When generating a description text based on an image acquired from an in-vehicle camera, a person with driving experience can explain the positional relationship with surrounding vehicles, etc., based on the analogy of the position and traveling direction of their own vehicle in the image. On the other hand, in the case of a large language model for image understanding, there is a problem that a description text different from the traffic situation described by a human is given because the description is made based on the position of surrounding vehicles in the image. That is, for example, when an operator creates a description text, the operator creates the description text based on the tacit knowledge that "the description text must be in the positional relationship based on the position of the driver". Such a text can be called a driver-based description text. However, since the generative AI does not necessarily have such tacit knowledge, the description text by the generative AI does not necessarily become the driver-based description text described above.

[0006] This invention was made in consideration of these circumstances, and one of its objectives is for the AI ​​to generate explanatory text in a positional relationship intended by the operator. For example, in the field of automobiles, the objective is to provide an image processing device that generates natural-sounding explanatory text about the surrounding situation, which can express the presence or absence of surrounding objects of interest according to the traffic scene in which the vehicle is placed.

[0007] One feature of this invention is, for example, the assignment of positional relationship standards to an image.

[0008] According to the present invention, for example, it becomes possible to generate descriptive text in a positional relationship intended by the operator, thereby improving the operator's search accuracy. For example, in the field of automobiles, when generating descriptive text for a vehicle image, image description data including the presence or absence of surrounding objects of interest based on the vehicle's position information within the image is generated, and driving condition description data is generated based on the image description data, thereby improving the search accuracy from search terms entered by a human.

[0009] This is a diagram of the image description system according to Example 1. This is a diagram of the image processing device according to Example 1. This is a detailed diagram of the image processing device according to Example 1. This is a diagram showing an example of vehicle driving image data extracted from vehicle driving video data according to Example 1. This is a flowchart of the explanatory text generation unit according to Example 1. This is a diagram showing an example of camera mounting position data according to Example 1. This is a diagram showing an example of vehicle driving control data according to Example 1. This is a flowchart of the process executed by the vehicle position information superposition unit according to Example 1. This is a diagram showing the procedure for correcting vehicle position information based on camera mounting position data according to Example 1. This is a diagram showing the procedure for correcting vehicle position information based on vehicle driving control data according to Example 1. This is a diagram showing vehicle position superimposed image data in which the corrected vehicle position based on camera mounting position data and vehicle driving control data according to Example 1 is superimposed. This is a diagram showing the details of the image description data model that defines explanatory information for an image according to Example 1. This is a flowchart showing the details of the process executed by the image description information generation unit according to Example 1. This is a diagram showing image description data in which the large-scale language model unit explains the vehicle position superimposed image data based on the image description data model according to Example 1. This is a flowchart detailing the process of generating driving condition description data executed by the driving condition description generation unit for Example 1. This is a diagram showing an example of driving condition description image data for Example 1. This is a flowchart of the search unit for Example 1. This is a diagram showing the image search interface of the input / output unit for Example 1. This is a display screen diagram showing details of the search results in Figure 18 for Example 1. This is a display screen diagram showing details of the search results in Figure 18 for Example 1. This is a plan view of the vehicle for Example 1. This is a configuration diagram of the image description system for Example 2. This is a diagram showing details of the driving log for Example 2. This is a diagram showing an example of still image data which is part of the vehicle driving video data for Example 2. This is a configuration diagram of the vehicle image analysis device for Example 2. This is a configuration diagram of the necessity table for Example 2. This is a configuration diagram of the recognition table for Example 2. This is a flowchart of the description generation unit for Example 2. This is a hardware configuration diagram of the vehicle image analysis device for Example 2.This is a detailed flowchart of the image description generation process for Example 2. This flowchart shows an example of the specific operation of the image description generation process described in Figure 30 for Example 2. This is a detailed flowchart of the image description generation process for general roads for Example 2. This is a table showing intermediate data resulting from the image description generation process performed on the still image data of Figure 24 for Example 2. This is a table showing output data resulting from the description generation unit deleting unnecessary data from the intermediate data of Figure 33 for Example 2. This is a diagram detailing the GPS description generation process for Example 2. This is a diagram detailing the control description generation process for Example 2. This is a table showing an example of the description generated in Figures 35 and 36 for Example 2. This is a table showing instructions to the large-scale language model unit used in the driving condition description generation process for Example 2. This is a table showing an example generated from an image different from the image description in Figure 34 for Example 2. This is a flowchart of the search unit's processing for Example 2. This is a diagram showing the image search interface of the input / output unit for Example 2. This is the playback screen when the search result (scene 1) in Figure 41 for Example 2 is clicked. This is the playback screen that appears when the search result (scene 2) in Figure 41 related to Example 2 is clicked.

[0010] Examples 1 and 2 will be described below with reference to the drawings.

[0011] Figure 1 is a diagram of the configuration of the image description system 100. The image description system 100 consists of an image processing device 1P and connected cars 91, 92, and 93, which are connected cars, connected via a communication line 8. Vehicles 91 to 93 capture images of their surroundings using an on-board camera 99. This on-board camera 99 may capture the forward view (indicated by an arrow in Figure 1), as illustrated in Figure 1, or it may capture any view.

[0012] Vehicles 91 to 93 then transmit the following various data (vehicle driving log data) acquired while driving or stopped via the communication line 8: • Vehicle driving video data 11AP is video data captured by the on-board camera 99. • Camera mounting position data 11BP (details in Figure 6) indicates the mounting position data of the on-board camera 99. • Vehicle driving control data 11CP (details in Figure 7) is vehicle control data such as speed, acceleration / deceleration, and steering angle. The communication line 8 collects the various data transmitted from vehicles 91, 92, and 93 via the communication network and transmits it to the image processing device 1P. The image processing device 1P collects the various data received via the communication line 8 and stores it in the driving log DB 11P.

[0013] Figure 2 is a diagram of the configuration of the image processing device 1P. The driving log DB 11P stores vehicle driving video data 11AP, camera mounting position data 11BP, vehicle driving control data 11CP, and image description data model 11DP (see Figure 3 for details). The description text generation unit 12P analyzes the driving status of each vehicle using the large-scale language model unit 13P based on the vehicle driving video data 11AP, camera mounting position data 11BP, and vehicle driving control data 11CP, and generates natural language text that describes the driving status of the vehicles. Therefore, the description text generation unit 12P has a vehicle position information superimposition unit 12AP, an image description information generation unit 12BP, and a driving status description text generation unit 12CP (see Figure 3 for details).

[0014] The large-scale language model unit 13P is a generating AI called by the descriptive text generation unit 12P. The large-scale language model unit 13P generates image description data 15AP in text form for images and generates driving situation description data 15BP which integrates multiple sentences and pieces of information. The descriptive text DB 15P stores the image description data 15AP, the driving situation description data 15BP, and the driving situation description image data 15CP (see Figure 3 for details). The search unit 16P performs a search of the descriptive text DB 15P according to instructions from the input / output unit 17P and returns the results to the input / output unit 17P. The input / output unit 17P is operated by the user, receives a search statement from the user, issues a query instruction to the search unit 16P, and provides a response to the user based on the search results from the search unit 16P.

[0015] Figure 3 is a detailed configuration diagram of the image processing device 1P. The driving log DB 11P stores vehicle driving video data 11AP, camera mounting position data 11BP, vehicle driving control data 11CP, and image explanation data model 11DP. The vehicle position information superimposition unit 12AP generates vehicle position superimposed image data 114 (a superimposed image on which the large-scale language model unit 13P superimposes positional relationship references for explanation) by superimposing vehicle position information onto vehicle driving image data 111 (images captured by the onboard camera 99) extracted from the vehicle driving video data 11AP.

[0016] The description DB 15P stores image description data 15AP, driving situation description data 15BP, and driving situation description image data 15CP. The image description information generation unit 12BP generates image description data 15AP, which is description information for the vehicle position superimposed image data 114, based on the image description data model 11DP (see Figure 14 for details). Therefore, the image description data model 11DP defines the description information for the vehicle driving image data 111 (objects to be described in the driving situation description) (see Figure 12 for details).

[0017] The driving situation description generation unit 12CP generates driving situation description data 15BP, which is a driving situation description for the vehicle position superimposed image data 114, based on the image description data 15AP (see Figure 16 for details). Furthermore, the driving situation description generation unit 12CP generates driving situation description image data 15CP by adding the driving situation description data 15BP to the vehicle position superimposed image data 114 (see Figure 16 for details), and stores the generation result in the description DB 15P.

[0018] Figure 4 shows an example of vehicle driving image data 111 extracted from vehicle driving video data 11AP. This is still image data extracted from video footage from an onboard camera 99 of a vehicle driving on a highway. The vehicle driving image data 111 can be extracted, for example, at regular time intervals such as every 10 seconds, or by extracting 10 images at equal time intervals from a single video file.

[0019] Figure 5 is a flowchart of the explanatory text generation unit 12P. Processing starts from S21. In S22, the vehicle position information superimposition unit 12AP generates vehicle position superimposed image data 114 by superimposing vehicle position information onto vehicle driving image data 111 (see Figure 8 for details). In S23, the image explanatory information generation unit 12BP generates image explanatory data 15AP of the vehicle position superimposed image data 114 based on the image explanatory data model 11DP (see Figure 13 for details). In S24, the driving situation explanatory text generation unit 12CP generates driving situation explanatory text data 15BP based on the image explanatory data 15AP (see Figure 15 for details). Processing is completed in S25. As a result, the image processing device 1P generates image explanatory data 15AP that explains surrounding objects to be noticed and confirmed according to the vehicle's driving situation, based on the vehicle position information in the vehicle driving image data 111, which is an image from the in-vehicle camera 99. As a result, the image processing device 1P can generate natural driving situation description data 15BP according to the traffic scene, improving the search accuracy for user-submitted search queries.

[0020] Figure 6 shows an example of camera mounting position data 11BP. The camera mounting position data 11BP stores the horizontal deviation Δ of the mounting position of the in-vehicle camera 99 as -10% and the horizontal angular deviation θ as -5 degrees.

[0021] Figure 7 shows an example of vehicle driving control data 11CP. The vehicle driving control data 11CP stores the vehicle's speed v as 65 [km] and the vehicle's steering angle φ as 3 [deg].

[0022] Figure 21 is a plan view of vehicle 91. An onboard camera 99 is mounted on the front of vehicle 91, with the front wheels 97 slightly to the right and the rear wheels 98 facing the direction of travel. Camera mounting position data 11BP shows the horizontal deviation Δ and the horizontal angular deviation θ of the onboard camera 99. In a typical vehicle 91, a rearview mirror is installed on the central axis of the vehicle 91, so the drive recorder is often installed in a position that does not obstruct the view of the occupants, resulting in a horizontal deviation Δ. The horizontal deviation Δ shows the horizontal deviation of the mounting position of the onboard camera 99 from the central axis of the vehicle. Also, when installing the drive recorder, a horizontal angular deviation θ occurs due to the curvature of the windshield and misalignment during installation. Vehicle driving control data 11CP shows the vehicle's speed v and the steering angle φ of the front wheels 97. The vehicle's speed v shows the speed of vehicle 91 in the direction of travel. The steering angle φ indicates the angle between the steering wheel and the direction of travel of the vehicle 91.

[0023] Figure 8 is a detailed flowchart of the process (S22) executed by the vehicle position information superposition unit 12AP. From S221, the vehicle position information superposition unit 12AP starts processing. In S222, the vehicle position information superposition unit 12AP acquires vehicle driving image data 111. In S223, the vehicle position information superposition unit 12AP corrects the vehicle position information based on the camera mounting position data 11BP (see Figure 9 for details). In S224, the vehicle position information superposition unit 12AP corrects the vehicle position information based on the vehicle driving control data 11CP (see Figure 10 for details). In S225, the vehicle position information superposition unit 12AP generates vehicle position superimposed image data 114 (see Figure 11 for details). In S226, the vehicle position information superposition unit 12AP completes processing.

[0024] Figure 9 shows the procedure for correcting the vehicle's position information based on camera mounting position data 11BP. In the vehicle driving image data 112, the initial vehicle position 1121 is the predicted path (vehicle's driving line) of the initial vehicle position in the vehicle driving image data 111, and is defined as a line extending from bottom to top from the center point in the left-right direction on the camera image. Then, as a horizontal correction 1122, a correction equivalent to the horizontal deviation Δ is performed, resulting in the vehicle position after horizontal correction 1123. Subsequently, as a horizontal angle correction 1124, a correction equivalent to the horizontal angle deviation θ is performed, resulting in the vehicle position after horizontal angle correction 1125.

[0025] Figure 10 shows the procedure for correcting the vehicle's position information based on vehicle driving control data 11CP. In the vehicle driving image data 113, the vehicle position after camera mounting position correction 1131 is the vehicle's position after the camera mounting position has been corrected, and in Embodiment 1, it is the same as the vehicle position after horizontal angle correction 1125 in Figure 9. In contrast, the vehicle position after speed correction 1132, which indicates the driving position after a certain period of time in the direction of travel of 1131, is determined based on the speed v. Subsequently, the vehicle position after steering angle correction 1133 is determined based on the steering angle φ.

[0026] Figure 11 shows superimposed vehicle position image data 114, which is obtained by superimposing the vehicle's position 1141, which has been corrected based on camera mounting position data 11BP and vehicle driving control data 11CP. The vehicle's position 1141 in Figure 11 is the same as the vehicle's position 1133 after steering angle correction in Figure 10. On the other hand, the vehicle's position 1141 only needs to be a position based on at least one of the camera mounting position data 11BP and the vehicle driving control data 11CP. The vehicle position information superimposition unit 12AP generates superimposed vehicle position image data 114 by superimposing the vehicle's position 1141 onto the vehicle driving image data 111. Note that the vehicle body itself is not captured in the field of view of the camera mounted on the vehicle. Therefore, the vehicle's position 1141 is the predicted path (vehicle's driving line) that the vehicle will move in the future from the time the superimposed vehicle position image data 114 was taken. Furthermore, the vehicle position 1141 is data that indicates the positional relationship reference for the vehicle driving image data 111 captured by the onboard camera 99, which the large-scale language model unit 13P uses to provide explanations.

[0027] As a result, the large-scale language model unit 13P can generate natural image description data 15AP and driving situation description data 15BP based on the vehicle position superimposed image data 114, according to the positional relationship between the vehicle and surrounding vehicles and pedestrians. Therefore, the search accuracy from the driving situation description data 15BP for user search queries is improved. For example, the vehicle position superimposed image data 114 shows the left lane 1142 and the right lane 1143 in which truck 1144 is driving, and truck 1145 is changing lanes from the left lane 1142 to the right lane 1143. Here, since the vehicle position 1141 is in the right lane 1143 in the vehicle position superimposed image data 114, the large-scale language model unit 13P can recognize that "truck 1145 is located ahead in the same lane as the vehicle."

[0028] Figure 12 shows the details of the image description data model 11DP, which defines descriptive information for an image. The image description data model 11DP defines the information to be described as driving condition description data 15BP and examples of its description, as follows: ・"Image summary" is natural language that shows a summary of the entire image. ・"Road information" is information that shows the characteristics of the road on which the vehicle is traveling, such as "road type" and "road shape". ・"Surrounding vehicles" is information that shows the characteristics of surrounding vehicles, such as "vehicle type," "location," "color," and "behavior". ・"Surrounding pedestrians" is information that shows the characteristics of surrounding pedestrians, such as "location," "clothing," and "behavior". ・"Environment" is information that shows the characteristics of the surrounding environment that change depending on the date and time, such as "time," "weather," and "road surface condition." This road information, surrounding object information (surrounding vehicles, surrounding pedestrians), and surrounding environment information (environment) are included in the image description data model 11DP as camera surrounding information.

[0029] The image description data model 11DP is prepared in advance by an administrator or other person and registered in the image processing device 1P. The image description data model 11DP is prepared for each applicable task (driving test, safety instruction, etc.), and the items used in that task (for example, objects to be avoided in a driving test) are registered in it.

[0030] Figure 13 is a flowchart detailing the process (S23) executed by the image description information generation unit 12BP. According to this flowchart, the image description information generation unit 12BP generates image description data 15AP based on the vehicle position superimposed image data 114. From S231, the image description information generation unit 12BP starts processing. In S232, the image description information generation unit 12BP reads the vehicle position superimposed image data 114. In S233, the image description information generation unit 12BP reads the image description data model 11DP. In S234, the image description information generation unit 12BP instructs the large-scale language model unit 13P to describe the vehicle position superimposed image data 114 based on the image description data model 11DP, and obtains image description data 15AP from the large-scale language model unit 13P. In S235, the image description information generation unit 12BP ends processing.

[0031] As a result, the image processing device 1P defines the information to be explained according to the use case of image search and the explanatory parts that experts focus on as an image explanation data model 11DP, and can generate image explanation data 15AP and driving condition explanation text data 15BP, thereby improving the search accuracy for user search queries.

[0032] Figure 14 shows the image description data 15AP generated by the large-scale language model unit 13P, which describes the vehicle position superimposed image data 114 based on the image description data model 11DP. The large-scale language model unit 13P recognizes the position of the small truck 1145 (#3-A-1 in Figure 14) as "forward" (#3-A-3 in Figure 14) as the result of recognizing various information described in the image description data model 11DP in Figure 12 from the vehicle position superimposed image data 114 in Figure 11. This recognition is possible because the truck 1145 is located on the extension of the vehicle position 1141 that is visible in the vehicle position superimposed image data 114. As a result, descriptions of "position" such as #3-A-2 and #3-B-2 in Figure 14 are described in terms of relative position to the vehicle position, improving the accuracy of the traffic situation description in the driving situation description data 15BP generated from the image description data 15AP. On the other hand, the large-scale language model unit 13P compares this to the case where it recognizes various information described in the image description data model 11DP in Figure 12 from the vehicle driving image data 111 in Figure 4. In this case, the position of the truck, which appears small, remains abstract information that cannot narrow down its relative position to the vehicle's position, such as "right side in the image".

[0033] Figure 15 is a flowchart detailing the process (S24) for generating driving situation description data 15BP, which is performed by the driving situation description generation unit 12CP. According to this flowchart, the driving situation description generation unit 12CP generates driving situation description data 15BP based on the image description data 15AP. From S241, the driving situation description generation unit 12CP starts processing. In S242, the driving situation description generation unit 12CP reads the image description data 15AP. In S243, the driving situation description generation unit 12CP instructs the large-scale language model unit 13P to generate driving situation description data 15BP based on the image description data 15AP, and obtains the driving situation description data 15BP from the large-scale language model unit 13P. In S244, the driving situation description generation unit 12CP ends processing.

[0034] As a result, the image processing device 1P can generate driving condition description data 15BP with different languages, character counts, and writing styles based on the image description data 15AP, without having to perform image understanding processing using a computationally expensive large-scale language model each time.

[0035] Figure 16 shows an example of driving situation description image data 15CP. Driving situation description image data 15CP is data obtained by superimposing information on surrounding vehicles obtained from image description data 15AP onto the vehicle position superimposed image data 114. The input / output unit 17P may display the driving situation description image data 15CP alone, or it may display the driving situation description image data 15CP and the driving situation description text data 15BP together. This makes it possible for the image processing device 1P to allow the user to understand the surrounding traffic objects in the image data described in the image description data 15AP more intuitively.

[0036] Figure 17 is a flowchart of the search unit 16P. Processing starts in S51. In S52, the input / output unit 17P receives a search statement entered by the user. In S53, the similarity between the search statement entered by the user in S52 and the driving situation description data 15BP generated in S24 is evaluated, and video information containing the description with a high similarity is obtained. For similarity evaluation, cosine similarity search based on document vectorization may be used. In S54, vehicle driving video data 11AP (vehicle driving image data 111 or self-position superimposed image data 114 generated from it), driving situation description data 15BP, related information, etc. related to the video information searched in S53 are searched, and a response statement including the search results is generated. Related information includes, for example, information on surrounding vehicles from the self-position and information on surrounding pedestrians. In S55, the response statement is sent to the input / output unit 17P. Processing is completed in S56. As a result, the image processing device 1P can quickly present the user with the desired video by having it search for video data containing driving condition description data 15BP that is similar to the search query entered by the user in natural language.

[0037] Figure 18 shows the image search interface 61 of the input / output unit 17P. The image search interface 61 consists of a search text input unit 62 and a search result display unit 63. The search text input unit 62 receives the search text received from the user via the input / output unit 17P in S52. The search result display unit 63 displays multiple search results 631, 632, 633, and 634 as vehicle driving video data 11AP (video thumbnail images) in S54.

[0038] Figure 19 is a display screen diagram showing details of the search results 631 in Figure 18. The search results 631 also display the following information: - Display field 631a of the vehicle driving video data 11AP that was found to have a high similarity to the search query. In Figure 19, as an example of the vehicle driving video data 11AP, the vehicle position superimposed image data 114, in which the vehicle position 631d is superimposed, is displayed. - Display field 631b of the driving situation description data 15BP associated with the video in display field 631a. - Display field 631c of related information (image, time, weather, location, vehicle type, driving speed information) to the content of display field 631b.

[0039] Figure 20 is a display screen diagram showing details of the search results 632 in Figure 18. The search results 632 also display the following information: - Display field 632a of the vehicle driving video data 11AP that was found to have a high similarity to the search query. In Figure 20, as an example of the vehicle driving video data 11AP, vehicle driving image data 111 before the vehicle's position was superimposed is displayed. - Display field 632b of the driving situation description data 15BP associated with the video in display field 632a. - Display field 632c of related information (image, time, weather, location, vehicle type, driving speed information) for the content of display field 632b.

[0040] As shown in Figures 19 and 20, the image processing device 1P prompts the user to input a search query in natural language into the search query input unit 62. The image processing device 1P searches for videos by referring to driving condition description data 15BP, vehicle driving video data 11AP, and related information that are similar to the input search query, and displays the search results on the search result display unit 63. This makes it possible to improve search efficiency.

[0041] In the first embodiment described above, the vehicle's position and driving trajectory were superimposed on the image from the in-vehicle camera 99, and the positional relationship with surrounding vehicles, pedestrians, etc., was explained based on the superimposed image of the vehicle's position.

[0042] Next, we will describe Example 2. For example, when constructing an image database used in vehicle development, it is desirable to pre-associate a descriptive text with each image in the database, so that a descriptive text in the database that is similar to the search query entered by the user is assigned. Example 2 was made with these circumstances in mind, and describes an image description system 100Q (Figure 22) that databases images with naturalistic descriptive texts that match search queries entered by humans. With this image description system 100Q, it is possible to database images with naturalistic descriptive texts that match search queries entered by humans.

[0043] Figure 22 is a diagram of the configuration of the image description system 100Q. In the image description system 100Q, there are vehicles 91-93, each a connected car, that can communicate with the vehicle image analysis device 1Q via the communication line 8. Each vehicle 91-93 transmits various measurement data measured by itself to the vehicle image analysis device 1Q's driving log DB 11Q via the communication line 8. Each vehicle 91-93 is equipped with an on-board camera that takes images. The vehicle image analysis device 1Q generates a descriptive text of the situation of the images taken by the on-board camera. The communication line 8 can be a general public network, whether wired or wireless, such as the 5th generation mobile communication system, or 5G (5th Generation), which enables "massive simultaneous connections" and "ultra-low latency". Furthermore, by taking advantage of the features of new mobile phone systems after 5G, effects such as online (real-time while driving) descriptive text generation can also be expected.

[0044] Figure 23 is a diagram showing details of the driving log DB 11Q. The driving log DB 11Q stores measurement data from each vehicle 91 - 93 according to the type of data as follows. - Vehicle driving video data 11AQ, which is video data captured by an in-vehicle camera (not shown) during vehicle driving or parking. - Vehicle driving GPS data 11BQ, which is in-vehicle GPS (Global Positioning System) data. - Vehicle driving control data 11CQ, which is control data of a vehicle such as speed, acceleration, and steering angle, as vehicle driving log data obtained from an in-vehicle ECUQ (Electronic Control Unit) or the like. Note that GPS is an example of a satellite positioning system.

[0045] Figure 24 is a diagram showing an example of still image data 111Q, which is a part of the vehicle driving video data 11AQ. The still image data 111Q is an example of data extracted from the vehicle driving video data 11AQ captured by each vehicle 91 - 93 traveling on a highway.

[0046] Figure 25 is a configuration diagram of the vehicle image analysis device 1Q. In addition to the driving log DB 11Q described in Figure 23, the vehicle image analysis device 1Q includes an explanatory text generation unit 12Q, a large language model unit 13Q, an explanation target setting unit 14Q, an explanatory text DB (image database) 15Q, a search unit 16Q, and an input / output unit 17Q. The explanation target setting unit 14Q includes a necessity table 14AQ and a recognition table 14BQ. The large language model unit 13Q includes a VQA (Visual Question Answer) unit 13AQ and a summary generation unit 13BQ. Note that various data (the driving log DB 11Q, the necessity table 14AQ, and the recognition table 14BQ) stored in the vehicle image analysis device 1Q may be stored in a storage device (not shown) outside the vehicle image analysis device 1Q and may be configured to be accessible from the vehicle image analysis device 1Q via a network.

[0047] Based on the data in the driving log DB 11Q, the description generation unit 12Q analyzes the driving status of each vehicle using the large language model unit 13Q and the description target setting unit 14Q, and generates a natural sentence that describes the driving status of the vehicle according to the analysis result. Therefore, the description generation unit 12Q analyzes whether the driving status of the vehicle corresponds to any of the traffic scenes that have been pre-classified. The large language model unit 13Q is called by the description generation unit 12Q. The large language model unit 13Q is realized by LAVIS (LAnguage VISion) etc. which uses the natural language conversation method of the exchange of questions and answers as an interface, and has the following processing units. - The VQA unit 13AQ answers a natural language query about an image with a natural language. Therefore, the VQA unit 13AQ prepares an image recognition model such as a CNN (Convolutional Neural Network) in advance using learning data etc., and inputs a query sentence to the image recognition model to obtain a description sentence of the corresponding image. - The summary generation unit 13BQ answers a summary sentence (integrating multiple sentences) of the content of the input natural sentence (prompt) as a traffic situation description sentence. The summary generation unit 13BQ may use existing services such as GPT-4, CLIP, BLIP, BLIP-2 etc. as a service of a text generation AI that performs summary processing, translation processing etc.

[0048] The description target setting unit 14Q refers to the following tables, sets an object to be described corresponding to the traffic scene for the description generation unit 12Q, and stores it in the description text DB 15Q. - The necessity table 14AQ (Fig. 26) associates an object to be described for each traffic scene, and defines the degree of importance of whether to mention it in the description text for an individual object to be described. - The recognition table 14BQ (Fig. 27) defines the detailed content to be mentioned in the description text for an individual object to be described. The description text DB 15Q stores the driving status description text generated by the description generation unit 12Q.

[0049] In this way, the description generation unit 12Q performs the following (process 1) to (process 3). (Process 1) It identifies the scenes in the image received from the camera and reads recognition necessity information for each object in the image from the necessity table 14AQ for each identified scene. (Process 2) It recognizes the objects that are designated as requiring recognition based on the read recognition necessity information from within the image and generates a description for each object from the recognition result. (Process 3) It generates a situation description for the image based on the identified scenes and the description for each object, and stores the situation description for the image in the description DB 15Q, associating it with the image.

[0050] Figure 26 is a diagram of the configuration of the necessity table 14AQ. The explanation target setting unit 14Q sets the explanation targets for each traffic scene, such as highways, general roads, and parking lots. For example, the combination of the traffic scene "highway" and the explanation target "pedestrian" is "necessary / necessary". The recognition necessity information "necessary" on the left side of the "necessary / necessary" notation indicates that it is necessary to recognize the explanation target from the image. The explanation necessity information "necessary" on the right side of the "necessary / necessary" notation indicates that it is necessary to explain the explanation target, whether it was recognized or not, in an explanatory text. Therefore, in the traffic scene "general road", pedestrian recognition is "necessary", and even if it is not recognized, an explanation is set to be "necessary". On the other hand, in the traffic scene "general road", pedestrian crossing recognition is "necessary", but an explanation is not necessary.

[0051] In this way, the explanatory text generation unit 12Q reads information on whether an object in the image needs to be explained from the necessity table 14AQ. If an object designated as needing an explanation based on the read explanation necessity information cannot be recognized in the image, it generates an explanatory text for that object stating that the object does not exist in the image. This allows the system to generate explanatory text only for objects that are usually not present in a given traffic scene where a vehicle is driving, and to generate explanatory text even when objects that are usually present are not present. As a result, more natural traffic situation explanatory texts can be generated, which improves the search accuracy for user search queries.

[0052] Figure 27 is a diagram of the recognition table 14BQ. The recognition table 14BQ defines as detailed items items to be analyzed using the VQA unit 13AQ for the recognized object to be described, and items to be included in the description text based on the analysis results. For example, if a pedestrian is recognized, the VQA unit 13AQ analyzes their location, clothing color, and actions. In this way, the description text generation unit 12Q reads the detailed items about the object in the image from the recognition table 14BQ and generates a description text for each object based on the detailed items read from the image for the recognized object. This makes it possible to individually set the information to be added depending on the object to be described, and a natural traffic situation description text can be generated. As a result, the search accuracy for user search queries is improved.

[0053] Figure 28 is a flowchart of the description generation unit 12Q. The description generation unit 12Q generates an image description by having the VQA unit 13AQ perform image analysis on the still image data 111Q extracted from the vehicle driving video data 11AQ as part of the image description generation process (S21Q). The extraction process of the still image data 111Q is, for example, to extract images taken at regular intervals such as every 10 seconds, or to extract 10 images at equal time intervals from a single video file.

[0054] The description generation unit 12Q generates a GPS description based on the vehicle driving GPS data 11BQ as part of the GPS description generation process (S22Q). In this process, the target GPS data (positioning data) is either GPS data from the same time as the still image data 111Q in S21Q, or the GPS data with the closest time. In other words, the description generation unit 12Q adds a description to the image status description that relates to at least one of the following pieces of information: information about the time of day when the image was taken, and information about the driving position when the image was taken, based on the GPS data read from the vehicle's onboard GPS.

[0055] The explanatory text generation unit 12Q generates a control explanatory text based on the vehicle driving control data 11CQ as part of the control explanatory text generation process (S23Q). In this process, the control data to be used is the control data from the same time as, or the closest time to, the still image data 111Q in S21Q. In other words, the explanatory text generation unit 12Q adds an explanatory text to the image situation explanatory text, based on the vehicle driving control data read from the vehicle's onboard ECU (Electronic Control Unit) of the vehicle 91, relating to at least one of the following pieces of information: speed information, acceleration / deceleration information, and steering angle information.

[0056] The description generation unit 12Q generates a driving situation description by having the summary generation unit 13BQ summarize the image description generated in S21Q, the GPS description generated in S22Q, and the control description generated in S23Q, as part of the driving situation description generation process (S24Q). The description generation unit 12Q associates the generated driving situation description with the still image data 111Q that is the subject of the driving situation description and saves it in the description database 15Q. In other words, the summary generation unit 13BQ generates a summary according to the input prompt. The description generation unit 12Q then generates an image situation description by inputting a prompt to the summary generation unit 13BQ that includes information to be included in the situation description, information to be excluded from the situation description, and a prompt that includes a generation instruction for the situation description based on the information of example situation descriptions, as well as a description for each object.

[0057] Figure 29 is a hardware configuration diagram of the vehicle image analysis device 1Q. The vehicle image analysis device 1Q is configured as a computer 900 having a CPU 901, RAM 902, ROM 903, HDD 904, communication I / F 905, input / output I / F 906, and media I / F 907. The communication I / F 905 is connected to an external communication device 915. The input / output I / F 906 is connected to an input / output device 916. The media I / F 907 reads and writes data to the recording medium 917. Furthermore, the CPU 901 controls each processing unit by executing a program (also called an application or app) loaded into the RAM 902. This program can also be distributed via a communication line or recorded on a recording medium 917 such as a CD-ROM and distributed.

[0058] Figure 30 is a detailed flowchart of the image caption generation process (S21Q in Figure 28). The caption generation unit 12Q classifies the traffic scenes depicted in the still image data 111Q by querying the VQA unit 13AQ (S211Q). The VQA unit 13AQ receives still image data 111Q and a traffic scene inquiry (for example, "Where is this scene? For example, is it a road, highway, or parking lot?") and responds to the explanation unit 12Q with an answer (for example, highway) (line A01 in Figure 33). Note that the explanation unit 12Q reads the traffic scene inquiry that has been set in the vehicle image analysis device 1Q in advance by the administrator, so users of the vehicle image analysis device 1Q do not need to generate the traffic scene inquiry themselves. Subsequently, the explanation unit 12Q executes the loop processing from S212Q to S217Q by sequentially selecting the objects to be explained (pedestrians, bicycles, cars, etc.) that are registered in the necessity table 14AQ. Hereafter, the objects to be explained selected in this loop processing will be referred to as the selected objects.

[0059] The descriptive text generation unit 12Q determines whether recognition of the selected object is necessary in the traffic scene identified in S221Q by referring to the necessity table 14AQ (S212Q). If the answer in S212Q is Yes (necessary), the process proceeds to S213Q; otherwise, it proceeds to S217Q. The descriptive text generation unit 12Q queries the VQA unit 13AQ of the large-scale language model unit 13Q to determine whether the selected object exists in the still image data 111Q. This query is, for example, "Are there pedestrians in this scene?" (Row A02 in Figure 33). Based on the response to this query (object recognition result), the descriptive text generation unit 12Q recognizes the object that was designated as needing recognition from within the image. As a result of this recognition, the descriptive text generation unit 12Q determines whether the selected object exists in the still image data 111Q (S213Q). If the answer in S213Q is Yes (exists), the process proceeds to S214Q; otherwise, it proceeds to S215Q.

[0060] The description generation unit 12Q obtains detailed items about the selected object present in the still image data 111Q (S214Q). Therefore, the description generation unit 12Q refers to the recognition table 14BQ to obtain the detailed items corresponding to the selected object. Then, the description generation unit 12Q queries the VQA unit 13AQ of the large-scale language model unit 13Q for each of the detailed items about the selected object in the still image data 111Q. For example, the description generation unit 12Q determines that the selected object is a car, and the detailed items corresponding to a car in the recognition table 14BQ include "color," so it generates the query "What color is the car?".

[0061] The description generation unit 12Q determines whether a description (mention) of the selected object is necessary by referring to the necessity table 14AQ (S215Q). If the answer in S215Q is Yes (necessary), proceed to S216Q; otherwise, proceed to S217Q. The description generation unit 12Q generates a description of the selected object by one of the following methods (S216Q): - If the answer in S213Q is Yes, generate a description of the selected object from the combination of the query statement for the detailed items of the selected object obtained in S214Q and its answer statement. - If the answer in S213Q is No, generate a description of the selected object stating that the selected object was not recognized in the still image data 111Q (e.g., row D07 in Figure 39). The description generation unit 12Q determines whether processing for all selected objects has been completed now that the processing for this selected object is finished (S217Q). If the answer in S217Q is Yes (completed), terminate the processing; otherwise, switch to the unprocessed selected object and return to S212Q. This enables the generation of natural-sounding traffic situation descriptions that include explanations of surrounding objects that should be noticed and checked depending on the traffic scene in which the vehicle is traveling, thereby improving the accuracy of searches based on user queries.

[0062] Figure 31 is a flowchart showing a specific example of the operation of the image caption generation process described in Figure 30. The caption generation unit 12Q branches into processing for each traffic scene as follows (S301Q) depending on the result of the process of classifying the traffic scene captured in the still image data 111Q (S211Q in Figure 30): - If the classification result is a general road, an image caption for general roads is generated (S302Q). - If the classification result is a highway, an image caption for highways is generated (S303Q). - If the classification result is a parking lot, an image caption for parking lots is generated (S304Q). - If the classification result is other, an image caption specified for other is generated (S305Q).

[0063] Figure 32 is a detailed flowchart of the process for generating image captions for public roads (S302Q in Figure 31). The caption generation unit 12Q generates a traffic scene question and answer (S211Q in Figure 30) and a question and answer about the presence or absence of traffic lights (S311Q). As the determination process in Figure 30 (S213Q), the caption generation unit 12Q asks whether there are pedestrians in the image (S312Q), and proceeds to S313Q if there are, and to S314Q if there are no pedestrians. As the question and answer for detailed items of the pedestrian (S214Q in Figure 30), the caption generation unit 12Q generates an answer indicating the presence of pedestrians, a question and answer about the pedestrian's location, a question and answer about the pedestrian's color, and a question and answer about the pedestrian's actions (S313Q). The caption generation unit 12Q generates an answer indicating that there are no pedestrians (S314Q).

[0064] The descriptive text generation unit 12Q performs a determination process (S213Q) as shown in Figure 30, asking whether there is a car in the image (S315Q). If there is a car, it proceeds to S316Q; otherwise, it proceeds to S317Q. The descriptive text generation unit 12Q generates a response to the questions about the details of the car (S214Q in Figure 30), including a response indicating the presence of the car, a question and answer regarding the car's location, a question and answer regarding the car's make and model, a question and answer regarding the car's color, and a question and answer regarding the car's operation (S316Q). The descriptive text generation unit 12Q performs a determination process (S213Q) as shown in Figure 30, asking whether there is a bicycle in the image (S317Q). If there is a bicycle, it proceeds to S318Q; otherwise, it proceeds to S319Q. The explanatory text generation unit 12Q generates, as answers to questions about the detailed items of the bicycle (S214Q in Figure 30), an answer indicating the presence of the bicycle, a question and answer regarding the bicycle's location, a question and answer regarding the bicycle's color, and a question and answer regarding the bicycle's operation (S318Q).

[0065] The descriptive text generation unit 12Q performs a determination process (S213Q) as shown in Figure 30, asking whether there is a pedestrian crossing in the image (S319Q). If there is, it proceeds to S320Q; otherwise, it terminates the process. The descriptive text generation unit 12Q generates a response indicating the presence of a pedestrian crossing (S320Q). The above describes the process of generating image descriptive text for general roads, but the process of generating image descriptive text for other traffic scenes (S303Q to S305Q) is similar.

[0066] Figure 33 is a table showing intermediate data resulting from the image description generation process (S21Q) performed on the still image data 111Q of Figure 24. In this table, each row consists of a question and its corresponding answer, which together constitute the image description. For example, row A01 is the result of the traffic scene classification process (S211Q). Rows A02 to A04 are the results of the process (S213Q) determining whether the selected object exists in the still image data 111Q. Rows A05 to A13 are the results of the process (S214Q) obtaining detailed information about the selected object.

[0067] Figure 34 is a table showing the output data after the explanatory text generation unit 12Q has removed unnecessary data from the intermediate data in Figure 33. The difference from Figure 33 is that the explanatory text for pedestrians (row A02), the explanatory text for bicycles (row A03), and the explanatory text for toll booths (row A13), which were deemed unnecessary when not recognized in the necessity table 14AQ of the explanatory target setting unit 14Q, have been removed in Figure 34.

[0068] Figure 35 shows the details of the GPS description generation process (S22Q). The description generation unit 12Q generates descriptions that classify the time of shooting of the still image data 111Q into morning, noon, evening, night, etc., according to the time of day of the GPS information obtained from the vehicle driving GPS data 11BQ (S221Q). The description generation unit 12Q generates descriptions that specify the shooting location of the still image data 111Q as a major arterial road or city name, according to the latitude and longitude of the GPS information obtained from the vehicle driving GPS data 11BQ (S222Q). As a result, traffic condition descriptions that include information about the time of day and location of the vehicle's driving can be generated. Therefore, the search accuracy for user search queries is improved.

[0069] Figure 36 shows the details of the control description generation process (S23Q). The description generation unit 12Q generates a description of the driving speed control from the vehicle driving control data 11CQ (S231Q). The description includes, for example, low speed for speeds below 20 km / h, medium speed for 20 km / h to 60 km / h, and high speed for speeds above 60 km / h. The description generation unit 12Q generates a description of the acceleration / deceleration control from the vehicle driving control data 11CQ (S232Q). The description includes, for example, acceleration state for acceleration above a certain level, deceleration state for deceleration above a certain level, and constant speed state otherwise. The description generation unit 12Q generates a description of the steering control from the vehicle driving control data 11CQ (S233Q). The description includes, for example, left turn state for steering angles above a certain level, right turn state, and straight-ahead state otherwise. This enables the generation of more natural traffic situation descriptions that include explanations of the vehicle's control state and behavior. Therefore, this has the effect of improving the accuracy of searches based on user search terms.

[0070] Figure 37 is a table showing an example of the descriptive text generated in Figures 35 and 36. In row B01, the descriptive text for the time period is generated at S221Q. In row B02, the descriptive text for the driving location is generated at S222Q. In row B03, the descriptive text for the driving speed is generated at S231Q. In row B04, the descriptive text for the acceleration / deceleration state is generated at S232Q. In row B05, the descriptive text for the steering state is generated at S233Q.

[0071] Figure 38 is a table showing the instructions for the large-scale language model unit 13Q used in the process of generating a driving situation description (S24Q). Row C01 contains the instruction to generate a traffic situation description. Row C02 contains the information to be included in the traffic situation description. Row C03 contains the information to be excluded from the traffic situation description. Row C04 contains an example of a traffic situation description.

[0072] Then, in the process of generating a driving status description (S24Q), the description generation unit 12Q generates a prompt by sequentially combining the following texts from (item 1) to (item 4). Note that at least one of (item 2) and (item 3) may be omitted. (item 1) Instruction to the large-scale language model unit 13Q (Figure 38) (item 2) GPS description (lines B01 and B02 in Figure 37) (item 3) Control description (lines B03, B04 and B05 in Figure 37) (item 4) Image description (Figure 34)

[0073] The description generation unit 12Q inputs the generated prompt to the large-scale language model unit 13Q to obtain a traffic situation description written in natural language. Below is an example of a traffic situation description generated by the large-scale language model unit 13Q from the prompt (corresponding to situation description 732Q in Figure 42 below). Traffic situation description = "It is noon, and you are driving straight on the Metropolitan Expressway Route 5 at high speed while slowing down. In this scene, there are solid orange lanes on the expressway. Furthermore, a white truck is driving forward on the road ahead of you." This improves the search accuracy for user searches by generating natural traffic situation descriptions that are tailored to the intended use and user preferences.

[0074] Figure 39 is a table showing an example generated from a different image than the one in Figure 34. The image caption in Figure 39 is based on an image taken from a vehicle traveling on a public road. Therefore, the combination of a public road and a pedestrian in the necessity table 14AQ is "necessary" for explanation, and the information "No pedestrian was detected" is recorded in row D07. The traffic situation description generated by the large-scale language model unit 13Q through the driving situation description generation process (S24Q) from the prompt containing the image caption in Figure 39 is as follows: Traffic situation description = "You are currently driving at a slow, steady speed on a city road in Mito City at night and are preparing to turn left. There are no pedestrians in this scene."

[0075] The above describes the process of creating a database of driving condition descriptions (the process up to saving them to the description DB 15Q) with reference to Figure 39. The following describes examples of how the database information can be used. Figure 40 is a flowchart of the process of the search unit 16Q. The search unit 16Q searches for images that match the input search statement from the images stored in the description DB 15Q. Specifically, the search unit 16Q outputs images with a high degree of similarity between the input search statement and the situation descriptions stored in the description DB 15Q as search results. The search unit 16Q may, for example, list the images in descending order of similarity and output the top 1 to X images (relative similarity judgment), or it may output search results images whose similarity is higher than a pre-set threshold value (threshold Y) (absolute similarity judgment). The details of the process of the search unit 16Q will be explained below.

[0076] The search unit 16Q receives a search query entered by the user from the input / output unit 17Q (S61Q). In this case, the user is, for example, a commentator at a traffic control center that manages highways, and the search query is, for example, a request to collect images from the database of situations similar to an accident that occurred at a specific location on a highway at a specific time. The commentator plans to edit the image materials obtained from the database to produce a news program about the accident. The search unit 16Q evaluates the similarity between the user's search query and the descriptions stored in the description database 15Q (S62Q), and obtains video information (image information) associated with the descriptions with high similarity from the description database 15Q. For similarity evaluation, cosine similarity search based on document vectorization may be used.

[0077] The search unit 16Q retrieves not only the driving video and driving situation description acquired in S62Q, but also related information about the driving situation (such as weather information that is not in the description DB 15Q but can be obtained from the weather database by specifying the location and date of the driving situation description), and generates a response based on these results (S63Q). In other words, the search unit 16Q may also output related information acquired from a database other than the description DB 15Q as a search result, based on the information contained in the situation description corresponding to the image with a high degree of similarity. The search unit 16Q transmits the response from S63Q to the input / output unit 17Q (S64Q). This allows the system to quickly find the target video when searching for video data with a traffic situation description similar to the natural language search query entered by the user.

[0078] Figure 41 shows the image search interface 71Q of the input / output unit 17Q. The image search interface 71Q consists of a search input unit 72Q, which is the input field for the search text of S61Q, and a search result display unit 73Q, which is the display field for the answer text of S64Q. Multiple search results (scenes 1 to 4) are displayed as icons or thumbnail images in the search result display unit 73Q.

[0079] Figure 42 shows the playback screen when the search result (Scene 1) in Figure 41 is clicked. The following information is displayed on this playback screen from top to bottom: • Image 731Q of the search results. • The situation description 732Q of image 731Q was extracted due to its high similarity to the search query. • Additional information 733Q such as GPS description (time, location), control description (vehicle type, driving speed), and weather.

[0080] Figure 43 shows the playback screen when the search result (Scene 2) in Figure 41 is clicked. Similar to Figure 42, this playback screen displays the image 741Q of the search result, its situation description 742Q, and additional information 743Q, just as in Figure 42. This allows the user to improve search efficiency by entering a search query in natural language and then searching for videos by referring to traffic situation descriptions, videos, and related information that are similar to the search query.

[0081] According to the embodiment 2 described above, when generating a description of a vehicle image, the description generation unit 12Q refers to the necessity table 14AQ and generates a natural description that includes the presence or absence of surrounding objects of interest based on the traffic scene in which the vehicle is placed. As a result, a natural description is created in the database that mentions necessary surrounding objects while omitting unnecessary ones, thereby improving the accuracy of database searches from search queries entered by humans.

[0082] The contents of Example 2 and Example 1 can be combined or partially substituted, as illustrated below. ・In the image description generation process by the description generation unit 12Q of Example 2 (S21Q in Figure 28), the still image data 111Q extracted from the vehicle driving video data 11AQ is replaced with the vehicle position superimposed image data 114, which is obtained by superimposing the vehicle position information onto the vehicle driving image data 111 of Example 1. ・Each process of S22Q to S24Q in Figure 28 of Example 2 (such as the process of causing the large-scale language model unit 13Q in S24Q to generate a traffic situation description) also processes the vehicle position superimposed image data 114 instead of the still image data 111Q. In other words, the vehicle position superimposed image data 114 generated by the image processing device 1P executing S21 and S22 in Figure 5 may be considered as the still image data 111Q for the vehicle image analysis device 1Q, and the vehicle image analysis device 1Q may execute S21Q to S24Q in Figure 28.

[0083] Furthermore, the hardware configuration of the image processing device 1P in Example 1 is the same as the hardware configuration of the vehicle image analysis device 1Q in Example 2 shown in Figure 29. Moreover, the image processing device 1P in Example 1 and the vehicle image analysis device 1Q in Example 2 may be configured as the same device housed in the same enclosure. For example, the large-scale language model unit 13P may have the same functions as the large-scale language model unit 13Q. Also, each processing unit in the image processing device 1P in Example 1 and each processing unit in the vehicle image analysis device 1Q in Example 2 may operate on a single computer 900 (Figure 29), or they may be distributed and operated across multiple computers 900.

[0084] Furthermore, by combining the contents of Example 2 and Example 1, the following descriptive text generation system can be constructed. The descriptive text generation system comprises a processing unit (descriptive text generation unit 12P) of the image processing device 1P and a database (descriptive text DB 15P, descriptive text DB 15Q) that can be searched by the operator using natural language. The descriptive text generation unit 12P performs the following processes: [First process] Receives scene information indicating a scene recognized by the generation model (large-scale language model unit 13Q) from an captured image (vehicle driving image data 111) or superimposed image (vehicle position superimposed image data 114) (S211Q in Figure 30, process 1 in Figure 25). [Second process] Inputs a prompt corresponding to the scene information to the generation model (large-scale language model unit 13P) based on whether or not an object needs to be described (necessity table 14AQ) set for each scene (S24Q in Figure 28, process 2 in Figure 25). [Third process] Receives a situational descriptive text (driving situation descriptive text data 15BP) related to the superimposed image generated by the generation model in response to the input prompt. [Fourth Processing] At least one image from the captured image and superimposed image is associated with the situation description text and stored in the database (driving situation description image data 15CP).

[0085] Furthermore, the present invention is not limited to the embodiments described above, and it goes without saying that various other applications and modifications can be taken as long as they do not depart from the gist of the present invention as described in the claims. For example, the embodiments described above describe in detail and specifically the configuration of the image explanation systems 100P and 100Q in order to explain the present invention in an easy-to-understand manner, and are not necessarily limited to having all the described components. Also, it is possible to replace a part of the configuration of one embodiment with a component of another embodiment. It is also possible to add a component of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, replace, or delete other components for a part of the configuration of each embodiment. For example, superimposed images may be created by other image processing methods. In the examples, the reference for positional relationships was described as the driver, but this is an example, and the reference for positional relationships can be changed according to the use and purpose of the explanatory text. Also, reflecting tacit knowledge of a certain field in the operation of the generating AI is within the scope of disclosure of this specification. Examples 1 and 2 are embodiments relating to the automotive field, but the application to the automotive field is just one example, and the present invention is broadly applicable to other fields as well.

[0086] Furthermore, some or all of the above configurations, functions, and processing units may be implemented in hardware, for example, by designing them as integrated circuits. Broadly defined processor devices such as FPGAs (Field Programmable Gate Arrays) and ASICs (Application Specific Integrated Circuits) may be used as hardware. In addition, each component of the image explanation systems 100P and 100Q according to the above-described embodiments may be implemented on any hardware, as long as the respective hardware can send and receive information from each other via a network. Moreover, the processing performed by a certain processing unit may be implemented by a single piece of hardware, or by distributed processing by multiple pieces of hardware.

[0087] 1P Image processing device (explanatory text generation system) 1Q Vehicle image analysis device (explanatory text generation system) 8 Communication line 11P Driving log DB 12P Explanatory text generation unit 13P Large-scale language model unit 15P Explanatory text DB (database) 16P Search unit 17P Input / output unit 11AP Vehicle driving video data 11BP Camera mounting position data (camera mounting position information) 11CP Vehicle driving control data (control information for moving object) 11DP Image explanation data model 12AP Self-vehicle position information superimposition unit (processing unit) 12BP Image explanation information generation unit (processing unit) 12CP Driving situation explanation text generation unit (processing unit) 15AP Image explanation data 15BP Driving situation explanation text data 15CP Driving situation explanation image data 91 Vehicle (moving object) 99 Onboard camera (camera) 100P, 100Q Image explanation system 111 Vehicle driving image data (captured image) 114 Superimposed image data of vehicle position (superimposed image) 1141 Vehicle position (reference for positional relationship)

Claims

1. An image processing device characterized by having a processing unit that outputs a superimposed image in which a reference for the positional relationship used by a generative model to explain the captured image taken by a camera is superimposed.

2. The image processing apparatus according to claim 1, characterized in that the processing unit outputs the superimposed image based on camera mounting position information or control information of a mobile body on which the camera is mounted.

3. The image processing apparatus according to claim 2, characterized in that the camera mounting position information includes information indicating a horizontal deviation of the mounting position, or information indicating a horizontal angular deviation of the mounting position.

4. The image processing apparatus according to claim 2, characterized in that the moving body is a vehicle, and the control information of the moving body includes information indicating the speed of the vehicle or information indicating the steering angle of the vehicle.

5. The image processing apparatus according to claim 4, wherein the image description data model is defined by summary information of the captured image or peripheral information of the camera, and the processing unit instructs the generation model to generate a description of the superimposed image based on the image description data model.

6. The image processing apparatus according to claim 5, characterized in that the surrounding information of the camera includes any of road information, surrounding object information, and surrounding environment information.

7. The image processing apparatus according to claim 6, characterized in that the road information includes information on the shape of the road.

8. The image processing apparatus according to claim 6, characterized in that the surrounding object information includes any of the following: information on the type of surrounding vehicle, information on the color of the surrounding vehicle, information on the behavior of the surrounding vehicle, information on the location of surrounding pedestrians, information on the clothing of surrounding pedestrians, and information on the behavior of surrounding pedestrians.

9. The image processing apparatus according to claim 1, wherein the processing unit receives a search statement entered by a user, and displays as a search result the following: an explanatory statement similar to the search statement among the explanatory statements generated by the generation model based on the superimposed image, the superimposed image or the captured image corresponding to the similar explanatory statement, and related information including peripheral information of the camera corresponding to the similar explanatory statement.

10. A description generation system comprising an image processing device as described in claim 1 and a database searchable by an operator using natural language, wherein the processing device receives scene information indicating a scene recognized by the generation model from the captured image or the superimposed image, inputs a prompt to the generation model corresponding to the scene information based on whether or not an explanation of an object set for each scene is necessary, receives a situation description related to the superimposed image generated by the generation model in response to the prompt, and stores at least one image from the captured image and the superimposed image in association with the situation description in the database.

Citation Information

Patent Citations

  • Periphery monitoring system

    JP2018056953A

  • Signal processing system, evaluation system therefor, and signal processing device used in signal processing system

    JP2019158390A

  • Recognition processing device, vehicle control device, recognition control method and program

    JP2019214320A

  • Obstacle identification apparatus and obstacle identification program

    JP2021064156A