Description text generation system and method for producing a description text database

The database system generates natural-sounding descriptions for vehicle images based on traffic scenes, addressing the limitations of conventional technologies by associating descriptive text with images, enhancing search accuracy through scene-specific object explanations.

JP7846656B2Active Publication Date: 2026-04-15HITACHI LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-26
Publication Date
2026-04-15

AI Technical Summary

Technical Problem

Conventional image caption generation technologies do not provide a suitable system for creating a database of vehicle images collected from connected cars and searching for images that match natural language instruction conditions input by users, failing to generate natural-sounding descriptions that accurately reflect the surrounding environment based on traffic scenes.

Method used

A database system that uses a generative model for image analysis and natural language output, generating descriptive text based on scene information, and associating it with images, including a processing unit that stores text and a necessity table to determine which objects in the image require explanation, improving the accuracy of user searches.

Benefits of technology

Enables the creation of a database with natural-sounding descriptions that match user search terms, enhancing the accuracy of searches by including or omitting explanations of relevant objects based on traffic scenarios, thereby improving search efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007846656000001
    Figure 0007846656000001
  • Figure 0007846656000002
    Figure 0007846656000002
  • Figure 0007846656000003
    Figure 0007846656000003
Patent Text Reader

Abstract

To make a data-base of images associated with a natural descriptive text such that it matches a retrieval sentence entered by a person.SOLUTION: A descriptive text generating part 12 of a vehicle image analyzer 1 identifies a scene captured in an image received from a camera, reads recognition necessity information on a target in the image for each of the identified scene from a necessity table 14A, identifies the target being specified as necessary to be recognized in the read recognition necessity information, generates a descriptive text for each target from the recognition result, generates a situation descriptive text of the image on the basis of the identified scene and the descriptive text for each target, and memories the image associated with the situation descriptive text of the image into a descriptive text DB 15.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a description generation system and a method for producing a database. Description

Background Art

[0002] Conventionally, a technology called image caption generation has been disclosed that recognizes images and videos and generates a description text of the image (for example, Patent Document 1). In the technology described in Patent Document 1, by recognizing peripheral objects moving into the image of an in-vehicle camera and outputting text including the positional relationship between the vehicle and the peripheral objects, it is provided as a peripheral situation description text.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, conventional technologies such as Patent Document 1 do not provide a system suitable for the purpose of creating a database of vehicle images collected from a connected car and searching for images that match the natural language instruction conditions input by the user from the database. When constructing such an image database, it is desirable to associate a description text of each image in the database in advance, and that the description texts in the database similar to the search sentence input by the user are assigned.

[0005] On the other hand, conventional technologies such as those described in Patent Document 1 did not particularly consider generating natural-sounding descriptions of the surrounding environment that are similar to the search terms entered by the user. For example, simply detecting objects in an image and using the names of those objects as the description may not mention the presence or absence of surrounding objects that should be of interest depending on the traffic scene in which the vehicle is located, and as a result, it may not be possible to match the search terms entered by the user.

[0006] This invention was made in consideration of these circumstances, and its main objective is to create a database of images associated with natural-sounding descriptions that match search terms entered by humans. [Means for solving the problem]

[0007] To solve the above problems, the explanatory text generation system of the present invention has the following features. This invention provides a database that can be searched by workers using natural language. ,before It has a processing unit that stores text in a database, The aforementioned processing unit, A generative model capable of image analysis and natural language output receives scene information representing the scene recognized from the camera image. Based on whether or not an explanation of the object is required for each scene, the scene information is as follows Generate an image description, A prompt containing the aforementioned image description is generated, The aforementioned A prompt is input to the generation model, and the generation model receives the situation description generated in response to the prompt. The system is characterized by storing the aforementioned situational description in the database in association with the aforementioned image. Other features will be described later. [Effects of the Invention]

[0008] According to the present invention, it is possible to create a database of images associated with natural-sounding descriptions that match search terms entered by humans. [Brief explanation of the drawing]

[0009] [Figure 1] It is a configuration diagram of an image description system according to this embodiment. [Figure 2] It is a diagram showing details of a travel log DB according to this embodiment. [Figure 3] It is a diagram showing an example of still image data that is part of vehicle travel video data according to this embodiment. [Figure 4] It is a configuration diagram of a vehicle image analysis device according to this embodiment. [Figure 5] It is a configuration diagram of a necessity table according to this embodiment. [Figure 6] It is a configuration diagram of a recognition table according to this embodiment. [Figure 7] It is a flowchart of an explanatory text generation unit according to this embodiment. [Figure 8] It is a hardware configuration diagram of a vehicle image analysis device according to this embodiment. [Figure 9] It is a detailed flowchart of the generation process of an image explanatory text according to this embodiment. [Figure 10] It is a flowchart showing an example of a specific operation of the generation process of the image explanatory text described in FIG. 9 according to this embodiment. [Figure 11] It is a detailed flowchart of the generation process of an image explanatory text for general roads according to this embodiment. [Figure 12] It is a table showing intermediate data of the result of executing the generation process of an image explanatory text for the still image data of FIG. 3 according to this embodiment. [Figure 13] It is a table showing output data of the result of deleting unnecessary data from the intermediate data of FIG. 12 by the explanatory text generation unit according to this embodiment. [Figure 14] It is a diagram showing details of the generation process of a GPS explanatory text according to this embodiment. [Figure 15] It is a diagram showing details of the generation process of a control explanatory text according to this embodiment. [Figure 16]A table showing an example of the description text generated in FIGS. 14 and 15 related to this embodiment. [Figure 17] A table showing an instruction statement to the large language model unit used in the generation process of the driving situation description text related to this embodiment. [Figure 18] A table showing an example generated from an image different from the image description text of FIG. 13 related to this embodiment. [Figure 19] A flowchart showing the processing of the search unit related to this embodiment. [Figure 20] A diagram showing the image search interface of the input / output unit related to this embodiment. [Figure 21] A playback screen when the search result (scene part 1) of FIG. 20 related to this embodiment is clicked. [Figure 22] A playback screen when the search result (scene part 2) of FIG. 20 related to this embodiment is clicked.

Mode for Carrying Out the Invention

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0011] FIG. 1 is a configuration diagram of an image description system 100. In the image description system 100, as vehicles (connected cars) that can communicate with the vehicle image analysis device 1 via the communication line 8, there are vehicles 91 - 93. Each of the vehicles 91 - 93 transmits various measurement data measured by the vehicle itself to the driving log DB 11 of the vehicle image analysis device 1 via the communication line 8. Each of the vehicles 91 - 93 is equipped with an in - vehicle camera that captures images. The vehicle image analysis device 1 generates a situation description text of the image captured by the in - vehicle camera. Furthermore, the communication line 8 can be a general public network, whether wired or wireless, such as the fifth-generation mobile communication system, or 5G (5th Generation), which enables "massive simultaneous connections" and "ultra-low latency." By taking advantage of the features of newer mobile phone systems beyond 5G, it is also possible to expect effects such as online (real-time while driving) generation of explanatory text.

[0012] Figure 2 shows the details of the driving log DB11. The driving log DB11 stores measurement data from each vehicle (91-93) according to the following data types. • Vehicle driving video data 11A, which is video data captured by an onboard camera (not shown) while the vehicle is in motion or stopped. • Vehicle GPS (Global Positioning System) data, specifically vehicle driving GPS data 11B. • Vehicle driving log data obtained from the in-vehicle ECU (Electronic Control Unit), etc., including vehicle driving control data 11C such as speed, acceleration / deceleration, and steering angle. GPS is one example of a satellite positioning system.

[0013] Figure 3 shows an example of still image data 111, which is part of the vehicle driving video data 11A. Still image data 111 is an example of data extracted from vehicle driving video data 11A, which is captured by each vehicle 91-93 traveling on the highway.

[0014] Figure 4 is a diagram showing the configuration of the vehicle image analysis device 1. The vehicle image analysis device 1 includes, in addition to the driving log DB 11 described in Figure 2, an explanatory text generation unit 12, a large-scale language model unit 13, an explanatory target setting unit 14, an explanatory text DB (image database) 15, a search unit 16, and an input / output unit 17. The explanatory target setting unit 14 includes a necessity table 14A and a recognition table 14B. The large-scale language model unit 13 includes a VQA (Visual Question Answer) unit 13A and a summary generation unit 13B. Furthermore, the various data stored within the vehicle image analysis device 1 (driving log DB11, necessity table 14A, and recognition table 14B) may be stored in a storage device (not shown) outside the vehicle image analysis device 1, and may be configured to be accessible from the vehicle image analysis device 1 via a network from that storage device.

[0015] The descriptive text generation unit 12 analyzes the driving conditions of each vehicle based on the data in the driving log DB 11, using the large-scale language model unit 13 and the description target setting unit 14, and generates natural language text that describes the driving conditions of the vehicles according to the analysis results. Therefore, the descriptive text generation unit 12 analyzes whether the driving conditions of the vehicles fall into one of the pre-classified traffic scenes. The large-scale language model unit 13 is called by the explanatory text generation unit 12. The large-scale language model unit 13 is implemented by LAVIS (LAnguage VISion), which uses a natural language conversation method that interacts between question sentences and answer sentences as an interface, and has the following processing units. The VQA unit 13A responds to natural language queries about images in natural language. Therefore, the VQA unit 13A prepares an image recognition model such as a CNN (Convolutional Neural Network) in advance using training data, and obtains a description of the corresponding image by inputting a query to that image recognition model. The summary generation unit 13B responds with a summary sentence, which is a description of the traffic situation, obtained by summarizing (integrating multiple sentences) the content of the input natural language text (prompt). The summary generation unit 13B may use existing services such as GPT-4, CLIP, BLIP, and BLIP-2 as text generation AI services that perform summarization and translation processing.

[0016] The explanation target setting unit 14 refers to the following table and sets the explanation target corresponding to the traffic scene in the explanation text generation unit 12, and stores it in the explanation text DB 15. • The necessity table 14A (Figure 5) associates explanatory objects with each traffic scene and defines the degree of importance of whether or not to mention each explanatory object in the explanatory text. • Recognition table 14B (Figure 6) defines the detailed information to be mentioned in the explanatory text for each individual object. The description DB15 stores the driving condition descriptions generated by the description generation unit 12.

[0017] Thus, the explanatory text generation unit 12 executes the following (process 1) to (process 3). (Process 1) The system identifies the scenes in the image received from the camera, and for each identified scene, reads recognition requirements for the objects in that image from the requirements table 14A. (Process 2) Based on the recognition necessity information read, the system recognizes objects that require recognition from within the image and generates a descriptive text for each object from the recognition results. (Process 3) Based on the identified scene and the descriptive text for each object, a situational description of the image is generated, and the situational description of the image is associated with the image and stored in the description DB15.

[0018] Figure 5 is a diagram showing the configuration of the necessity table 14A. The explanation target setting unit 14 sets the explanation targets for each traffic scene, such as expressways, general roads, and parking lots. For example, the combination of the traffic scene "highway" and the object to be explained "pedestrian" is marked "Required / Required". The recognition requirement information "Required" on the left side of this "Required / Required" notation indicates that the object to be explained needs to be recognized from the image. The explanation requirement information "Required" on the right side of the "Required / Required" notation indicates that the object to be explained, whether recognized or not recognized from the image, needs to be explained in the explanatory text. Therefore, in the "general road" traffic scenario, pedestrian recognition is "necessary," and even if pedestrians are not recognized, an explanation is "necessary." On the other hand, in the "general road" traffic scenario, pedestrian crossings are "necessary," but an explanation is not required.

[0019] In this manner, the explanatory text generation unit 12 reads information on whether an object in the image needs to be explained from the necessity table 14A. If an object designated as requiring an explanation based on the read explanation necessity information cannot be recognized from within the image, it generates an explanatory text for that object stating that the object does not exist in the image. This allows the system to generate descriptive text only for objects that are not normally present in a given traffic scenario, and to generate descriptive text even when objects that are normally present are not present. As a result, more natural traffic situation descriptions can be generated, improving the accuracy of user searches.

[0020] Figure 6 is a diagram showing the configuration of the recognition table 14B. The recognition table 14B defines, as detailed items, the items to be analyzed using the VQA unit 13A for the recognized explanatory object, and the items to be included in the explanatory text based on the analysis results. For example, if a pedestrian is recognized, the VQA unit 13A analyzes their location, clothing color, and actions. In this way, the descriptive text generation unit 12 reads detailed items about objects in the image from the recognition table 14B and generates descriptive texts for each object based on the detailed items read from the image. This makes it possible to individually set the information to be added depending on the object being described, and thus natural traffic condition descriptions can be generated. As a result, the search accuracy for user search queries is improved.

[0021] Figure 7 is a flowchart of the explanatory text generation unit 12. The description generation unit 12 generates an image description by having the VQA unit 13A perform image analysis on the still image data 111 extracted from the vehicle driving video data 11A as part of the image description generation process (S21). The extraction process for the still image data 111 is, for example, to extract images taken at regular intervals such as every 10 seconds, or to extract 10 images at equal time intervals from a single video file.

[0022] The description generation unit 12 generates a GPS description based on the vehicle driving GPS data 11B as part of the GPS description generation process (S22). In this process, the target GPS data (positioning data) is either GPS data from the same time as the still image data 111 in S21, or the GPS data with the closest time. In other words, the explanatory text generation unit 12 adds an explanatory text to the image situation description that relates to at least one of the following pieces of information: information about the time of day when the image was taken, and information about the driving position when the image was taken, based on GPS data read from the vehicle's onboard GPS.

[0023] The explanatory text generation unit 12 generates a control explanatory text based on the vehicle driving control data 11C as part of the control explanatory text generation process (S23). In this process, the control data to be used is the control data from the same time as the still image data 111 in S21, or the control data from the closest time. In other words, the descriptive text generation unit 12 adds a descriptive text to the image's situation description, based on vehicle driving control data read from the vehicle's onboard ECU (Electronic Control Unit) of the vehicle 91, relating to at least one of the following pieces of information: speed information, acceleration / deceleration information, and steering angle information.

[0024] The description generation unit 12 generates a driving situation description by having the summarization generation unit 13B summarize the image description generated in S21, the GPS description generated in S22, and the control description generated in S23, as part of the driving situation description generation process (S24). The description generation unit 12 associates the generated driving situation description with the still image data 111 that the driving situation description describes and saves it in the description database 15. In other words, the summary generation unit 13B generates a summary text that conforms to the input prompt. The descriptive text generation unit 12 then inputs a prompt to the summary generation unit 13B that includes information to be included in the situational descriptive text, information to be excluded from the situational descriptive text, and a prompt that includes instructions for generating a situational descriptive text based on example sentences of situational descriptive texts, as well as descriptive texts for each object, thereby generating a situational descriptive text for the image.

[0025] Figure 8 is a hardware configuration diagram of the vehicle image analysis device 1. The vehicle image analysis device 1 is configured as a computer 900 having a CPU 901, RAM 902, ROM 903, HDD 904, communication I / F 905, input / output I / F 906, and media I / F 907. The communication interface 905 is connected to an external communication device 915. The input / output interface 906 is connected to the input / output device 916. The media interface 907 reads and writes data to the recording medium 917. Furthermore, the CPU 901 improves and controls each processing unit by executing a program (also called an application or app) loaded into the RAM 902. This program can also be distributed via a communication line or by recording it on a recording medium 917 such as a CD-ROM and distributing it that way.

[0026] Figure 9 is a detailed flowchart of the image caption generation process (S21 in Figure 7). The description generation unit 12 classifies the traffic scenes captured in the still image data 111 by querying the VQA unit 13A (S211). The VQA unit 13A receives the still image data 111 and a traffic scene query (for example, "Where is this scene? For example, is it a road, highway, or parking lot?") and responds to the description generation unit 12 with a reply (for example, highway) (line A01 in Figure 12). Note that the description generation unit 12 reads the traffic scene query that has been set in the vehicle image analysis device 1 in advance by the administrator, so users of the vehicle image analysis device 1 do not need to generate the traffic scene query themselves. The explanatory text generation unit 12 then executes the loop processing from S212 to S217 by sequentially selecting the items to be explained from the necessity table 14A (pedestrians, bicycles, cars, etc.). Hereafter, the items to be explained selected in this loop processing will be referred to as the selected items.

[0027] The explanatory text generation unit 12 determines whether recognition of the selected object is necessary in the traffic scene identified in S221 by referring to the necessity table 14A (S212). If the answer in S212 is Yes (necessary), the process proceeds to S213; otherwise, it proceeds to S217. The description generation unit 12 queries the VQA unit 13A of the large-scale language model unit 13 to determine whether the selected object exists in the still image data 111. This query is, for example, "Are there any pedestrians in this scene?" (row A02 in Figure 12). Based on the response to this query (the object recognition result), the description generation unit 12 recognizes the object that needs to be recognized from within the image. As a result of this recognition, the description generation unit 12 determines whether the selected object exists in the still image data 111 (S213). If the answer in S213 is Yes (it exists), the process proceeds to S214; otherwise, it proceeds to S215.

[0028] The description generation unit 12 obtains detailed items about the selected object present in the still image data 111 (S214). Therefore, the description generation unit 12 refers to the recognition table 14B to obtain the detailed items corresponding to the selected object. Then, the description generation unit 12 queries the VQA unit 13A of the large-scale language model unit 13 for each of the detailed items about the selected object in the still image data 111. For example, the description generation unit 12 determines that the selected object is a car, and the detailed items corresponding to a car in the recognition table 14B include "color," so it generates the query "What color is the car?".

[0029] The explanatory text generation unit 12 determines whether or not an explanation (mention) of the selected object is necessary by referring to the necessity table 14A (S215). If the answer in S215 is Yes (necessary), proceed to S216; otherwise, proceed to S217. The description generation unit 12 generates a description of the selected object by one of the following means (S216): If the answer to S213 is Yes, a description of the selected object is generated from the combination of the query statement and the answer statement for the detailed items of the selected object obtained in S214. If the answer to S213 is No, a descriptive text is generated indicating that the selected object was not recognized in the still image data 111 (e.g., line D07 in Figure 18). The explanatory text generation unit 12 determines whether processing for all selected objects has been completed after finishing processing for the current selected object (S217). If the answer in S217 is Yes (completed), the processing ends; otherwise, it switches to the unprocessed selected object and returns to S212. This enables the generation of natural-sounding traffic situation descriptions that include explanations of surrounding objects that should be noticed and checked depending on the traffic scene in which the vehicle is traveling, thereby improving the accuracy of searches based on user queries.

[0030] Figure 10 is a flowchart showing an example of the specific operation of the image caption generation process described in Figure 9. The explanatory text generation unit 12 branches into processing for each traffic scene as follows (S301), depending on the result of the processing to classify the traffic scenes captured in the still image data 111 (S211 in Figure 9). If the classification result is for a public road, generate an image description for public roads (S302). If the classification result is for a highway, generate an image description for highways (S303). If the classification result is for a parking lot, generate an image description for the parking lot (S304). • If the classification result is "Other," generate an image description specified for "Other" (S305).

[0031] Figure 11 is a detailed flowchart of the process for generating image captions for general roads (S302 in Figure 10). The explanatory text generation unit 12 generates questions and answers regarding traffic scenes (S211 in Figure 9) and questions and answers regarding the presence or absence of traffic signals (S311). The explanatory text generation unit 12 performs the determination process (S213) shown in Figure 9, asking whether there are pedestrians in the image (S312). If there are, it proceeds to S313; otherwise, it proceeds to S314. The explanatory text generation unit 12 generates, as answers to questions about the detailed items of the pedestrian (S214 in Figure 9), an answer indicating the presence of the pedestrian, a question and answer regarding the pedestrian's location, a question and answer regarding the pedestrian's color, and a question and answer regarding the pedestrian's actions (S313). The explanatory text generation unit 12 generates a response indicating that there are no pedestrians (S314).

[0032] The explanatory text generation unit 12 performs the determination process (S213) shown in Figure 9, asking whether there is a car in the image (S315). If there is, it proceeds to S316; otherwise, it proceeds to S317. The explanatory text generation unit 12 generates, as answers to questions about the details of the automobile (S214 in Figure 9), an answer to a question about the automobile's location, an answer to a question about the automobile's make and model, an answer to a question about the automobile's color, and an answer to a question about the automobile's operation (S316). The explanatory text generation unit 12 performs the determination process shown in Figure 9 (S213), which asks whether there is a bicycle in the image (S317). If there is, it proceeds to S318; otherwise, it proceeds to S319. The explanatory text generation unit 12 generates, as answers to questions about the detailed items of the bicycle (S214 in Figure 9), an answer indicating the presence of the bicycle, a question and answer regarding the bicycle's location, a question and answer regarding the bicycle's color, and a question and answer regarding the bicycle's operation (S318).

[0033] The explanatory text generation unit 12 performs the determination process shown in Figure 9 (S213), asking whether there is a pedestrian crossing in the image (S319). If there is, it proceeds to S320; otherwise, it terminates the process. The explanatory text generation unit 12 generates a response indicating the presence of a pedestrian crossing (S320). As explained above in Figure 10, the process for generating image captions for general roads is similar, but the process for generating image captions for other traffic scenes (S303-S305) is the same.

[0034] Figure 12 is a table showing the intermediate data obtained by performing the image description generation process (S21) on the still image data 111 in Figure 3. In this table, each row consists of a question and its corresponding answer, which together form an image description. For example, line A01 is the result of the process of classifying traffic scenes (S211). Lines A02 to A04 are the results of the process (S213) that determines whether or not the selected object exists in the still image data 111. Rows A05 to A13 are the results of the process (S214) for obtaining detailed information about the selected object.

[0035] Figure 13 is a table showing the output data after the explanatory text generation unit 12 has removed unnecessary data from the intermediate data in Figure 12. The difference from Figure 12 is that the explanatory text for pedestrians (row A02), the explanatory text for bicycles (row A03), and the explanatory text for toll booths (row A13), which were deemed unnecessary when not recognized in the necessity table 14A of the explanatory target setting unit 14, have been removed in Figure 13.

[0036] Figure 14 shows the details of the GPS description generation process (S22). The explanatory text generation unit 12 generates an explanatory text that classifies the time of shooting of the still image data 111 into morning, noon, evening, night, etc., according to the time of day of the GPS information obtained from the vehicle driving GPS data 11B (S221). The explanatory text generation unit 12 generates an explanation that identifies the location where the still image data 111 was taken, such as a major arterial road or a city name, according to the latitude and longitude of the GPS information obtained from the vehicle driving GPS data 11B (S222). This allows for the generation of traffic condition descriptions that include information about the time of day and location of vehicle travel. As a result, the accuracy of user searches is improved.

[0037] Figure 15 shows the details of the control explanation text generation process (S23). The explanatory text generation unit 12 generates explanatory text related to driving speed control from the vehicle driving control data 11C (S231). The explanatory text may include, for example, that speeds below 20 km / h are low speed, 20 km / h to 60 km / h are medium speed, and 60 km / h and above are high speed. The explanatory text generation unit 12 generates explanatory texts regarding acceleration and deceleration control from the vehicle driving control data 11C (S232). The explanatory texts include, for example, "acceleration state" when the acceleration is above a certain level, "deceleration state" when the deceleration is above a certain level, and "constant speed state" otherwise. The explanatory text generation unit 12 generates explanatory text related to steering control from the vehicle driving control data 11C (S233). The explanatory text may include, for example, left turn state, right turn state when the steering angle is above a certain level, and straight-ahead state in all other cases. This allows for the generation of more natural traffic situation descriptions that include explanations of the vehicle's control state and behavior. As a result, it has the effect of improving the accuracy of searches based on user queries.

[0038] Figure 16 is a table showing an example of the explanatory text generated in Figures 14 and 15. In line B01, the time zone description is generated in S221. In line B02, a description of the driving location is generated in S222. In line B03, the description of the driving speed is generated in S231. In line B04, a descriptive text about the acceleration and deceleration state is generated in S232. In line B05, a description of the steering state is generated at S233.

[0039] Figure 17 is a table showing the instructions given to the large-scale language model unit 13 used in the process of generating a driving condition description (S24). Line C01 contains the instructions for generating a traffic condition description. Line C02 contains the information to be included in the traffic condition description. Line C03 contains information to be excluded from the traffic condition description. Line C04 contains an example of a traffic situation description.

[0040] Then, in the process of generating the driving condition description (S24), the description generation unit 12 generates a prompt by sequentially combining the following texts from (item 1) to (item 4). Note that at least one of (item 2) and (item 3) may be omitted. (Item 1) Instructions to the large-scale language model unit 13 (Figure 17) (Item 2) GPS explanation (lines B01 and B02 in Figure 16) (Item 3) Control description (lines B03, B04, and B05 in Figure 16) (Item 4) Image caption (Figure 13)

[0041] The description generation unit 12 inputs the generated prompt to the large-scale language model unit 13 to obtain a traffic situation description written in natural language. Below is an example of a traffic situation description (corresponding to situation description 732 in Figure 21 below) generated by the large-scale language model unit 13 from the prompt. Traffic situation description: "It is noon, and you are driving straight on the Metropolitan Expressway Route 5 at high speed, slowing down considerably. In this scene, there are solid orange lane markings on the expressway. Additionally, a white truck is driving ahead of you on the road." This improves search accuracy for user searches by generating natural-sounding traffic condition descriptions tailored to usage and user preferences.

[0042] Figure 18 is a table showing an example generated from an image different from the image caption in Figure 13. The image caption for Figure 18 is for images taken from a vehicle traveling on a public road. Therefore, since the combination of public road and pedestrian in the necessity table 14A requires explanation, the information "Pedestrian not recognized" is included in row D07. The traffic situation description generated by the large-scale language model unit 13 through the driving situation description generation process (S24) from the prompt containing the image description in Figure 18 is as follows: Traffic situation description: "You are currently driving at a slow, steady speed on a city road in Mito City at night, preparing to turn left. There are no pedestrians in this scene."

[0043] The above explains the process of creating a database of driving condition descriptions (up to saving them to the description DB15), referring to Figure 18. The following describes examples of how the database information can be used. Figure 19 is a flowchart showing the processing of the search unit 16. The search unit 16 searches for images that match the input search query from the images stored in the description database 15. Specifically, the search unit 16 outputs images as search results that have a high degree of similarity between the input search query and the situational descriptions stored in the description database 15. The search unit 16 may, for example, list images in descending order of similarity in the search results and output the top 1 to X images (relative similarity determination), or it may output search results images whose similarity is higher than a pre-set threshold value (threshold Y) (absolute similarity determination). The following describes the details of the processing in the search unit 16.

[0044] The search unit 16 receives a search query entered by the user from the input / output unit 17 (S61). In this case, the user is, for example, a commentator at a traffic control center that manages highways, and the search query is, for example, a request to collect images from the database of situations similar to an accident that occurred at a specific location on a highway at a specific time. The commentator plans to edit the image materials obtained from the database to produce a news program about the accident that occurred. The search unit 16 evaluates the similarity between the user's search query and the descriptions stored in the description database 15 (S62), and retrieves video information (image information) associated with the descriptions with high similarity from the description database 15. For similarity evaluation, cosine similarity search based on document vectorization may be used.

[0045] The search unit 16 retrieves the driving video and driving situation description acquired in S62, as well as related information about the driving situation (such as weather information that is not in the description DB 15 but can be obtained from the weather database by specifying the location and date of the driving situation description), and generates a response based on these results (S63). In other words, the search unit 16 may output related information obtained from a database other than the description DB 15 as part of the search results, based on the information contained in the situational description corresponding to the image with a high degree of similarity. The search unit 16 sends the response text from S63 to the input / output unit 17 (S64). This allows users to quickly find the desired video when searching for video data containing traffic condition descriptions similar to the natural language search query they entered.

[0046] Figure 20 shows the image search interface 71 of the input / output unit 17. The image search interface 71 consists of a search input section 72, which is the input field for the search text in S61, and a search result display section 73, which is the display field for the answer text in S64. Multiple search results (scenes 1 to 4) are displayed as icons or thumbnail images in the search result display section 73.

[0047] Figure 21 shows the playback screen when the search result (Scene 1) in Figure 20 is clicked. The following information is displayed on this playback screen from top to bottom. • Image 731 from the search results. The situation description 732 for image 731 was extracted because it had a high degree of similarity to the search query. • Additional information such as GPS description (time, location), control description (vehicle type, driving speed), and weather (733 items).

[0048] Figure 22 shows the playback screen when the search result (Scene 2) in Figure 20 is clicked. Similar to Figure 21, this playback screen displays the image 741 of the search result, its situation description 742, and additional information 743, just as in Figure 21. This allows users to improve search efficiency by entering search terms in natural language and then searching for videos by referencing traffic condition descriptions, videos, and related information that are similar to their search terms.

[0049] According to the embodiment described above, when generating a description of a vehicle image, the description generation unit 12 refers to the necessity table 14A and generates a natural description that includes the presence or absence of surrounding objects of interest based on the traffic scene in which the vehicle is placed. As a result, a natural description is created in the database that mentions necessary surrounding objects while omitting unnecessary ones, thereby improving the accuracy of database searches from search queries entered by humans.

[0050] Furthermore, the present invention is not limited to the embodiments described above, and it goes without saying that various other applications and modifications can be taken as long as they do not depart from the gist of the present invention as described in the claims. For example, the embodiments described above are detailed and specific descriptions of the configuration of the vehicle image analysis device 1 in order to explain the present invention in an easy-to-understand manner, and are not necessarily limited to those comprising all the components described. Also, it is possible to replace a part of the configuration of one embodiment with a component of another embodiment. It is also possible to add a component of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, replace, or delete other components for a part of the configuration of each embodiment.

[0051] Furthermore, some or all of the above configurations, functions, and processing units may be implemented in hardware, for example, by designing them as integrated circuits. Broadly defined processor devices such as FPGAs (Field Programmable Gate Arrays) and ASICs (Application Specific Integrated Circuits) may be used as hardware. Furthermore, each component of the vehicle image analysis device 1 according to the above-described embodiment may be implemented on any hardware, as long as the respective hardware can send and receive information from each other via a network. Also, the processing performed by a certain processing unit may be implemented by a single piece of hardware, or by distributed processing by multiple pieces of hardware. [Explanation of symbols]

[0052] 1. Vehicle image analysis device 8. Communication lines 11. Driving Log Database 11A Vehicle driving video data 11B Vehicle GPS data 11C Vehicle Driving Control Data 12. Description Generation Unit 13. Large-scale language model section 13A VQA section 13B Summary generator 14. Section for setting the subject of explanation 14A Necessity / Dependency Table 14B Recognition Table 15 Description Database (Image Database) 16 Search Section 17 Input / output section 91-93 Vehicles 100 Image Description System 111 Still image data 732,742 Situation description

Claims

1. A database that workers can search using natural language, The system includes a processing unit for storing documents in the aforementioned database, The aforementioned processing unit, A generative model capable of image analysis and natural language output receives scene information representing the scene recognized from the camera image. Based on whether or not an explanation is required for each object set for each scene, an image description corresponding to the scene information is generated. A prompt containing the aforementioned image description is generated, The prompt is input to the generation model, The generation model receives the situation description generated in response to the prompt, A description generation system characterized by associating the aforementioned situation description with the aforementioned image and storing it in the aforementioned database.

2. The explanatory text generation system according to claim 1, The aforementioned generation model is a descriptive text generation system characterized by including a language model.

3. The explanatory text generation system according to claim 1, The prompt is generated based on whether an explanation is necessary and whether recognition is necessary, and is a description generation system.

4. The explanatory text generation system according to claim 1, The processing unit is characterized by generating the image description text by inputting a query statement to the generation model.

5. The explanatory text generation system according to claim 1, The prompt is characterized by including a control explanation.

6. The explanatory text generation system according to claim 1, The prompt is characterized by including a GPS description in the description generation system.

7. The explanatory text generation system according to claim 2, The prompt is generated based on whether an explanation is necessary and whether recognition is necessary, and is a description generation system.

8. The explanatory text generation system according to claim 7, The processing unit is characterized by generating the image description text by inputting a query statement to the generation model.

9. The explanatory text generation system according to claim 8, The prompt is characterized by including a control explanation.

10. The explanatory text generation system according to claim 9, The prompt is characterized by including a GPS description in the description generation system.

11. The processing unit is, A generative model capable of image analysis and natural language output receives scene information representing the scene recognized from the camera image. Based on whether or not an explanation is required for each object set for each scene, an image description corresponding to the scene information is generated. A prompt containing the aforementioned image description is generated, The prompt is input to the generation model, The generation model receives the situation description generated in response to the prompt, A method for producing a descriptive text database, characterized by associating the aforementioned situational descriptive text with the aforementioned image and storing it in a database.

12. A method for producing an explanatory text database according to claim 11, The aforementioned generative model is a method for producing an explanatory text database, characterized in that it includes a language model.

13. A method for producing an explanatory text database according to claim 11, A method for producing an explanatory text database, characterized in that the prompt is generated based on whether an explanation is necessary and whether recognition is necessary.

14. A method for producing an explanatory text database according to claim 11, The method for producing an image description database is characterized in that the processing unit generates the image description by inputting a query into the generation model.

15. A method for producing an explanatory text database according to claim 11, A method for producing an explanatory text database, characterized in that the prompt includes a control explanatory text.

16. A method for producing an explanatory text database according to claim 11, The method for producing a descriptive text database is characterized in that the prompt includes GPS descriptive text.

17. A method for producing an explanatory text database according to claim 12, A method for producing an explanatory text database, characterized in that the prompt is generated based on whether an explanation is necessary and whether recognition is necessary.

18. A method for producing an explanatory text database according to claim 17, The method for producing an image description database is characterized in that the processing unit generates the image description by inputting a query into the generation model.

19. A method for producing an explanatory text database according to claim 18, A method for producing an explanatory text database, characterized in that the prompt includes a control explanatory text.

20. A method for producing an explanatory text database according to claim 19, The method for producing a descriptive text database is characterized in that the prompt includes GPS descriptive text.

Citation Information

Patent Citations

  • Recognition processing device, vehicle control device, recognition control method and program

    JP2019214320A

  • Information processing device, program and information processing method

    JP2024158436A

  • Information processing device, information processing method and information processing program

    JP2025038912A

  • Program, information processing device, method and system

    JP2025050370A