Article generation system, article generation program, and report database production method

CN122804234APending Publication Date: 2026-09-22HITACHI LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480087452.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2024-12-17
Publication Date
2026-09-22

AI Technical Summary

Benefits of technology

[0018]根据本发明,能够提高基于图像生成报告书的作业的效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122804234A_ABST
    Figure CN122804234A_ABST
Patent Text Reader

Abstract

The present application provides an article generation system, an article generation program, and a report book database production method. An operation assistance system (1000) includes: a vehicle image analysis device (1) that generates a situation explanation text using generative AI (Artificial Intelligence) from an image captured by a camera and saves the situation explanation text in a database; a scene specification unit (1001) that specifies an explanation target scene as an explanation target; a scene search unit (1002) that searches for a situation explanation text associated with the explanation target scene specified by the scene specification unit (1001) from the situation explanation texts saved in the database; an abstract condition setting unit (1003) that saves a time and an explanation target object corresponding to the explanation target scene specified by the scene specification unit (1001); and a report book generation device (1004) that generates a report book of the explanation target object using a situation explanation text corresponding to the saved time and explanation target object from among the situation explanation texts searched by the scene search unit (1002).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an article generation system, an article generation program, and a method for producing a report database. Background Technology

[0002] Patent Document 1 can be cited as an example of a technique for extracting desired scenes from moving images. Patent Document 1 specifically discloses a method for generating explanatory text about sports video content. Patent Document 1 also specifically discloses a method for generating summary images with additional explanations for video regions of interest to the user.

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2005-109566 Summary of the Invention

[0006] The technical problem that the invention aims to solve

[0007] However, Patent Document 1 does not consider the development of autonomous vehicles.

[0008] For example, in the development of autonomous vehicles, there are situations where developers create reports describing the results of test drives. For instance, developers of autonomous vehicles analyze dynamic images captured by sensors such as cameras mounted on the vehicle to determine the scenarios that should be reported, and create reports so that the results of the test drives can be applied in subsequent development.

[0009] If an accident or other incident occurs during the test drive, a report detailing the situation before and after the incident is required. Additionally, a full-day test drive may require reports lasting several hours.

[0010] This invention was proposed in consideration of the following situation, and the main technical problem is to improve the efficiency of image-based report generation.

[0011] Technical means to solve the problem

[0012] To address the aforementioned problems, the article generation system of the present invention has the following features.

[0013] The present invention is characterized by comprising: an image analysis unit that generates condition descriptions using generative AI (Artificial Intelligence) based on images captured by a camera and stores them in a database; a scene designation unit that designates a description object scene as the object of description; a scene retrieval unit that retrieves condition descriptions associated with the description object scene designated by the scene designation unit from the condition descriptions stored in the database; a summary condition setting unit that stores the period and description object corresponding to the description object scene designated by the scene designation unit; and a report generation unit that uses the condition descriptions retrieved by the scene retrieval unit that correspond to the period and description object stored by the summary condition setting unit to generate a report about the description object using generative AI.

[0014] The article generation program of the present invention enables a computer to perform the following steps: generating a condition description text using generative AI based on images captured by a camera; storing the condition description text in a database; specifying a description object scene as the object of description; retrieving condition description texts associated with the specified description object scene from the condition description texts stored in the database; storing the period and description object corresponding to the specified description object scene; and using the retrieved condition description texts corresponding to the time and description object of the description object scene, generating a report about the description object using generative AI.

[0015] The method for producing a report database according to the present invention is characterized by comprising: a step of generating a condition description text using generative AI based on images captured by a camera; a step of storing the condition description text in a database; a step of a scene designation unit designating a description object scene as the subject of the description; a step of a scene retrieval unit retrieving condition description texts associated with the description object scene designated by the scene designation unit from the condition description texts stored in the database; a step of a summary condition setting unit storing the period and description object corresponding to the designated description object scene; a step of generating a report about the description object using generative AI based on the condition description texts retrieved by the scene retrieval unit that correspond to the period and description object stored by the summary condition setting unit; and a step of generating a database including the report texts.

[0016] Other features will be described later.

[0017] Invention Effects

[0018] According to the present invention, the efficiency of image-based report generation can be improved. Attached Figure Description

[0019] Figure 1 This is a structural diagram of the work assistance system of this embodiment.

[0020] Figure 2 This is a flowchart of the report generation process in this embodiment.

[0021] Figure 3 This is a diagram showing the details of the driving log database in this embodiment.

[0022] Figure 4 This is a diagram illustrating an example of still image data as part of the vehicle driving dynamic image data in this embodiment.

[0023] Figure 5 This is a structural diagram of the vehicle image analysis device according to this embodiment.

[0024] Figure 6 This is a structural diagram of the yes / no table in this embodiment.

[0025] Figure 7 This is a structural diagram of the identification table in this embodiment.

[0026] Figure 8 This is a flowchart of the explanatory text generation section of this embodiment.

[0027] Figure 9 This is a hardware structure diagram of the work assistance system in this embodiment.

[0028] Figure 10 This is a detailed flowchart of the image description generation process in this embodiment.

[0029] Figure 11 This refers to the implementation method. Figure 10 The flowchart illustrates an example of the specific actions involved in generating and processing image captions.

[0030] Figure 12 This is a detailed flowchart of the image description generation process for general roads in this embodiment.

[0031] Figure 13 This indicates that the implementation method is... Figure 4 The table contains intermediate data resulting from the processing of still image data to generate image descriptions.

[0032] Figure 14 This indicates that the description text generation unit of this embodiment is... Figure 13 The intermediate data is deleted, and the resulting output data is displayed in a table.

[0033] Figure 15 This is a diagram showing the details of the GPS description text generation process in this embodiment.

[0034] Figure 16 This is a diagram showing the details of the control specification document generation process in this embodiment.

[0035] Figure 17 This refers to the implementation method. Figure 15 and Figure 16 An example table of explanatory text generated in the middle.

[0036] Figure 18 This is a table representing the instruction text of the large-scale language model unit used in the traffic condition description text generation process of this embodiment.

[0037] Figure 19 This indicates the basis of this embodiment and Figure 14 The image description is an example of a table generated from different images.

[0038] Figure 20 This is a flowchart illustrating the processing of the retrieval unit in this embodiment.

[0039] Figure 21 This is a diagram showing the image retrieval interface of the input / output unit in this embodiment.

[0040] Figure 22 This is the implementation method Figure 21 The screen that appears when the search result (Scenario 1) is clicked.

[0041] Figure 23 This is the implementation method Figure 21 The screen that appears when the search result (Scenario 2) is clicked.

[0042] Figure 24 This is a flowchart representing the retrieval process.

[0043] Figure 25 It is a diagram illustrating the period of the scene.

[0044] Figure 26 It is a diagram illustrating a table that records the period of a scene.

[0045] Figure 27 It is a diagram illustrating the table defining the objects to be described.

[0046] Figure 28 It is a diagram illustrating the program's startup status.

[0047] Figure 29 It is a diagram illustrating the detection status of an object. Detailed Implementation

[0048] The following description uses the accompanying drawings to illustrate one embodiment of the present invention.

[0049] Figure 1 This is a structural diagram of the 1000 work assistance system.

[0050] The work assistance system 1000 includes a driving log database 11 (which serves as a database), a vehicle image analysis device 1, a scene designation unit 1001, a scene retrieval unit 1002, a summary condition setting unit 1003, and a report generation device 1004. This work assistance system 100 operates as a document generation system that assists in generating reports.

[0051] The vehicle image analysis device 1 includes a first generative AI (Artificial Intelligence) 1101. The vehicle image analysis device 1 uses the first generative AI 1101 to generate a traffic condition description 1006 based on images captured by a camera and stores it in [the relevant database / system]. Figure 4 The functions of the image analysis section in the explanatory text database 15.

[0052] The scene designation unit 1001 includes a second generative formula AI 1102, which is used to designate the description object scene.

[0053] Scene retrieval unit 1002 includes a second generative AI 1102, which retrieves data from stored... Figure 4 The traffic condition description 1006 associated with the description object scene specified by the scene designation unit 1001 is retrieved from the traffic condition description description database 15.

[0054] When specifying the preceding and following periods of a scene and the objects to be described (e.g., the other party in an accident, the location of an emergency stop), the summary condition setting unit 1003 saves the preceding and following periods and the objects to be described corresponding to the specified scene.

[0055] The report generation device 1004 includes a third generating AI 1103. The first generating AI 1101, the second generating AI 1102, and the third generating AI 1103 can be different generating AIs or a single generating AI. When the report generation device 1004 generates a report 1005, it appropriately generates a report database 1007 and saves the report 1005.

[0056] The first generative AI 1101 includes a model that learns using images to determine a scene from images captured by a camera. The second generative AI 1102 and the third generative AI 1103 include models that learn using articles.

[0057] The first generative AI1101, the second generative AI1102, or the third generative AI1103 can use existing services such as GPT-4, CLIP, BLIP, and BLIP-2.

[0058] Figure 2 This is a flowchart of the report generation process in this embodiment.

[0059] In this embodiment, in order to generate Figure 1 The report 1005 includes the following four steps. Details of each step are described below.

[0060] In step S11, the vehicle image analysis device 1 uses images captured by sensors such as cameras to generate a traffic condition description 1006 using the first generative formula AI1101.

[0061] In step S12, the scene retrieval unit 1002 retrieves the scene required to generate the report 1005 based on the traffic condition description text 1006 generated in step S11 using the second generative AI 1102.

[0062] In step S13, the abstract condition setting unit 1003 specifies the abstract period and the objects to be described in the abstract period.

[0063] In step S14, the report generation device 1004 generates a report 1005 using the third generative formula AI1103 based on the results of the retrieval process, the summary period, and the descriptive objects in the summary period. This improves the efficiency of image-based report generation.

[0064] In step S15, when the report generation device 1004 saves the driving log corresponding to the generated report 1005 in the report database 1007, Figure 2 The processing is now complete.

[0065] Step S11: Details of the explanatory text generation process

[0066] Figure 1 The vehicles 91-93 shown are so-called connected cars that can communicate with the work assistance system 1000 via communication line 8. Each vehicle 91-93 sends various measurement data measured by the vehicle to the driving log database 11 via communication line 8.

[0067] Each vehicle 91-93 is equipped with an onboard camera that captures images. The vehicle image analysis device 1 generates a traffic condition description based on the images captured by the onboard camera.

[0068] Furthermore, communication line 8 can utilize general public networks, whether wired or wireless, such as the fifth-generation mobile communication system, 5G, which enables "multiple simultaneous connections" and "ultra-low latency." Moreover, by applying the features of new mobile phone systems after 5G, it is expected that effects such as online (real-time while driving) text generation can be achieved.

[0069] Figure 3 This is a diagram showing the details of the driving log database 11.

[0070] The driving log database 11 stores measurement data from each vehicle 91 to 93 according to each data type, as described below.

[0071] Vehicle driving dynamic image data 11A is dynamic image data captured by an onboard camera (illustration omitted) while the vehicle is in motion or parked.

[0072] Vehicle GPS data 11B is vehicle-mounted GPS (Global Positioning System) data.

[0073] Vehicle driving control data 11C is vehicle driving log data obtained from on-board ECU (Electronic Control Unit), etc., which includes vehicle control data such as speed, acceleration and deceleration, and steering angle.

[0074] GPS is an example of a satellite positioning system.

[0075] Figure 4 This is a diagram showing an example of still image data 111, which is part of vehicle driving dynamic image data 11A.

[0076] Still image data 111 is an example of data extracted from dynamic vehicle image data 11A captured by vehicles 91 to 93 traveling on a highway.

[0077] Figure 5 This is a structural diagram of the vehicle image analysis device 1.

[0078] Vehicle image analysis device 1 has Figure 3 The description includes a driving log database 11, a descriptive text generation unit 12, a large-scale language model unit 13, a description object setting unit 14, a descriptive text database (image database) 15, a retrieval unit 16, and an input / output unit 17. The description object setting unit 14 has a yes / no table 14A and an identification table 14B. The large-scale language model unit 13 has a VQA (Visual Question Answering) unit 13A and a summary generation unit 13B.

[0079] In addition, various data stored in the vehicle image analysis device 1 (driving log database 11, optional table 14A, identification table 14B) can also be stored in a storage device (not shown) outside the vehicle image analysis device 1, so that the vehicle image analysis device 1 can access the storage device via a network.

[0080] The explanatory text generation unit 12 analyzes the driving conditions of each vehicle using the large-scale language model unit 13 and the description object setting unit 14, based on the data in the driving log database 11, and generates natural language articles describing the driving conditions of the vehicles based on the analysis results. Therefore, the explanatory text generation unit 12 analyzes whether the driving conditions of the vehicles correspond to a certain traffic scenario pre-classified.

[0081] The large-scale language model unit 13 is called by the explanatory text generation unit 12. The large-scale language model unit 13 is implemented through LAVIS (LAnguage VISion, visual language framework), which uses a natural language conversation mode for exchanging questions and answers, and has the following processing units.

[0082] • The VQA unit 13A responds to inquiries about images in natural language using natural language. Therefore, the VQA unit 13A prepares image recognition models such as CNNs (Convolutional Neural Networks) in advance using learning data, and obtains corresponding descriptions of the images by inputting the query text into the image recognition model.

[0083] The summary generation unit 13B will summarize the content of the input natural language article (prompt words) (by aggregating multiple articles) to obtain a summary text, which will then be used as a traffic condition description to provide an answer. The summary generation unit 13B can use existing services such as GPT-4, CLIP, BLIP, and BLIP-2 as text generation AI services for summarizing and translating.

[0084] The description object setting unit 14 refers to the following table, sets the description object corresponding to the traffic scene in the description text generation unit 12, and stores it in the description text database 15.

[0085] ·Do you need Form 14A ( Figure 6 It associates the objects of description with various traffic scenarios and defines whether the importance of individual objects of description needs to be mentioned in the explanatory text.

[0086] • Identification Table 14B ( Figure 7 For individual objects being described, the detailed information mentioned in the description is provided.

[0087] The explanatory text database 15 stores traffic condition explanatory texts generated by the explanatory text generation unit 12.

[0088] Thus, the explanatory text generation unit 12 performs the following processes (1) to (3).

[0089] (1) Determine the scene captured in the image received from the camera, and for each determined scene, read the information on whether the objects in the image need to be identified from the yes / no table 14A.

[0090] (2) Identify objects in the image that need to be identified based on the read information on whether identification is required, and generate a description of each object based on the identification results.

[0091] (3) Generate traffic condition descriptions for images based on the defined scene and descriptions for each object, and store the traffic condition descriptions for images in association with the images in the description database 15.

[0092] Figure 6 Is the structural diagram in Table 14A required? The object setting section 14 sets the objects to be described for various traffic scenarios such as highways, general roads, and parking lots.

[0093] For example, the combination of the traffic scene "highway" and the described object "pedestrian" is "required / required". The "required" information on the left side of the "required / required" indicates that the described object needs to be identified from the image. Moreover, the "required" information on the right side of the "required / required" indicates that the described object, whether identified or not, needs to be explained in writing.

[0094] Therefore, for the traffic scenario of "general roads," pedestrians are "required" to be identified, and if they are not identified, an explanation is required. On the other hand, for the traffic scenario of "general roads," pedestrian crossings are "required" to be identified, but no explanation is required.

[0095] Thus, the explanatory text generation unit 12 reads whether or not an object in the image needs to be explained from the dossier 14A. If it cannot identify an object in the image that is specified as needing explanation based on the read dossier information, it generates an explanatory text indicating that the object does not exist in the image as an explanatory text for each object.

[0096] Therefore, corresponding to traffic scenarios involving vehicles, descriptive text is generated only when objects that are usually nonexistent exist, and descriptive text is also generated even when objects that are usually present do not exist. This allows for the generation of more natural traffic condition descriptions, thus improving the accuracy of user search queries.

[0097] Figure 7 This is a structural diagram of Identification Table 14B. For the identified explanatory objects, Identification Table 14B defines the items analyzed using VQA section 13A and the items whose analysis results are included in the explanatory text as detailed items.

[0098] For example, VQA Section 13A analyzes the location, clothing color, and actions of pedestrians when identifying them.

[0099] In this way, the description generation unit 12 reads detailed items about objects within the image from the recognition table 14B and generates descriptions of the detailed items read from the objects identified in the image, serving as descriptions for each object. Therefore, it is possible to individually set the information to be added based on the object being described, and to generate natural traffic condition descriptions. Consequently, the accuracy of the search for the user's search query is improved.

[0100] Figure 8 This is a flowchart of the explanatory text generation section 12.

[0101] As part of the image description generation process (step S21), the description generation unit 12 causes the VQA unit 13A to perform image analysis on the still image data 111 extracted from the vehicle driving dynamic image data 11A, thereby generating image descriptions. The extraction process of the still image data 111 may be, for example, extracting continuously captured images at regular intervals of 10 seconds, or extracting 10 images from a dynamic image file at equal time intervals.

[0102] As part of the GPS description text generation process (step S22), the description text generation unit 12 generates GPS description text based on the vehicle driving GPS data 11B. At this time, the GPS data (positioning data) being used is GPS data from the same time as the still image data 111 in S21, or GPS data from the time closest to that time.

[0103] That is, the description generation unit 12 adds descriptions of at least one of the following information to the traffic condition description of the image based on GPS data read from the vehicle 91’s onboard GPS: information about the time period when the image was taken and information about the vehicle’s driving position when the image was taken.

[0104] As part of the control description document generation process (step S23), the description document generation unit 12 generates a control description document based on the vehicle driving control data 11C. At this time, the control data to be used is the control data at the same time as or closest to the time of the still image data 111 in S21.

[0105] That is, the description generation unit 12 adds descriptions of at least one of the following information—speed, acceleration / deceleration, and steering angle—to the traffic condition description of the image based on the vehicle driving control data read from the vehicle's on-board ECU (Electronic Control Unit).

[0106] As part of the traffic condition description text generation process (step S24), the description text generation unit 12 causes the summary generation unit 13B to summarize the image description text generated in step S21, the GPS description text generated in step S22, and the control description text generated in step S23, thereby generating a traffic condition description text. The description text generation unit 12 stores the generated traffic condition description text in association with the still image data 111, which is the subject of the description, in the description text database 15.

[0107] That is, the summary generation unit 13B generates a summary text that matches the input prompt. Furthermore, the description text generation unit 12 inputs a prompt containing the following content into the summary generation unit 13B and generates a traffic condition description text for the image: a traffic condition description text generation instruction text based on information contained in the traffic condition description text, information excluded from the traffic condition description text, and example information of the traffic condition description text; and a description text for each object.

[0108] Figure 9 This is a hardware structure diagram of the vehicle image analysis device 1.

[0109] The vehicle image analysis device 1 is configured as a computer 900 having a CPU 901, RAM 902, ROM 903, HDD 904, communication I / F (interface) 905, input / output I / F (interface) 906, and media I / F (interface) 907.

[0110] Communication I / F 905 is connected to an external communication device 915. Input / output I / F 906 is connected to input / output device 916. Medium I / F 907 reads and writes data to recording medium 917 and installs the document generation program stored in recording medium 917. Then, CPU 901 executes the document generation program (application program, also simply called application) read from RAM 902, thereby realizing... Figure 1 Each processing unit. Furthermore, the document generation program can be distributed via communication lines or recorded on recording media 917 such as CD-ROM.

[0111] Figure 10 It is the generation and processing of image descriptions ( Figure 8 The detailed flowchart of step S21 is as follows.

[0112] The description text generation unit 12 queries the VQA unit 13A, thereby classifying the traffic scene captured in the still image data 111 (step S211). The VQA unit 13A receives input of the still image data 111 and a query text about the traffic scene (e.g., "Where is this scene? For example, is it a road, highway, or parking lot?"), and returns a response text (e.g., highway) to the description text generation unit 12. Figure 13(Line A01). In addition, the explanatory text generation unit 12 reads the traffic scene query text that the administrator has set in advance for the vehicle image analysis device 1, so the user of the vehicle image analysis device 1 does not need to generate the traffic scene query text himself.

[0113] Next, the explanatory text generation unit 12 sequentially selects the objects to be described (pedestrians, bicycles, motor vehicles, etc.) registered in the no-need table 14A and performs the loop processing of steps S212 to S217. Hereinafter, the objects to be described selected in this loop processing will be referred to as the selected objects.

[0114] The explanatory text generation unit 12 refers to the yes / no table 14A to determine whether it is necessary to identify the selected object in the traffic scene determined in step S221 (step S212). If the result in step S212 is Yes (meaning it is necessary), the process proceeds to step S213; if the result is No, the process proceeds to step S217.

[0115] The descriptive text generation unit 12 queries the VQA unit 13A of the large-scale language model unit 13 to determine whether the selected object exists in the still image data 111. The query text is, for example, "Are there pedestrians in this scene?" Figure 13 (Line A02). The description text generation unit 12 identifies the object to be identified from the image based on the response text (object identification result) to the query text. As the identification result, the description text generation unit 12 determines whether the selected object exists in the still image data 111 (step S213). If the result in step S213 is Yes (meaning it exists), the process proceeds to step S214; if the result is No, the process proceeds to step S215.

[0116] The descriptive text generation unit 12 obtains detailed items regarding the selected object present in the still image data 111 (step S214). To do this, the descriptive text generation unit 12 refers to the recognition table 14B to obtain the detailed items corresponding to the selected object. Then, the descriptive text generation unit 12 queries the VQA unit 13A of the large-scale language model unit 13 item by item regarding the detailed items of the selected object in the still image data 111. For example, since the selected object = motor vehicle, and the detailed items corresponding to motor vehicles in the recognition table 14B include "color," the descriptive text generation unit 12 generates a query such as "What color is the car?"

[0117] The explanatory text generation unit 12 refers to the yes / no table 14A to determine whether it is necessary to explain (mention) the selected object (step S215). If the answer in step S215 is Yes (meaning it is necessary), the process proceeds to step S216; if the answer is No, the process proceeds to step S217.

[0118] The description text generation unit 12 generates a description text for the selected object by any of the following methods (step S216).

[0119] If the condition in step S213 is yes, an explanatory text for the selected object is generated based on the combination of the query text and the answer text for the detailed items of the selected object obtained in step S214.

[0120] If step S213 is not true, a description is generated indicating that no selected object was identified in the still image data 111. Figure 19 (e.g., line D07).

[0121] The description generation unit 12 ends the processing of the selected objects and determines whether the processing of all selected objects has been completed (step S217). If the result in S217 is Yes (meaning it is completed), the processing ends; if the result is No, the process switches to the unprocessed selected objects and returns to step S212.

[0122] As a result, it can generate natural traffic condition descriptions that include information about surrounding objects that should be noted and identified, in accordance with the traffic scene in which the vehicle is traveling, thereby improving the accuracy of the user's search query.

[0123] Figure 11 It means Figure 10 The flowchart illustrates an example of the specific actions involved in generating and processing image captions.

[0124] The explanatory text generation unit 12 performs classification processing based on the traffic scenes captured in the still image data 111. Figure 10 The result of step S211) is branched to the processing of each traffic scenario as described below (step S301).

[0125] • If the classification result is general road, then generate image description text for general road (step S302).

[0126] • If the classification result is highway, then generate image descriptions for highways (step S303).

[0127] • If the classification result is parking lot, then generate an image description for the parking lot (step S304).

[0128] • If the classification result is "Other", then generate an image description text specified as "Other" (step S305).

[0129] Figure 12 This is the generation and processing of image captions for general roads. Figure 11 Detailed flowchart of S302.

[0130] The explanatory text generation section 12 generates questions and answers for traffic scenarios. Figure 10 Step S211), and the question and answer text regarding whether there is a signal (step S311).

[0131] Explanatory text generation section 12 as Figure 10 The judgment process (step S213) asks whether there is a pedestrian in the image (step S312). If there is, the process proceeds to step S313; if there is no pedestrian, the process proceeds to step S314.

[0132] Explanatory text generation section 12 provides answers to questions about detailed items concerning pedestrians. Figure 10 (Step S214) generates a response text indicating the presence of a pedestrian, a question and response text indicating the pedestrian's location, a question and response text indicating the pedestrian's color, and a question and response text indicating the pedestrian's action (Step S313).

[0133] The explanatory text generation unit 12 generates a response text indicating that there are no pedestrians (step S314).

[0134] Explanatory text generation section 12 as Figure 10 The judgment process (step S213) asks whether there is a motor vehicle in the image (step S315). If there is, the process proceeds to step S316; if there is no motor vehicle, the process proceeds to step S317.

[0135] Section 12 of the explanatory text provides answers to questions regarding detailed items about motor vehicles. Figure 10 (Step S214) generates a response text indicating the presence of a motor vehicle, a question and response text indicating the location of the motor vehicle, a question and response text indicating the type of motor vehicle, a question and response text indicating the color of the motor vehicle, and a question and response text indicating the movement of the motor vehicle (Step S316).

[0136] Explanatory text generation section 12 as Figure 10 The judgment process (step S213) asks whether there is a bicycle in the image (step S317). If there is, the process proceeds to step S318; if there is no bicycle, the process proceeds to step S319.

[0137] Section 12 of the explanatory text provides answers to questions about detailed items related to bicycles. Figure 10 (Step S214) generates a response text indicating the presence of a bicycle, a question and response text indicating the bicycle's location, a question and response text indicating the bicycle's color, and a question and response text indicating the bicycle's movement (Step S318).

[0138] Explanatory text generation section 12 as Figure 10The judgment process (step S213) asks whether a pedestrian crossing exists in the image (step S319). If it exists, proceed to step S320; otherwise, end the process. The explanatory text generation unit 12 generates an answer indicating that a pedestrian crossing exists (step S320).

[0139] The above, in Figure 11 The process of generating image captions for general roads is explained. The process of generating image captions for other traffic scenarios (steps S303 to S305) is the same.

[0140] Figure 13 It means for Figure 4 The table of intermediate data resulting from the image description generation process (step S21) performed on the still image data 111.

[0141] The table uses a combination of a question and an answer to that question to create a caption for the image in each row.

[0142] For example, line A01 is the result of classifying traffic scenarios (step S211).

[0143] Rows A02 to A04 are the results of the process (step S213) to determine whether the selected object exists in the still image data 111.

[0144] Lines A05 to A13 are the results of the processing (step S214) to obtain detailed information about the selected object.

[0145] Figure 14 This indicates that the explanatory text generation section 12 is for Figure 13 The intermediate data is deleted, and the resulting output data is displayed in a table. Figure 13 The difference lies in the fact that, in the "Whether to Explain" table 14A of the "Explain Object Setting Unit 14", the following descriptions are set to not require explanation when not identified: pedestrian descriptions (line A02), bicycle descriptions (line A03), and tollbooth descriptions (line A13). Figure 14 The text has been deleted.

[0146] Figure 15 This is a diagram showing the details of the GPS description text generation process (step S22).

[0147] The description generation unit 12 generates a description that classifies the shooting time of the still image data 111 into morning, daytime, evening, night, etc., based on the time period (time zone) of the GPS information obtained from the vehicle driving GPS data 11B (step S221).

[0148] The description generation unit 12 generates a description that identifies the location of the still image data 111 as a main road and city name based on the latitude and longitude of the GPS information obtained from the vehicle driving GPS data 11B (step S222).

[0149] This allows for the generation of traffic condition descriptions that include information about the time and location of vehicle travel. Consequently, the accuracy of user search queries is improved.

[0150] Figure 16 This is a diagram showing the details of the control description text generation process (step S23).

[0151] The explanatory text generation unit 12 generates an explanatory text about driving speed control based on the vehicle driving control data 11C (step S231). The explanatory text may specify, for example, that speeds below 20 km / h are low speeds, 20 km / h to 60 km / h are medium speeds, and speeds above 60 km / h are high speeds.

[0152] The explanatory text generation unit 12 generates an explanatory text about acceleration and deceleration control based on the vehicle driving control data 11C (step S232). The explanatory text may specify, for example, that the acceleration state is when the acceleration is above a certain value, the deceleration state is when the deceleration is above a certain value, and the constant speed state is otherwise.

[0153] The explanatory text generation unit 12 generates an explanatory text about steering control based on the vehicle driving control data 11C (step S233). The explanatory text may specify, for example, whether the steering angle is above a certain value, whether it is a left turn state or a right turn state, and whether it is a straight-going state otherwise.

[0154] This allows for the generation of more natural traffic condition descriptions, including explanations of vehicle control status and actions. Consequently, it improves the accuracy of user search queries.

[0155] Figure 17 It means Figure 15 and Figure 16 An example table of explanatory text generated in the middle.

[0156] In line B01, the description text for the time period was generated in step S221.

[0157] In line B02, a description of the driving location is generated in step S222.

[0158] In line B03, a description of the driving speed is generated in step S231.

[0159] In line B04, the description text of the acceleration / deceleration state is generated in step S232.

[0160] In line B05, a description of the turning status is generated in step S233.

[0161] Figure 18 It is a table of instructions for the large-scale language model unit 13 used in the process of generating traffic condition description text (step S24).

[0162] Line C01 contains instructions for generating a traffic condition description.

[0163] Line C02 contains information contained in the traffic condition description.

[0164] Line C03 contains information excluded from the traffic condition description.

[0165] Example of a description of traffic conditions in line C04.

[0166] Therefore, in the process of generating traffic condition description text (step S24), the description text generation unit 12 generates prompt words that combine the texts from (1) to (4) below in sequence. Among them, at least one of (2) and (3) can be omitted.

[0167] (1) Instructions for the large-scale language model section 13 ( Figure 18 )

[0168] (2) GPS description text ( Figure 17 (Lines B01 and B02)

[0169] (3) Control description text ( Figure 17 (Lines B03, B04, and B05)

[0170] (4) Image description text Figure 14 )

[0171] The description text generation unit 12 inputs the generated prompt words into the large-scale language model unit 13, thereby obtaining a traffic condition description text expressed in natural language. Hereinafter, an example is shown of a traffic condition description text generated by the large-scale language model unit 13 based on the prompt words (equivalent to the description below). Figure 22 Traffic conditions description (732).

[0172] Traffic situation description: "You are driving straight at very high speed on Metropolitan Expressway 5 at noon. In this scenario, there are solid orange lane markings on the expressway. Also, there is a white truck driving forward in the road ahead of you."

[0173] This generates natural traffic condition descriptions that match the intended use and user preferences, thereby improving the accuracy of the user's search query.

[0174] Figure 19 It means according to and Figure 14 The image description is an example of a table generated from different images.

[0175] Figure 19 The image description will take images taken by vehicles traveling on ordinary roads as the subject. Therefore, whether the combination of ordinary roads and pedestrians in Table 14A is "needed" to be explained is recorded, hence the information "no pedestrians identified" is recorded in line D07.

[0176] Large-scale language modeling section 13, based on including the Figure 19 The prompts for the image descriptions and the traffic condition descriptions generated through the traffic condition description generation process (step S24) are as follows.

[0177] Traffic situation description: "You are currently driving at a slow, steady speed on a city road in Mito City at night, preparing to turn left. There are no pedestrians in this scene."

[0178] The above is for reference until... Figure 19 The accompanying diagram illustrates the database processing of traffic condition descriptions (up to the processing of data stored in description database 15). The following describes examples of the application of this database information.

[0179] Figure 20 This is a flowchart illustrating the processing of the retrieval unit 16.

[0180] The retrieval unit 16 retrieves images from the images stored in the descriptive text database 15 that match the input search text. Specifically, the retrieval unit 16 outputs images that are highly similar to the traffic condition descriptions stored in the descriptive text database 15 as search results.

[0181] In addition, as "images with high similarity", the retrieval unit 16 may, for example, list images in descending order of similarity of the search results and output the images ranked from 1st to Xth (relative ranking similarity judgment), or output images of the search results with similarity higher than a pre-set benchmark value (threshold Y) (absolute similarity judgment).

[0182] The following describes the details of the processing by the retrieval unit 16.

[0183] The retrieval unit 16 receives the search text input by the user from the input / output unit 17 (step S61). The user here is, for example, a commentator at a traffic control center managing highways, and the search text is, for example, images collected from a database showing situations similar to accidents occurring at designated locations on highways at specified times. The commentator plans to edit the image material obtained from the database to produce a report program about the accident.

[0184] The retrieval unit 16 evaluates the similarity between the user's search text and the explanatory texts stored in the explanatory text database 15 (step S62), and retrieves dynamic image information (image information) associated with the explanatory texts with high similarity from the explanatory text database 15. Similarity evaluation can utilize methods such as cosine similarity retrieval based on document vectorization.

[0185] In addition to the driving dynamic images and traffic condition descriptions obtained in step S62, the retrieval unit 16 also retrieves related information about the driving condition (such as weather information that is not found in the description database 15 and can be obtained by specifying the location, date and time of the traffic condition description from the meteorological database), and generates a response based on these results (step S63).

[0186] That is, the retrieval unit 16 can also output retrieval results based on information contained in traffic condition descriptions corresponding to images with high similarity, and related information obtained from a database different from the description database 15.

[0187] The retrieval unit 16 sends the response text of step S63 to the input / output unit 17 (step S64).

[0188] Therefore, when retrieving dynamic image data with traffic condition descriptions similar to the natural language search text input by the user, the target dynamic image can be retrieved in a short time.

[0189] Figure 21 This is a diagram showing the image retrieval interface 71 of the input / output unit 17.

[0190] The image retrieval interface 71 consists of a retrieval text input box (i.e., a retrieval text input unit 72) for the retrieval text in step S61 and a retrieval result display unit (i.e., a retrieval result display unit 73) for the retrieval text in step S64. In the retrieval result display unit 73, multiple retrieval results (scenarios 1-4) are displayed using icons or thumbnails.

[0191] Figure 22 yes Figure 21 The screen that appears when the search result (Scenario 1) is clicked. In this screen, the following information is displayed sequentially from top to bottom.

[0192] Image 731 of the search results.

[0193] • Traffic condition description text 732 of image 731, which is extracted as content with high similarity to the search text.

[0194] • GPS description (time, location), control description (vehicle type, speed), weather and other additional information 733.

[0195] Figure 23 yes Figure 21The screen displayed when the search result (Scenario 2) is clicked. (And) Figure 22 Similarly, in this reproduced image, with Figure 22 Similarly, the search results are displayed as image 741, its traffic condition description 742, and additional information 743.

[0196] Therefore, after users input their search terms in natural language, they can refer to similar traffic condition descriptions, dynamic images, and related information to search for dynamic images, thereby improving search efficiency.

[0197] According to the explanatory text generation process described above, when generating explanatory text for vehicle images, the explanatory text generation unit 12 refers to the optional table 14A and generates natural explanatory text based on the traffic scene in which the vehicle is located, including whether there are any surrounding objects that should be of concern. As a result, a database of natural explanatory texts that mention necessary surrounding objects and do not mention unnecessary ones is created, thus improving the accuracy of database retrieval based on human input search text.

[0198] Step S12: Details of the retrieval process

[0199] Next, details of the retrieval process of the scene designation unit 1001 will be explained. The retrieval process is the process of extracting image descriptions associated with a specific scene based on the traffic condition descriptions generated in step S11.

[0200] The scene designation unit 1001, as described above, includes a second generative AI 1102. Therefore, in cases where the operator is not a skilled operator, [the following is achieved]... Figure 24 The dialogue between the operators and the second generative AI1102 can also extract specific scenarios. The following is an appropriate reference. Figure 1 illustrate Figure 24 The flowchart.

[0201] In step S71, the scenario designation unit 1001 accepts a prompt word input by the operator. The prompt word is, for example, "Please tell me the events that occurred during the drive from 10:00 AM to 8:00 PM".

[0202] In step S72, the scene designation unit 1001 provides a prompt word to the second generative AI 1102, causing the second generative AI 1102 to access the description text database 15. The second generative AI 1102 generates a response text to the prompt word (step S73). An example response is "The event that occurred during driving from 10:00 AM to 8:00 PM was a collision with a pedestrian."

[0203] In step S73, the scene designation unit 1001 accepts a prompt word input by the operator. The prompt word could be, for example, "Please search for a description of traffic conditions involving pedestrian collisions."

[0204] In step S74, the scene designation unit 1001 provides prompts to the second generative AI 1102, causing the second generative AI 1102 to access the description text database 15 and extract the associated traffic condition description text 1006. When the processing in step S74 ends, Figure 24 The processing is now complete.

[0205] The prompts entered by the operator include, for example, the following: Figure 6 The system records vehicle locations such as "highway," "general road," and "parking lot." Additionally, the system includes prompts entered by operators for events such as "sudden braking," "collision with other vehicles," "leaving the lane," and "vehicle malfunction." Furthermore, the system includes prompts for times such as "from 10:00 AM to 8:00 PM" and "from 3:00 PM to 3:10 PM."

[0206] Furthermore, in this invention, the second generative AI 1102 is not mandatory. The scene designation unit 1001 may also include an interface for operators to designate a specified scene, and similarly to step S11, it directly accesses the description text database 15 based on the natural language input by the operator, without going through the second generative AI 1102, and extracts the associated traffic condition description text 1006.

[0207] Step S13: Details of condition setting processing

[0208] Next, refer to Figure 1 This section explains the details of the condition setting process in the summary condition setting unit 1003. The condition setting process is the process of specifying the traffic condition description text 1006 that meets the conditions from the search results of step S12 in order to generate the report 1005, and specifying the description objects to be mentioned in the report 1005.

[0209] For example, two conditions are stored in the summary condition setting unit 1003.

[0210] The first condition is that the period includes the scene. For example, the period including the scene represents... Figure 25 The time interval is from time (-T2) to time T1. Depending on the scenario, there is a possibility that future information at time T0 is important, and there is also a possibility that past information at time T0 is important. Therefore, there are cases where time T1 and time (-T2) take the same value, and there are also cases where they take different values.

[0211] Figure 26 It is a diagram illustrating a table that records the period of a scene.

[0212] like Figure 26The information recorded, for example, allows setting the duration of time during the period including the scene, before the scene, and after the scene for each scenario, and is stored as Table 2401 in the summary condition setting unit 1003. Furthermore, generating the report 1005 (described later) using only time T1 or time (-T2) is also within the scope of the implementation method.

[0213] In addition, a scenario includes both events and non-events. Examples of events include collisions between the vehicle and pedestrians, collisions between the vehicle and other vehicles, and collisions between the vehicle and buildings. Examples of non-events include sudden braking, leaving the lane, vehicle malfunction, violations of traffic laws, and speeding.

[0214] Figure 27 It is a diagram illustrating the table of objects to be defined for each scenario.

[0215] The second condition is to describe the object. The object is the one mentioned in report 1005. For example... Figure 27 The description, for example, can be set for each scene and stored as Table 2501 in the summary condition setting section 1003.

[0216] The first and second conditions can be set in advance by the operators. As described later, report 1005 is generated according to the conditions specified in step S13. Therefore, the conditions in step S13 function as prompts for the third generative AI 1103. That is, the prompt for the third generative AI 1103 is "Please record buildings, intersections, pedestrians, and crosswalks as the objects described in the report." By pre-storing prompt information in a certain format in the table, the generative AI can be accurately instructed.

[0217] Step S14: Details of report generation process

[0218] Next, details of the report generation process performed by the report generation device 1004 will be explained. In step S14, the third generation formula AI1103 in the report generation device 1004 generates a summary of the traffic condition description 1006 that meets the conditions, namely the report 1005, according to the conditions in step S13.

[0219] For example, if the operator selects a pedestrian collision scenario in step S12, the third generative algorithm AI1103 generates report 1005. The generated report 1005 is stored in the descriptive text database 15 in a format that the operator can retrieve using natural language. The report is described below, for example.

[0220] "The video begins with a driving scenario on a city road. This is a typical driving scene in an urban environment. As the journey progresses, new high-rise buildings appear along the way. This suggests that the area is located in a developed part of the city. This is a typical location where traffic volume and pedestrian activity increase as you approach an intersection near this landmark."

[0221] Shortly after, while driving near a crosswalk, a pedestrian appears on your right. This indicates a potential collision or distraction. The situation escalates rapidly; in the next instant, someone can be seen positioned on your windshield. This suggests a sudden and surprising scenario indicating a collision with a pedestrian. In the final scenario, while driving on the crosswalk, a person can be seen lying on the road. This indicates a high probability of a serious accident due to previous interaction with the pedestrian.

[0222] Overall, the video documents the journey from ordinary city driving to the occurrence of a major incident involving pedestrians, suggesting a serious accident at a crosswalk.

[0223] As described above, the explanatory text generation unit 12 generates control explanatory text based on the vehicle driving control data 11C as part of the control explanatory text generation process (step S23). Furthermore, the explanatory text generation unit 12 can add explanatory text regarding vehicle control to the traffic condition explanatory text 1006 of the image based on vehicle driving control data read from the vehicle's onboard ECU (Electronic Control Unit). Therefore, information from the onboard ECU can also be reflected in the report 1005.

[0224] The information from the vehicle's ECU includes, for example, information from sensors such as cameras. This sensor information includes whether the program used by the cameras to detect pedestrians or other objects has been activated, and if so, whether pedestrians or other objects are detected. Furthermore, the information from the vehicle's ECU includes, for example, information about vehicle movement. This information includes speed information, acceleration / deceleration information, acceleration information, the time derivative of acceleration (jerk), yaw rate, brake hydraulic pressure, and / or steering angle information.

[0225] For example, if Report 1005 describes a situation where the program for detecting objects such as pedestrians does not start, the developer can improve the program to ensure it starts in the same scenario. Additionally, if Report 1005 describes a situation where the program starts but does not detect objects, the developer can improve the program to enable object detection. Furthermore, if Report 1005 describes a situation where objects are successfully detected, the developer can determine if there is room for improvement in vehicle control.

[0226] In addition, similar to image descriptions, GPS descriptions, and control descriptions, the description generation unit 12 generates a third generative AI1103 to more reliably describe the program's startup status and / or the object's detection status. Figure 28 Program startup status description document 2801 Figure 29 The object detection status description 2901 is input as one of the prompt words to the third generative formula AI1103, which is also within the scope of the disclosure of this embodiment.

[0227] In addition, in order to generate Figure 28 Program startup status description document 2801 Figure 29 The object detection status description 2901 collects the program startup status and / or object detection status from at least one of the vehicles 91 to 93 in association with time information and stores it in the driving log database 11, which is also within the scope of this embodiment. Furthermore, the program startup status description 2801 and / or object detection status description 2901 are output as one of the additional information 733 and 743 in association with the image, which is also within the scope of this embodiment.

[0228] According to this embodiment, the efficiency of generating report 1005 can be improved. Because the preceding and following times and / or the objects to be described can be specified, the generation of lengthy report 1005 can be avoided. This invention is particularly effective for the development of autonomous vehicles.

[0229] The structure and effects of the present invention are described below. [1]

[0231] An article generation system (homework assistance system 1000) is characterized by comprising:

[0232] The image analysis unit (vehicle image analysis device 1) generates a status description text using generative AI (Artificial Intelligence) (first generative AI 1101) based on the images captured by the camera and stores it in the database;

[0233] The scene designation unit (1001) designates the scene as the object of description;

[0234] The scene retrieval unit (1002) retrieves a situation description text associated with the description object scene specified by the scene designation unit (1001) from the situation description text stored in the database;

[0235] The summary condition setting unit (1003) stores the time and the object of description corresponding to the scene specified by the scene designation unit (1001); and

[0236] The report generation unit (report generation device 1004) uses the situation description text retrieved by the scene retrieval unit (1002) that corresponds to the time and the object of description stored in the summary condition setting unit (1003) to generate a report (1005) about the object of description using a generative AI (third generative AI 1103).

[0237] This can improve the efficiency of image-based report generation. [2]

[0239] The article generation system as described in aspect 1 is characterized by:

[0240] The scene designation unit (1001) uses generative AI (second generative AI 1102) to designate the scene as the object of description.

[0241] By using generative AI, it is easy to specify the context of the object. [3]

[0243] The article generation system as described in aspect 1 is characterized by:

[0244] The camera in question is a vehicle-mounted camera.

[0245] This allows for the generation of effective reports on the development of autonomous vehicles. [4]

[0247] The article generation system as described in aspect 3 is characterized by:

[0248] The report (1005) includes information about the vehicle's ECU.

[0249] This allows for the generation of effective reports on the development of autonomous vehicles. [5]

[0251] The article generation system as described in aspect 4 is characterized by:

[0252] Information about the vehicle ECU includes information about the sensors connected to the vehicle ECU and information about vehicle movement.

[0253] This allows for the generation of effective reports on the development of autonomous vehicles. [6]

[0255] An article generation program that causes a computer to perform:

[0256] The steps for generating a condition description text using generative AI (first generative AI 1101) based on images captured by a camera;

[0257] The step of saving the aforementioned situation description in the database;

[0258] The steps to specify the scenario of the object to be described;

[0259] The step of retrieving a status description associated with the specified description object scenario from the status descriptions stored in the database;

[0260] The steps of saving the time and the object of description corresponding to the specified description object scene; and

[0261] The step of generating a report about the described object using generative AI (third generative AI 1103) based on the retrieved condition description texts corresponding to the time and the described object in the described object scene.

[0262] This can improve the efficiency of image-based report generation. [7]

[0264] A method for producing a report database, characterized by comprising:

[0265] The steps for generating a condition description text using generative AI (first generative AI 1101) based on images captured by a camera;

[0266] The step of saving the aforementioned situation description in the database;

[0267] The scene specification unit (1001) specifies the steps of the scene to be described as the object of description;

[0268] The scene retrieval unit (1002) retrieves a situation description text associated with the description object scene specified by the scene designation unit (1001) from the situation description text stored in the database;

[0269] The summary condition setting unit (1003) saves the time and steps of the described object corresponding to the specified described object scene;

[0270] The steps include: using the situation description text retrieved by the scene retrieval unit (1002) that corresponds to the time and object of description stored by the summary condition setting unit (1003), and using generative AI (third generative AI 1103) to generate a report (1005) about the object of description; and

[0271] The step of generating a database (1007) including the report (1005).

[0272] This can improve the efficiency of image-based report generation.

[0273] Furthermore, the content of this invention can also be applied to uses other than the development of autonomous vehicles. For example, the content of this invention can also be applied to the following fields.

[0274] (1) In the field of logistics, the present invention can be applied to generate daily reports for drivers.

[0275] (2) In the field of education, this invention can be applied to the summarization of related content based on keywords from online lectures, seminars, and academic conferences. Additionally, it can also be applied to the summarization of content from lecturers who ask questions.

[0276] (3) In the field of security, the present invention can be applied to reports summarizing images from security cameras. The reports can be used, for example, to detect vehicle theft, robbery, theft, and suspicious persons at night.

[0277] (4) In the retail sector, this invention can be applied to the generation of reports using data from burglar alarm cameras. These reports can be used to retrieve theft scenarios. Furthermore, this invention can also be applied to the generation of reports using camera data in unmanned stores and self-checkout systems.

[0278] (5) In the field of broadcasting, the present invention can be applied to scoring scenes in sports competitions and to the summarization of dynamic images before and after exciting moments in programs.

[0279] (6) In the medical field, this invention can be applied to the summarization of surgical dynamic images. The generated report can function as a textbook. Therefore, by referring to the report, the procedure can be confirmed. In addition, the required procedure can be retrieved. As a result, the efficiency of training can be improved.

[0280] (7) In the fields of factory, construction site, and warehouse management, this invention can be applied to generate accident reports. Accidents can include objects falling or tipping over. Reports can be generated based on dynamic images of the scene. Symptoms of malfunctions can be retrieved from the reports. In addition, this invention can also be applied to outputting work reports with dynamic images, generating inspection / repair reports, and generating process sheets that record maintenance procedures.

[0281] (8) In the manufacturing sector, this invention can summarize (abstract) the techniques of skilled technicians such as craftsmen.

[0282] (9) In the field of building elevator management, this invention can generate a report summarizing the number of people entering and exiting for each time period. The report can be used for user profile analysis.

[0283] (10) In the field of home appliances, the present invention can generate a consumption report based on consumption records obtained from a cold storage with sensors such as cameras.

[0284] (11) The present invention can also be applied to cooking video summaries, product / service introduction video summaries, and pet / child action history summaries.

[0285] Furthermore, the present invention is not limited to the embodiments described above. Various other applications and modifications can be adopted as long as they do not depart from the spirit of the invention as described in the claims. For example, the structure of the vehicle image analysis device 1 has been described in detail and specifically for the purpose of easily understanding the present invention, and it is not limited to having all the constituent elements described. In addition, a part of the structure of a certain embodiment can be replaced with the constituent elements of other embodiments. In addition, the constituent elements of other embodiments can be added to the structure of a certain embodiment. In addition, for a part of the structure of each embodiment, other constituent elements can be added, replaced, or deleted.

[0286] Furthermore, the aforementioned structures, functions, and processing units can be partially or entirely implemented in hardware, for example, through design in integrated circuits. As hardware, generalized processor devices such as FPGAs (Field Programmable Gate Arrays) and ASICs (Application Specific Integrated Circuits) can be used.

[0287] Furthermore, each component of the vehicle image analysis device 1 described above can be implemented in any hardware as long as each piece of hardware can send and receive information via a network. Additionally, processing performed by a single processing unit can be implemented using a single piece of hardware, or it can be implemented using distributed processing by multiple pieces of hardware.

[0288] Explanation of reference numerals in the attached figures

[0289] 1. Vehicle image analysis device (image analysis unit)

[0290] 8 communication lines

[0291] 11 Driving Log Database

[0292] 11A vehicle driving dynamic image data

[0293] 11B vehicle driving GPS data

[0294] 11C Vehicle Driving Control Data

[0295] 12. Explanatory Text Generation Section

[0296] 13 Large-scale language modeling department

[0297] 13A VQA Department

[0298] 13B Abstract Generation Department

[0299] 14. Description of Object Setting Section

[0300] 14A Requirement Form

[0301] 14B Identification Table

[0302] 15. Explanatory Text Databases (Image Databases)

[0303] 16. Retrieval Department

[0304] 17 Input / Output Section

[0305] Vehicles 91-93

[0306] 100 Image Description System

[0307] 111 still image data

[0308] 732, 742 Situation Description

[0309] 901 CPU

[0310] 902 RAM

[0311] 903 ROM

[0312] 904 HDD

[0313] 905 communication I / F

[0314] 906 Input / Output I / F

[0315] 907 Medium I / F

[0316] 1000 Operation Assistance System

[0317] 1001 Scenario Designation Department

[0318] 1002 Scene Retrieval Department

[0319] 1003 Summary Condition Setting Department

[0320] 1004 Report Generation Device

[0321] 1005 Report

[0322] 1006 Traffic Conditions Explanation

[0323] 1101 First Generative AI

[0324] 1102 Second Generative AI

[0325] 1103 Third Generative AI.

Claims

1. An article generation system, characterized in that, include: The image analysis department uses generative AI (Artificial Intelligence) to generate situation descriptions based on images captured by cameras and stores them in a database; The scene specification section specifies the scene as the object of description. The scene retrieval unit retrieves a situation description text associated with the description object scene specified by the scene designation unit from the situation description text stored in the database. The summary condition setting unit stores the period and the object of description corresponding to the object of description scene specified by the scene specification unit; and The report generation unit uses the situation description text retrieved by the scene retrieval unit that corresponds to the time and the object of description stored by the summary condition setting unit to generate a report about the object of description using generative AI.

2. The article generation system as described in claim 1, characterized in that: The scene designation unit utilizes generative AI to designate the scene as the object of description.

3. The article generation system as described in claim 1, characterized in that: The camera in question is a vehicle-mounted camera.

4. The article generation system as described in claim 3, characterized in that: The report includes information about the vehicle's ECU.

5. The article generation system as described in claim 4, characterized in that: The information about the vehicle ECU includes information about the sensors connected to the vehicle ECU and information about vehicle movement.

6. An article generation program, characterized in that: To make the computer perform: The steps for generating condition descriptions using generative AI based on images captured by a camera; The step of saving the aforementioned situation description in the database; The steps to specify the scenario of the object to be described; The step of retrieving a status description associated with the specified description object scenario from the status descriptions stored in the database; The steps for saving the time and the object of description corresponding to the specified scene; and The step of using generative AI to generate a report about the described object using the retrieved condition description texts that correspond to the time and the described object in the described object scenario.

7. A method for producing a report database, characterized in that, include: The steps for generating condition descriptions using generative AI based on images captured by a camera; The step of saving the aforementioned situation description in the database; The scene specification section specifies the steps of the scene as the object of description. The step of the scene retrieval unit retrieving a situation description text associated with the description object scene specified by the scene designation unit from the situation description text stored in the database; The summary condition setting unit saves the period and the steps of the described object corresponding to the specified described object scene; The step of generating a report about the object of description using generative AI, which is based on the situation description text retrieved by the scene retrieval unit and corresponds to the period and the object of description stored by the summary condition setting unit. and The step of generating a database containing the report.

Citation Information

Patent Citations

  • Video image summary instrument, explanation note forming instrument, video image summary method, explanation note forming method and program

    JP2005109566A