Text creation system, text creation program, and report database production method
The text generation system enhances report creation efficiency in autonomous vehicle development by using generative AI to analyze images, designate scenes, and generate reports, addressing the inefficiencies of existing technologies.
Patent Information
- Application Number
- JP2024024143
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-20
- Publication Date
- 2025-09-01
AI Technical Summary
Existing technologies do not efficiently support the creation of reports based on images, particularly in the context of autonomous vehicle development, where detailed reports are needed for test drives, including incidents and long durations, without considering the development of autonomous vehicles.
A text generation system utilizing generative AI to analyze images, designate scenes, search for relevant descriptions, and create reports based on summary conditions, incorporating scene designation, search, and report creation units to enhance efficiency.
Improves the efficiency of creating reports by generating detailed and accurate descriptions from images, enabling efficient report creation for autonomous vehicle development.
Smart Images

Figure 2025127401000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a text generation system, a text generation program, and a method for producing a report database. [Background technology]
[0002] Patent Document 1 below is an example of a technology for extracting desired scenes from a video. Patent Document 1 discloses a method for generating explanatory text, particularly for sports video content. In particular, Patent Document 1 discloses a method for creating a video summary with explanatory text added to video sections that are of interest to the user. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-109566 Summary of the Invention [Problem to be solved by the invention]
[0004] However, Patent Document 1 does not take into consideration the development of autonomous vehicles. For example, in the development of autonomous vehicles, developers may prepare reports explaining the results of test drives. For example, developers of autonomous vehicles may analyze videos captured by sensors such as cameras mounted on the autonomous vehicles, identify scenes that should be reported, and prepare reports so that the results of test drives can be used in subsequent development. If an accident or other incident occurs during a test drive, a report explaining the situation before and after the accident may be required. In addition, if the test drive lasts for an entire day, reports covering several hours may be required.
[0005] The present invention has been made in consideration of the above circumstances, and its main object is to improve the efficiency of the work of creating reports based on images. [Means for solving the problem]
[0006] In order to solve the above problems, the text generation system of the present invention has the following features. The present invention is characterized by comprising an image analysis unit that uses generative AI (Artificial Intelligence) to generate a situation description from an image taken by a camera and stores the generated description in a database; a scene designation unit that designates a scene to be described that is the subject of the description; a scene search unit that searches the situation descriptions stored in the database for a situation description related to the scene to be described designated by the scene designation unit; a summary condition setting unit that saves a period and an object to be described according to the scene to be described designated by the scene designation unit; and a report creation unit that uses generative AI to generate a report on the object to be described using a situation description corresponding to the time and object to be described stored by the summary condition setting unit from the situation descriptions searched by the scene search unit.
[0007] The text generation program of the present invention causes a computer to execute the following steps: generating a situation description from an image captured by a camera using a generation AI; storing the situation description in a database; specifying a target scene to be described; searching for a situation description related to the specified target scene from the situation descriptions stored in the database; saving a period and an object to be described according to the specified target scene; and generating a report on the object to be described using a generation AI using a situation description corresponding to the time and object to be described according to the target scene from the searched situation description.
[0008] The method for producing a report database of the present invention is characterized by comprising the steps of generating a situation description using a generation AI from an image taken by a camera, storing the situation description in a database, a scene designation unit designating a target scene to be described, a scene search unit searching for a situation description related to the target scene designated by the scene designation unit from the situation descriptions stored in the database, a summary condition setting unit saving a period and an object to be described according to the designated target scene, generating a report regarding the object to be described using a generation AI using a situation description corresponding to the time and object to be described stored by the summary condition setting unit among the situation descriptions searched by the scene search unit, and generating a database including the report. Other features will be described later. [Effects of the Invention]
[0009] According to the present invention, the work of creating a report based on an image can be made more efficient. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a configuration diagram of a work support system according to an embodiment of the present invention. [Figure 2] 10 is a flowchart of a report creation process according to the present embodiment. [Figure 3] FIG. 2 is a diagram showing details of a driving log database according to the present embodiment. [Figure 4] FIG. 2 is a diagram showing an example of still image data that is part of vehicle driving video data according to the present embodiment. [Figure 5] 1 is a configuration diagram of a vehicle image analysis device according to an embodiment of the present invention. [Figure 6] FIG. 10 is a diagram illustrating the configuration of a necessity table according to the present embodiment. [Figure 7] FIG. 2 is a diagram showing the configuration of a recognition table according to the present embodiment. [Figure 8]10 is a flowchart of the explanation generation unit according to the present embodiment. [Figure 9] FIG. 1 is a hardware configuration diagram of a work support system according to an embodiment of the present invention. [Figure 10] 10 is a detailed flowchart of the image description generation process according to the present embodiment. [Figure 11] 11 is a flowchart showing an example of a specific operation of the image caption generation process described with reference to FIG. 10 according to this embodiment. [Figure 12] 10 is a detailed flowchart of a process for generating an image description for an ordinary road according to the present embodiment. [Figure 13] 5 is a table showing intermediate data resulting from execution of image caption generation processing on the still image data of FIG. 4 according to this embodiment. [Figure 14] 14 is a table showing output data resulting from the intermediate data of FIG. 13 being deleted by an explanation generation unit according to this embodiment. [Figure 15] FIG. 10 is a diagram showing details of a process for generating a GPS description according to the present embodiment. [Figure 16] 10A to 10C are diagrams illustrating details of a process for generating a control explanation according to the present embodiment. [Figure 17] 17 is a table showing an example of the explanatory text generated in FIGS. 15 and 16 according to this embodiment. [Figure 18] 10 is a table showing instructions to a large-scale language model unit used in the process of generating traffic condition descriptions according to the present embodiment. [Figure 19] 15 is a table showing an example of the image caption generated from an image different from that of FIG. 14 according to this embodiment. [Figure 20] 10 is a flowchart showing the processing of a search unit according to the present embodiment. [Figure 21] FIG. 2 is a diagram showing an image search interface of an input / output unit according to the present embodiment. [Figure 22] This is a playback screen when the search result (scene 1) in FIG. 21 related to this embodiment is clicked. [Figure 23]This is a playback screen when the search result (scene 2) in FIG. 21 related to this embodiment is clicked. [Figure 24] 10 is a flowchart showing a search process. [Figure 25] FIG. 1 is a diagram illustrating a period including a scene. [Figure 26] FIG. 10 is a diagram illustrating a table in which periods including scenes are recorded. [Figure 27] FIG. 10 is a diagram illustrating a table in which explanation objects are set. [Figure 28] FIG. 10 is a diagram illustrating a program activation status description. [Figure 29] FIG. 10 is a diagram illustrating an object detection situation description. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0012] FIG. 1 is a configuration diagram of a work support system 1000. The work support system 1000 includes a travel log database 11, which is a database, a vehicle image analysis device 1, a scene designation unit 1001, a scene search unit 1002, a summary condition setting unit 1003, and a report creation device 1004. This work support system 100 operates as a document generation system that supports the creation of reports.
[0013] The vehicle image analysis device 1 includes a first generation AI (Artificial Intelligence) 1101. The vehicle image analysis device 1 functions as an image analysis unit that generates a traffic condition description 1006 from an image captured by a camera using the first generation AI 1101 and stores the generated description in the description database 15 shown in FIG.
[0014] The scene designation unit 1001 includes a second generated AI 1102, and uses the second generated AI 1102 to designate a scene to be explained. The scene search unit 1002 includes a second generation AI 1102 and searches for a traffic condition description 1006 related to the scene to be described specified by the scene specification unit 1001 from the traffic condition description stored in the description database 15 of Figure 4.
[0015] When a period before and after a scene and an object to be explained (for example, the other party in the accident or the point where sudden braking occurs) are specified, the summary condition setting unit 1003 saves the period before and after the specified scene to be explained and the object to be explained. The report creation device 1004 includes a third generation AI 1103. The first generation AI 1101, the second generation AI 1102, and the third generation AI 1103 may be separate generation AIs or may be a single generation AI. When the report creation device 1004 creates a report 1005, it appropriately creates a report database 1007 to store the report 1005.
[0016] The first generated AI 1101 includes a model trained using images to identify scenes from images captured by a camera, while the second generated AI 1102 and the third generated AI 1103 include models trained using text. The first generation AI 1101, the second generation AI 1102, or the third generation AI 1103 may use existing services such as GPT-4, CLIP, BLIP, or BLIP-2.
[0017] FIG. 2 is a flowchart of the report creation process according to this embodiment. In this embodiment, the following four steps are included in order to create the report 1005 in Fig. 1. Each step will be described in detail below. In step S11, the vehicle image analyzing device 1 uses an image captured by a sensor such as a camera to generate a traffic condition explanation 1006 using the first generation AI 1101. In step S12, the scene search unit 1002 searches for scenes necessary for creating the report 1005 from the traffic condition description 1006 generated in step S11 using the second generated AI 1102. In step S13, the summarization condition setting unit 1003 specifies a summarization period and an object to be explained within this summarization period. In step S14, the report creation device 1004 creates a report 1005 using the third creation AI 1103 based on the search result, the summary period, and the object to be explained during the summary period. This makes it possible to improve the efficiency of the work of creating reports based on images. In step S15, report creation device 1004 stores the created report 1005 and the corresponding driving log in report database 1007, and the processing in FIG. 2 ends.
[0018] Step S11: Details of the explanation generation process 1 are so-called connected cars that can communicate with the work support system 1000 via a communication line 8. Each of the vehicles 91 to 93 transmits various measurement data measured by the vehicle itself to a travel log database 11 via the communication line 8. Each vehicle 91 to 93 is equipped with an on-board camera that captures images. The vehicle image analysis device 1 generates a traffic situation description for the image captured by the on-board camera. The communication line 8 can be a general public line network, whether wired or wireless, such as a fifth-generation mobile communication system (5G), which enables multiple simultaneous connections and ultra-low latency. Furthermore, by utilizing the characteristics of new mobile phone systems beyond 5G, it is possible to expect effects such as online (real-time while driving) description generation.
[0019] FIG. 3 is a diagram showing the details of the driving log database 11. As shown in FIG. The travel log database 11 stores the measurement data from each of the vehicles 91 to 93 by data type as follows. The vehicle driving video data 11A is video data captured by an in-vehicle camera (not shown) while the vehicle is driving or stopped. The vehicle driving GPS data 11B is in-vehicle GPS (Global Positioning System) data. The vehicle driving control data 11C is vehicle control data such as speed, acceleration / deceleration, steering angle, etc., as vehicle driving log data obtained from an on-board ECU (Electronic Control Unit) or the like. Note that GPS is an example of a satellite positioning system.
[0020] FIG. 4 is a diagram showing an example of still image data 111, which is a part of the vehicle driving video data 11A. The still image data 111 is an example of data extracted from vehicle traveling video data 11A captured by each of the vehicles 91 to 93 traveling on the expressway.
[0021] FIG. 5 is a diagram showing the configuration of the vehicle image analysis device 1. 3, the vehicle image analysis device 1 includes an explanation generation unit 12, a large-scale language model unit 13, an explanation target setting unit 14, an explanation target database (image database) 15, a search unit 16, and an input / output unit 17. The explanation target setting unit 14 includes a necessity table 14A and a recognition table 14B. The large-scale language model unit 13 includes a VQA (Visual Question Answer) unit 13A and a summary generation unit 13B. The various data stored in the vehicle image analysis device 1 (driving log database 11, necessity table 14A, and recognition table 14B) may be stored in a storage device (not shown) outside the vehicle image analysis device 1, and may be configured to be accessible from the storage device via a network from the vehicle image analysis device 1.
[0022] The explanatory sentence generation unit 12 analyzes the driving situation of each vehicle using the large-scale language model unit 13 and the explanation target setting unit 14 based on the data in the driving log database 11, and generates natural sentences that explain the driving situation of the vehicle according to the analysis results. To this end, the explanatory sentence generation unit 12 analyzes whether the driving situation of the vehicle corresponds to any of the traffic scenes classified in advance.
[0023] The large-scale language model unit 13 is called by the explanation generation unit 12. The large-scale language model unit 13 is realized by LAVIS (LAnguage VISion) or the like, which uses a natural language conversational method as an interface for exchanging questions and answers, and has the following processing units.
[0024] The VQA unit 13A responds to queries about images in natural language in natural language. To this end, the VQA unit 13A prepares an image recognition model such as a CNN (Convolutional Neural Network) using training data in advance, and inputs a query to the image recognition model to obtain a description of the corresponding image.
[0025] The summary generation unit 13B responds by summarizing (integrating multiple sentences) the contents of the input natural language (prompt) as a traffic situation explanation. The summary generation unit 13B may use existing services such as GPT-4, CLIP, BLIP, and BLIP-2 as a text generation AI service that performs summarization and translation processing.
[0026] The explanation object setting unit 14 refers to the following table, sets explanation objects corresponding to traffic scenes in the explanation sentence generation unit 12, and stores them in the explanation sentence database 15. The necessity table 14A (FIG. 6) associates an explanation object with each traffic scene and defines the degree of importance of whether or not to mention the explanation for each explanation object. The recognition table 14B (FIG. 7) defines the detailed content to be mentioned in the description for each individual object to be described.
[0027] The explanation database 15 stores the traffic condition explanations generated by the explanation generation unit 12.
[0028] In this way, the explanation generation unit 12 executes the following steps (1) to (3). (1) Scenes appearing in the images received from the camera are identified, and for each identified scene, recognition necessity information for the object in the image is read from the necessity table 14A. (2) Recognizes objects in the image that require recognition based on the read recognition necessity information, and generates an explanatory text for each object from the recognition results. (3) A traffic situation description for the image is generated based on the identified scene and the description for each object, and the traffic situation description for the image and the image are associated with each other and stored in the description database 15.
[0029] 6 is a diagram showing the configuration of the necessity table 14A. The explanation object setting unit 14 sets explanation objects for each traffic scene such as an expressway, a general road, a parking lot, and the like. For example, the combination of a traffic scene "highway" and an explanation object "pedestrian" is "necessary / necessary." The recognition necessity information "necessary" on the left side of this "necessary / necessary" notation indicates that the explanation object needs to be recognized from the image. Also, the explanation necessity information "necessary" on the right side of the "necessary / necessary" notation indicates that the explanation object, whether recognized or not recognized from the image, needs to be explained in the explanation text. Therefore, in the traffic scene "general road", pedestrian recognition is "necessary", and even if pedestrians are not recognized, an explanation is "necessary". On the other hand, in the traffic scene "general road", crosswalk recognition is "necessary", but an explanation is not required.
[0030] In this way, the explanation generation unit 12 reads the explanation necessity information for the object in the image from the necessity table 14A, and if the object specified as requiring explanation in the read explanation necessity information cannot be recognized from within the image, an explanation is generated for each object to the effect that the object does not exist in the image. This allows for the generation of descriptions of objects that are not normally present in a traffic scene in which a vehicle is traveling, only when they are present, and for objects that are normally present in a traffic scene in which a vehicle is traveling, descriptions can be generated even when they are not present. This has the effect of generating more natural traffic situation descriptions and improving the search accuracy for user search queries.
[0031] 7 is a diagram showing the configuration of the recognition table 14B. The recognition table 14B defines, as detailed items, items to be analyzed by the VQA unit 13A for the recognized explanation object and items to include the analysis results in the explanation. For example, when a pedestrian is recognized, the VQA unit 13A analyzes the location, color of clothing, and movement of the pedestrian.
[0032] In this way, the description generator 12 reads detailed items about the objects in the image from the recognition table 14B, and generates a description for each object based on the read detailed items about the objects recognized in the image. This makes it possible to individually set the information to be added depending on the object being described, and to generate natural traffic situation descriptions. This improves the search accuracy for user search queries.
[0033] FIG. 8 is a flowchart of the explanation generation unit 12. As an image caption generation process (step S21), the caption generation unit 12 generates an image caption by having the VQA unit 13A perform image analysis on the still image data 111 extracted from the vehicle driving video data 11A. The extraction process of the still image data 111 is a process of extracting images continuously captured at a fixed interval, for example, every 10 seconds, or extracting 10 images at equal time intervals from one video file.
[0034] The description generator 12 generates a GPS description based on the vehicle driving GPS data 11B as a GPS description generation process (step S22). At this time, the target GPS data (positioning data) is the GPS data of the same time as the still image data 111 of S21 or the GPS data of the closest time. In other words, based on the GPS data read from the vehicle's onboard GPS 91, the description generation unit 12 adds a description regarding at least one of the information on the time period when the image was taken and the information on the driving position when the image was taken to the traffic situation description of the image.
[0035] The description generation unit 12 generates a control description based on the vehicle driving control data 11C as a control description generation process (step S23). At this time, the control data to be used is the control data that is the same time as the still image data 111 of S21 or the control data that is closest in time to the still image data 111 of S21. In other words, the description generation unit 12 adds a description regarding at least one of speed information, acceleration / deceleration information, and steering angle information to the traffic situation description of the image based on the vehicle driving control data read from the vehicle 91's onboard ECU (Electronic Control Unit).
[0036] As a traffic condition description generation process (step S24), the description generation unit 12 generates a traffic condition description by having the summary generation unit 13B summarize the image description generated in step S21, the GPS description generated in step S22, and the control description generated in step S23. The description generation unit 12 associates the generated traffic condition description with the still image data 111 that is the description target of the traffic condition description, and stores them in the description database 15. That is, summary generation unit 13B generates a summary in accordance with the input prompt. Then, description generation unit 12 generates a traffic condition description for the image by inputting to summary generation unit 13B a prompt including information to be included in the traffic condition description, information to be excluded from the traffic condition description, and an instruction for generating a traffic condition description based on information on example sentences of the traffic condition description, and a description for each object.
[0037] FIG. 9 is a hardware configuration diagram of the vehicle image analyzing device 1. As shown in FIG. The vehicle image analyzing device 1 is configured as a computer 900 having a CPU 901 , a RAM 902 , a ROM 903 , a HDD 904 , a communication I / F 905 , an input / output I / F 906 , and a media I / F 907 . The communication I / F 905 is connected to an external communication device 915. The input / output I / F 906 is connected to an input / output device 916. The media I / F 907 reads and writes data from a recording medium 917 and installs a document generation program stored in the recording medium 917. Furthermore, the CPU 901 embodies each processing unit in FIG. 1 by executing the document generation program (also called an application or an app for short) loaded into the RAM 902. This document generation program can also be distributed via a communication line or recorded on a recording medium 917 such as a CD-ROM and distributed.
[0038] FIG. 10 is a detailed flowchart of the image caption generation process (step S21 in FIG. 8). The description generation unit 12 classifies the traffic scenes captured in the still image data 111 by querying the VQA unit 13A (step S211). The VQA unit 13A receives the still image data 111 and a traffic scene query (e.g., "Where is this scene? For example, is it a road, a highway, or a parking lot?") and returns an answer (e.g., a highway) to the description generation unit 12 (line A01 in FIG. 13). Note that the description generation unit 12 reads the traffic scene query set in advance by an administrator in the vehicle image analysis device 1, so the user of the vehicle image analysis device 1 does not need to generate the traffic scene query by themselves. Thereafter, the explanation generation unit 12 executes the loop processing of steps S212 to S217 by sequentially selecting explanation objects registered in the necessity table 14A (pedestrian, bicycle, automobile, ...). Hereinafter, the explanation object selected in this loop processing will be referred to as the selected object.
[0039] The explanation generator 12 determines whether or not recognition of the selected object is necessary in the traffic scene identified in step S221 by referring to the necessity table 14A (step S212). If Yes (necessary) in step S212, the process proceeds to step S213, and if No, the process proceeds to step S217. The explanation generation unit 12 queries the VQA unit 13A of the large-scale language model unit 13 as to whether or not the selected object exists in the still image data 111. This query is, for example, "Are there any pedestrians in this scene?" (line A02 in FIG. 13). The explanation generation unit 12 recognizes the object designated as requiring recognition from within the image using a response sentence (object recognition result) to this query. As a result of this recognition, the explanation generation unit 12 determines whether or not the selected object exists in the still image data 111 (step S213). If the answer is Yes (exists) in step S213, the process proceeds to step S214; if the answer is No, the process proceeds to step S215.
[0040] The explanation generation unit 12 acquires detailed items about the selected object present in the still image data 111 (step S214). To do this, the explanation generation unit 12 refers to the recognition table 14B to acquire detailed items corresponding to the selected object. Then, the explanation generation unit 12 queries the VQA unit 13A of the large-scale language model unit 13 for detailed items about the selected object in the still image data 111 one by one. For example, since the selected object is a car and the detailed items corresponding to the car in the recognition table 14B include "color," the explanation generation unit 12 generates a query such as "What color is the car?"
[0041] The explanation generation unit 12 determines whether an explanation (mention) of the selected object is necessary by referring to the necessity table 14A (step S215). If Yes (necessary) in step S215, the process proceeds to step S216, and if No, the process proceeds to step S217. The explanation generator 12 generates an explanation for the selected object by one of the following methods (step S216). If the answer to step S213 is Yes, an explanation of the selected object is generated from a combination of the inquiry text and the response text for the detailed items of the selected object acquired in step S214. If the result of step S213 is No, an explanatory text for the selected object is generated to the effect that the selected object was not recognized in the still image data 111 (eg, line D07 in FIG. 19). The explanation generation unit 12 determines whether or not the processing for all selection objects has been completed by finishing the processing for the current selection object (step S217). If the answer is Yes (completed) in step S217, the processing ends, and if No, the processing switches to an unprocessed selection object and returns to step S212. This allows for the generation of natural traffic situation descriptions that include explanations of surrounding objects that should be noted and confirmed according to the traffic scene in which the vehicle is traveling, thereby improving the search accuracy for user search queries.
[0042] FIG. 11 is a flowchart showing an example of a specific operation of the image caption generation process described with reference to FIG. The description generator 12 branches into processing for each traffic scene as follows (step S301) depending on the result of the processing for classifying traffic scenes captured in the still image data 111 (step S211 in FIG. 10). If the classification result is that of an ordinary road, an image description for an ordinary road is generated (step S302). If the classification result is an expressway, an image description for the expressway is generated (step S303). If the classification result is a parking lot, an image description for the parking lot is generated (step S304). If the classification result is "other," an image description specified for "other" is generated (step S305).
[0043] FIG. 12 is a detailed flowchart of the process of generating image captions for general roads (S302 in FIG. 11). The explanation generator 12 generates a traffic scene question and answer (step S211 in FIG. 10) and a traffic light question and answer (step S311). As the determination process (step S213) of FIG. 10, the caption generating unit 12 asks whether there is a pedestrian in the image (step S312), and if there is, the process proceeds to step S313, and if there is no pedestrian, the process proceeds to step S314. The explanation generation unit 12 generates answer sentences indicating the presence of a pedestrian, a question and answer sentence about the location of the pedestrian, a question and answer sentence about the color of the pedestrian, and a question and answer sentence about the action of the pedestrian (step S313) as answers to questions about detailed items of the pedestrian (step S214 in FIG. 10). The explanation generator 12 generates a reply sentence indicating that there is no pedestrian (step S314).
[0044] As the determination process (step S213) of FIG. 10, the description generation unit 12 asks whether there is a car in the image (step S315), and if there is, the process proceeds to step S316, and if there is no car, the process proceeds to step S317. The explanatory text generation unit 12 generates, as questions and answers about the detailed items of the car (step S214 in FIG. 10), answer sentences indicating the existence of the car, a question and answer sentence about the location of the car, a question and answer sentence about the car's model, a question and answer sentence about the car's color, and a question and answer sentence about the car's operation (step S316). As the determination process (step S213) of FIG. 10, the description generator 12 asks whether a bicycle is present in the image (step S317). If present, the process proceeds to step S318, and if not, the process proceeds to step S319. The explanation generation unit 12 generates answer sentences indicating the existence of a bicycle, a question and answer sentence about the location of the bicycle, a question and answer sentence about the color of the bicycle, and a question and answer sentence about the operation of the bicycle (step S318) as answers to questions about the details of the bicycle (step S214 in FIG. 10).
[0045] 10 (step S213), the explanation generation unit 12 asks whether a crosswalk is present in the image (step S319), and if present, proceeds to step S320, and if not, ends the process. The explanation generation unit 12 generates a response sentence indicating the presence of a crosswalk (step S320). Although the process of generating an image caption for an ordinary road has been described above with reference to FIG. 11, the process of generating an image caption for other traffic scenes (steps S303 to S305) is similar.
[0046] FIG. 13 is a table showing intermediate data resulting from the image caption generation process (step S21) performed on the still image data 111 of FIG. In this table, each row is a combination of a question and an answer to that question, which constitutes an image description. For example, row A01 is the result of the process of classifying traffic scenes (step S211). Rows A02 to A04 are the results of the process of determining whether or not the object to be selected exists in the still image data 111 (step S213). Lines A05 to A13 are the results of the process of acquiring detailed items about the selected object (step S214).
[0047] Fig. 14 is a table showing output data resulting from the explanation generation unit 12 deleting unnecessary data from the intermediate data of Fig. 13. The difference from Fig. 13 is that the explanation for pedestrians (row A02), bicycles (row A03), and toll booths (row A13), which are deemed not to require explanation if not recognized in the necessity table 14A of the explanation target setting unit 14, have been deleted in Fig. 14.
[0048] FIG. 15 is a diagram showing details of the GPS description generation process (step S22). The explanation generator 12 generates an explanation that classifies the capture time of the still image data 111 into morning, afternoon, evening, night, etc. according to the time period of the GPS information obtained from the vehicle driving GPS data 11B (step S221). The description generator 12 generates a description specifying the location where the still image data 111 was taken as a main road or a city name according to the latitude and longitude of the GPS information obtained from the vehicle driving GPS data 11B (step S222). This allows the generation of traffic situation descriptions that include information about the time period and location of the vehicle, thereby improving the accuracy of searches performed by users.
[0049] FIG. 16 is a diagram showing details of the process of generating a control explanation sentence (step S23). The explanation generator 12 generates an explanation regarding the driving speed control from the vehicle driving control data 11C (step S231). The explanation may be, for example, that a speed of 20 km / h or less is a low speed, 20 km to 60 km is a medium speed, and 60 km or more is a high speed. The explanatory sentence generator 12 generates explanatory sentences related to acceleration / deceleration control from the vehicle driving control data 11C (step S232). The explanatory sentences may be, for example, an accelerating state when the acceleration is equal to or greater than a certain level, a decelerating state when the deceleration is equal to or greater than a certain level, and a constant speed state otherwise. The explanation generator 12 generates an explanation about steering control from the vehicle driving control data 11C (step S233). The explanation may indicate, for example, that a steering angle of a certain value or more indicates a left turn or a right turn, and that other cases indicate a straight drive state. This allows for the generation of more natural traffic situation descriptions that include descriptions of the vehicle's control state and behavior, which has the effect of improving the accuracy of searches performed by users.
[0050] FIG. 17 is a table showing an example of the explanation generated in FIG. 15 and FIG. In line B01, an explanation of the time period is generated in step S221. In line B02, an explanation of the driving location is generated in step S222. In line B03, an explanation of the driving speed is generated in step S231. In line B04, an explanation of the acceleration / deceleration state is generated in step S232. In line B05, an explanation of the steering state is generated in step S233.
[0051] FIG. 18 is a table showing instruction sentences to the large-scale language model unit 13 used in the process of generating traffic condition explanation sentences (step S24). Line C01 describes an instruction to generate a traffic situation description. Line C02 describes the information to be included in the traffic situation description. Line C03 describes information to be excluded from the traffic situation description. Line C04 lists an example of a traffic situation description.
[0052] Then, in the traffic condition explanation generation process (step S24), the explanation generation unit 12 generates a prompt by combining the following texts (1) to (4) in order. Note that at least one of (2) and (3) may be omitted. (1) Instructions to the large-scale language model unit 13 (Figure 18) (2) GPS description (lines B01 and B02 in Figure 17) (3) Control description (lines B03, B04, and B05 in Figure 17) (4) Image description (Figure 14)
[0053] The description generation unit 12 acquires a traffic condition description written in natural language by inputting this generated prompt into the large-scale language model unit 13. Below, an example of a traffic condition description (corresponding to traffic condition description 732 in FIG. 22 described later) generated by the large-scale language model unit 13 from the prompt is shown. Traffic description = "You are driving straight ahead on Metropolitan Expressway Route 5 at noon, slowing down at high speed. In this scene, the highway has solid orange lane lines. Additionally, a white truck is driving ahead of you on the road ahead." This allows for the generation of natural traffic situation descriptions that match the intended use and the user's preferences, thereby improving the search accuracy for user search queries.
[0054] FIG. 19 is a table showing an example generated from an image different from the image caption in FIG. 19 is for an image taken from a vehicle traveling on a public road. Therefore, the combination of a public road and a pedestrian in the necessity table 14A is "needs explanation," so the information "pedestrian was not recognized" is written in row D07. The traffic condition description generated by the large-scale language model unit 13 from the prompt including the image description in FIG. 19 through the traffic condition description generation process (step S24) is as follows: Traffic description = "You are currently driving at a slow and steady speed on a city road in Mito city at night, preparing to turn left. There are no pedestrians in this scene."
[0055] The process of creating a database of traffic condition explanations (the process of storing the information in the explanation database 15) has been described above with reference to up to Fig. 19. An example of how the information created in the database can be used will be described below. FIG. 20 is a flowchart showing the processing of the search unit 16. The search unit 16 searches for images that match the input search statement from the images stored in the description database 15. Specifically, the search unit 16 outputs images that have a high similarity between the input search statement and the traffic condition description stored in the description database 15 as search results. The search unit 16 may, for example, list the images in the search results in order of similarity and output the top X images as "images with high similarity" (judgment of similarity based on relative ranking), or may output images in the search results whose similarity is higher than a predetermined reference value (threshold value Y) (judgment of absolute similarity). The processing of the search unit 16 will be described in detail below.
[0056] The search unit 16 receives a search statement input by a user from the input / output unit 17 (step S61). The user in this case may be, for example, a commentator at a traffic control center that manages expressways, and the search statement may be, for example, a request to collect images from a database of situations similar to an accident that occurred at a specific location on the expressway at a specific time. The commentator plans to edit the image materials obtained from the database to produce a news program about the accident that occurred. The search unit 16 evaluates the similarity between the search query from the user and the descriptions stored in the description database 15 (step S62), and acquires video information (image information) associated with the description with high similarity from the description database 15. The similarity evaluation may use a cosine similarity search based on document vectorization, or the like.
[0057] In addition to the driving video and traffic condition description acquired in step S62, the search unit 16 searches for and acquires related information about the driving conditions (such as weather information that is not in the description database 15 but can be acquired from the weather database by specifying the location and date of the traffic condition description), and generates an answer based on these results (step S63). In other words, the search unit 16 may output, as a search result, related information acquired from a database other than the description database 15 based on information contained in the traffic condition description corresponding to the image with high similarity. The search unit 16 transmits the reply sentence of step S63 to the input / output unit 17 (step S64). This allows the user to search for the desired video in a short time when searching for video data that has a traffic condition description similar to a search statement in natural language entered by the user.
[0058] FIG. 21 is a diagram showing the image search interface 71 of the input / output unit 17. As shown in FIG. The image search interface 71 is made up of a search query input section 72, which is an input field for the search query in step S61, and a search result display section 73, which is a display field for the answer sentence in step S64. The search result display section 73 displays multiple search results (scenes 1 to 4) as icons or thumbnail images.
[0059] Figure 22 shows the playback screen when the search result (scene 1) in Figure 21 is clicked. This playback screen displays the following information from top to bottom: ·731 images in search results. The traffic condition description 732 in the image 731 was extracted as having a high similarity to the search sentence. Additional information such as GPS description (time, location), control description (vehicle type, driving speed), and weather 733.
[0060] Figure 23 shows the playback screen when the search result (scene 2) in Figure 21 is clicked. As in Figure 22, this playback screen displays an image 741 as a search result, its traffic condition description 742, and additional information 743, just like in Figure 22. This allows users to input a search query in natural language, and then search for videos by referring to traffic condition descriptions, videos, and related information that are similar to the search query, thereby improving search efficiency.
[0061] According to the description generation process described above, when generating a description for a vehicle image, the description generation unit 12 refers to the necessity table 14A to generate a natural description that includes the presence or absence of peripheral objects that should be noted based on the traffic scene the vehicle is in. This allows a database to be created of natural descriptions that mention necessary peripheral objects but do not mention unnecessary peripheral objects, thereby improving the accuracy of database searches using search queries entered by humans.
[0062] Step S12: Details of search process Next, a detailed description will be given of the search process performed by the scene designation unit 1001. The search process is a process for extracting an image description relating to a specific scene from the traffic condition description generated in step S11.
[0063] As described above, the scene designation unit 1001 includes the second generation AI 1102. Therefore, even if the worker is not an experienced worker, a specific scene can be extracted through the dialogue between the worker and the second generation AI 1102 in Fig. 24. The flowchart in Fig. 24 will be described below with reference to Fig. 1 as needed.
[0064] In step S71, the scene designation unit 1001 accepts a prompt input by the operator. The prompt may be, for example, "Please tell us about incidents that occurred during driving between 10:00 AM and 8:00 PM."
[0065] In step S72, the scene specification unit 1001 provides a prompt to the second generation AI 1102, causing the second generation AI 1102 to access the explanation sentence database 15. The second generation AI 1102 generates a response sentence to the prompt (step S73). An example of the response is, "The incident that occurred while driving between 10:00 a.m. and 8:00 p.m. was a collision with a pedestrian."
[0066] In step S73, the scene designation unit 1001 accepts a prompt input by the operator. The prompt may be, for example, "Please search for traffic situation descriptions related to collisions with pedestrians."
[0067] In step S74, the scene specification unit 1001 provides a prompt to the second generation AI 1102, causing the second generation AI 1102 to access the explanation database 15 and extract the relevant traffic condition explanation 1006. When the processing of step S74 is completed, the processing of Figure 24 ends.
[0068] The prompts input by the worker include, for example, the vehicle location, such as "highway," "general road," or "parking lot," as shown in FIG. 6. The prompts input by the worker also include, for example, incidents, such as "sudden braking," "collision with another vehicle," "lane departure," or "failure of equipment within the vehicle." The prompts input by the worker also include times, such as "10:00 AM to 8:00 PM" or "3:00 PM to 3:10 PM."
[0069] Furthermore, the present invention does not necessarily require the second generation AI 1102. The scene specification unit 1001 includes an interface for the operator to specify a predetermined scene, and can directly access the description database 15 based on the natural language input by the operator, similar to step S11, without going through the second generation AI 1102, and extract the related traffic condition description 1006.
[0070] Step S13: Details of condition setting process Next, details of the condition setting process by the summarization condition setting unit 1003 will be described with reference to Fig. 1. The condition setting process is a process for specifying traffic condition explanation sentences 1006 that meet the conditions from the search results of step S12 and specifying explanation subjects to be mentioned in the report 1005 in order to generate the report 1005. The summarization condition setting section 1003 stores, for example, two conditions.
[0071] The first condition is a period that includes a scene. The period that includes a scene means, for example, from time (-T2) to time T1 in FIG. 25. Depending on the scene, information from time T0 onward may be important, or information from time T0 onward may be important. Therefore, time T1 and time (-T2) may take the same value, or they may take different values.
[0072] FIG. 26 is a diagram illustrating a table in which periods including scenes are recorded. 26, for example, the extent to which the period including a scene includes the period before the scene and the period after the scene can be set for each scene and stored as a table 2401 in the summary condition setting unit 1003. Note that it is also within the scope of the disclosure of this embodiment to create a report 1005 (described later) using only either time T1 or time (-T2).
[0073] Note that scenes include both incidents and non-incidents. Examples of incidents include collisions between the vehicle and a pedestrian, collisions between the vehicle and another vehicle, and collisions between the vehicle and a structure. Examples of non-incidents include sudden braking, lane departure, vehicle breakdown, violations of road traffic laws, and speeding.
[0074] FIG. 27 is a diagram for explaining a table in which explanation objects are set for each scene. The second condition is an explanation object. An explanation object is an object mentioned in the report 1005. As shown in FIG. 27 , for example, explanation objects can be set for each scene and stored in the summary condition setting unit 1003 as a table 2501.
[0075] The first and second conditions can be set in advance by the worker. As will be described later, the report 1005 is created according to the conditions specified in step S13. Therefore, the conditions in step S13 function as a prompt to the third generation AI 1103. In other words, the prompt to the third generation AI 1103 is instructed to "Please include buildings, intersections, pedestrians, and crosswalks as the objects to be described in the report." By storing standard prompt information in a table in advance in this way, it is possible to give instructions to the generation AI without error.
[0076] Step S14: Details of report creation process Next, we will explain the details of the report creation process by the report creation device 1004. In step S14, the third generation AI 1103 in the report creation device 1004 creates a report 1005 that is a summary of the traffic condition explanation 1006 that meets the conditions in step S13.
[0077] For example, if the worker selects a pedestrian collision in step S12, the third generation AI 1103 generates a report 1005. The generated report 1005 is saved in the explanatory sentence database 15 in a format that the worker can search in natural language. The report may look something like the following, for example: "The video begins with driving down a city street, capturing a driving scenario in a typical urban environment. As it progresses, new high-rise buildings become prominent along the route, suggesting that you are in a developed area of the city. As you approach an intersection near this landmark, it is a common location for increased traffic and pedestrian activity. Shortly after, while driving near a crosswalk, a pedestrian appears on your right, indicating a potential collision or distraction point. The situation quickly escalates, and the next moment, you see a person in the windshield of your car, a sudden and startling scene suggesting a collision with a pedestrian. In the final scene, you are driving through the crosswalk and a person is seen lying in the street, indicating a serious accident likely resulting from your previous interaction with the pedestrian. Overall, the video captures the progression from normal city driving to a serious incident involving a pedestrian, suggesting a serious accident at a crosswalk."
[0078] As described above, the explanation generator 12 generates a control explanation based on the vehicle driving control data 11C as a control explanation generation process (step S23). The explanation generator 12 can also add an explanation about vehicle control to the traffic situation explanation 1006 for the image based on the vehicle driving control data read from an on-board ECU (Electronic Control Unit) of the vehicle 91. Therefore, the information about the on-board ECU can also be reflected in the report 1005.
[0079] The information of the in-vehicle ECU includes, for example, information about sensors such as cameras. The information about the sensors includes information suggesting whether a program for detecting objects such as pedestrians by the sensors such as cameras was running, and, if the program was running, information suggesting whether an object such as a pedestrian was detected. Furthermore, the information of the in-vehicle ECU includes, for example, information about vehicle behavior. The information about vehicle behavior includes information about speed, acceleration / deceleration, acceleration, jerk information which is the time derivative of acceleration, yaw rate, brake hydraulic pressure, and / or steering angle.
[0080] For example, if report 1005 states that a program for detecting objects such as pedestrians was not running, the developer can improve the program so that it will start in similar scenes. Also, if report 1005 states that the program was running but did not detect an object, the developer can improve the program so that it can detect objects. Also, if report 1005 states that the object was detected, the developer can know that there is room for improvement in vehicle control.
[0081] In addition, similar to the image description, GPS description, and control description, it is also within the scope of the disclosure of this embodiment that the description generation unit 12 creates the program launch status description 2801 in Figure 28 and the object detection status description 2901 in Figure 29 and inputs them as one of the prompts to the third generation AI 1103 so that the third generation AI 1103 can more reliably refer to the program launch status and / or object detection status.
[0082] 28 and the object detection status description 2901 in Fig. 29, it is also within the scope of the disclosure of this embodiment to collect the program startup status and / or object detection status from at least one of the vehicles 91 to 93 in association with time information and store them in the driving log database 11. Furthermore, it is also within the scope of the disclosure of this embodiment to output the program startup status description 2801 and / or the object detection status description 2901 in association with an image as one of the additional information 733, 743.
[0083] According to this embodiment, it is possible to improve the efficiency of the work of creating the report 1005. By making it possible to specify the time before and after and / or the subject of explanation, it is possible to avoid creating a redundant report 1005. The present invention is particularly effective in the development of autonomous driving vehicles.
[0084] The configuration and effects of the present invention will be described below.
[0085] [1] An image analysis unit (vehicle image analysis device 1) that generates a situation description from an image captured by a camera using a generation AI (artificial intelligence) (first generation AI 1101) and stores the generated description in a database; a scene designation unit (1001) for designating an explanation target scene that is the subject of the explanation; a scene search unit (1002) for searching the situation description sentences stored in the database for a situation description sentence related to the scene to be explained designated by the scene designation unit (1001); a summary condition setting unit (1003) for saving a time and an object to be explained according to the scene to be explained designated by the scene designation unit (1001); a report creation unit (report creation device 1004) that uses a situation description corresponding to the time and the description object stored by the summary condition setting unit (1003) among the situation descriptions searched by the scene search unit (1002) to create a report (1005) about the description object by a generation AI (third generation AI 1103); A text generation system (task support system 1000) comprising:
[0086] This allows for the efficient creation of reports based on images.
[0087] [2] The scene designation unit (1001) designates an explanation target scene that is the subject of the explanation by a generation AI (second generation AI 1102). 2. The text generation system according to claim 1,
[0088] By using generative AI, it is easy to specify the scene to be explained.
[0089] [3] The camera is an in-vehicle camera. 2. The text generation system according to claim 1,
[0090] This makes it possible to create reports in the development of autonomous vehicles in an ideal manner.
[0091] [4] The report (1005) includes information about the vehicle ECU. 4. The text generation system according to claim 3.
[0092] This makes it possible to create reports in the development of autonomous vehicles in an ideal manner.
[0093] [5] The information about the in-vehicle ECU includes information about sensors connected to the in-vehicle ECU and information about vehicle behavior. 5. The text generation system according to claim 4.
[0094] This makes it possible to create reports in the development of autonomous vehicles in an ideal manner.
[0095] [6] A procedure for generating a situation description using a generation AI (first generation AI 1101) from an image captured by a camera; storing the situation description in a database; a procedure for specifying an explanation target scene to be explained; a step of searching for a situation description related to the specified scene to be described from the situation descriptions stored in the database; a step of saving a time and an explanation object according to the specified explanation object scene; a step of generating a report on the object of explanation using a generation AI (third generation AI 1103) based on a situation description sentence corresponding to the time and object of explanation according to the scene to be explained, among the situation descriptions found; A text generation program for a computer to execute the above.
[0096] This allows for the efficient creation of reports based on images.
[0097] [7] A step of generating a situation description sentence using a generation AI (first generation AI 1101) from an image captured by a camera; storing the situation description in a database; A step in which a scene designation unit (1001) designates an explanation target scene to be explained; A scene search unit (1002) searches for a situation description related to the scene to be explained, which is designated by the scene designation unit (1001), from the situation description stored in the database; A step in which a summary condition setting unit (1003) stores the time and the object to be explained according to the specified scene to be explained; generating a report (1005) about the object of explanation using a generation AI (third generation AI 1103) based on a situation description sentence corresponding to the time and object of explanation stored by the summary condition setting unit (1003) among the situation descriptions searched by the scene search unit (1002); generating a database (1007) containing said report (1005); 1. A method for producing a report database, comprising:
[0098] This allows for the efficient creation of reports based on images.
[0099] The present invention can be applied to applications other than the development of autonomous vehicles. For example, the present invention can be applied to the following fields: (1) In the field of logistics, the present invention can be applied to creating daily reports for drivers. (2) In the field of education, the present invention can be applied to summarizing related content based on keywords from online lectures, seminars, and academic conferences, as well as summarizing the content of lecture questions.
[0100] (3) In the field of crime prevention, the present invention can be applied to reports summarizing security camera footage, which can be used to detect, for example, vehicle theft, snatching, shoplifting, and suspicious individuals during late-night hours.
[0101] (4) In the retail sector, the present invention can be applied to the creation of reports using anti-theft camera data. The reports can be used to search for theft scenes. The present invention can also be applied to the creation of reports using camera data from unmanned stores and self-checkout registers.
[0102] (5) In the broadcasting field, the present invention can be applied to summarizing the video before and after sporting events and program highlights. (6) In the medical field, this invention can be applied to summarizing surgical videos. The created report functions as a textbook. Therefore, by referring to the report, it is possible to confirm the procedure. It is also possible to search for the desired procedure. As a result, training can be made more efficient.
[0103] (7) In the fields of factory, construction site, and warehouse management, this invention can be applied to the creation of accident reports. Accidents can include objects falling or tipping over. Reports can be created from videos of the site. From the reports, it becomes possible to search for signs of malfunction. This invention can also be applied to the output of work reports with videos, the creation of inspection and maintenance reports, and the creation of procedure manuals that describe maintenance procedures.
[0104] (8) In the field of manufacturing, the present invention makes it possible to summarize the skills of skilled craftsmen and other skilled technicians. (9) In the field of building and elevator management, this invention makes it possible to create a report summarizing the number of people entering and exiting each time period. The report can be used for persona analysis.
[0105] (10) In the field of home appliances, the present invention can also create a consumption report from consumption records obtained from a refrigerator equipped with a sensor such as a camera. (11) The present invention can also be applied to summarizing cooking videos, summarizing product and service introduction videos, and summarizing the behavioral history of pets and children.
[0106] Furthermore, the present invention is not limited to the above-described embodiments, and various other applications and modifications are possible without departing from the spirit of the present invention as defined in the claims. For example, the above-described embodiments provide a detailed and specific description of the configuration of the vehicle image analysis device 1 in order to clearly explain the present invention, and the present invention is not necessarily limited to a system including all of the components described. Furthermore, it is possible to replace part of the configuration of one embodiment with a component of another embodiment. It is also possible to add a component of another embodiment to the configuration of one embodiment. It is also possible to add, replace, or delete other components from part of the configuration of each embodiment.
[0107] Furthermore, the above-described configurations, functions, processing units, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. As the hardware, a broad processor device such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit) may be used. Furthermore, each component of the vehicle image analyzing device 1 according to the above-described embodiment may be implemented in any hardware as long as the respective hardware can transmit and receive information to and from each other via a network. Furthermore, the processing executed by a certain processing unit may be realized by a single piece of hardware, or may be realized by distributed processing using multiple pieces of hardware. [Explanation of symbols]
[0108] 1. Vehicle image analysis device (image analysis unit) 8. Communication Lines 11 Driving log database 11A Vehicle driving video data 11B Vehicle driving GPS data 11C Vehicle driving control data 12 Description generation section 13 Large-scale language model section 13A VQA section 13B Summary generator 14 Explanation target setting section 14A Necessity Table 14B Recognition Table 15 Description database (image database) 16 Search section 17 Input / output section 91-93 vehicles 100 Image Description System 111 Still image data 732,742 Situation Description 901 CPU 902 RAM 903 ROM 904 HDD 905 Communication I / F 906 Input / Output Interface 907 Media I / F 1000 Work Support System 1001 Scene specification section 1002 Scene Search Unit 1003 Summary condition setting section 1004 Report writing device 1005 Report 1006 Traffic Condition Description 1101 First Generation AI 1102 Second Generation AI 1103 Third Generation AI
Claims
1. an image analysis unit that uses artificial intelligence (AI) to generate a situation description from an image captured by the camera and stores the description in a database; a scene designation unit that designates an explanation target scene that is the subject of the explanation; a scene search unit that searches the situation description sentences stored in the database for a situation description sentence related to the scene to be explained that is designated by the scene designation unit; a summary condition setting unit that stores a period and an object to be explained according to the scene to be explained designated by the scene designation unit; a report creation unit that uses a situation description sentence corresponding to the time and the description object stored by the summary condition setting unit among the situation descriptions searched by the scene search unit to create a report on the description object by a generation AI; A sentence generation system comprising:
2. The scene designation unit designates an explanation target scene that is a target of the explanation by the generation AI.
2. The text generation system according to claim 1.
3. The camera is an in-vehicle camera.
2. The text generation system according to claim 1.
4. The report includes information about the vehicle ECU.
4. The text generation system according to claim 3.
5. The information about the in-vehicle ECU includes information about sensors connected to the in-vehicle ECU and information about vehicle behavior.
5. The text generation system according to claim 4.
6. A procedure for generating a situation description using AI from an image taken by a camera; storing the situation description in a database; a procedure for specifying an explanation target scene to be explained; a step of searching for a situation description related to the specified scene to be described from the situation descriptions stored in the database; a step of saving a time and an explanation object according to the specified explanation object scene; a step of generating a report on the object of explanation using a generation AI based on a situation description sentence corresponding to a time and an object of explanation according to the scene to be explained, among the retrieved situation descriptions; A text generation program for a computer to execute the above.
7. A step of generating a situation description using a generation AI from an image taken by a camera; storing the situation description in a database; a step in which a scene designation unit designates an explanation target scene to be explained; a step in which a scene search unit searches for a situation description related to the scene to be explained, which is designated by the scene designation unit, from the situation description stored in the database; a step of a summary condition setting unit saving a period and an object to be explained according to the specified scene to be explained; generating a report on the object of explanation using a generation AI based on the situation description sentences retrieved by the scene search unit that correspond to the period and object of explanation stored by the summary condition setting unit; generating a database containing the report; 1. A method for producing a report database, comprising:
Citation Information
Patent Citations
Information processing device, information processing system, information processing method and program
JP2023127538A
Electric Tractor
US20230132970A1
Video image summary instrument, explanation note forming instrument, video image summary method, explanation note forming method and program
JP2005109566A
Cited By
Interactive time series analysis system and method of the same
JP2025162971A
Interactive time series analysis system and method
JP7844695B2