Learning data generation system, learning data generation program, and learning data generation method

The training data generation system addresses the lack of domain-specific data for autonomous driving by automating the generation of accurate explanatory text and videos, enhancing model accuracy through iterative evaluation and training, thus reducing manual effort and time.

JP2025150916AActive Publication Date: 2025-10-09SOFTBANK CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024052079
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-10-09
Estimated Expiration
2044-03-27

AI Technical Summary

Technical Problem

Existing AI models for autonomous driving lack sufficient domain-specific training data and accurate text descriptions of vehicle driving conditions, necessitating manual generation of large amounts of data and text, which is time-consuming and inefficient.

Method used

A training data generation system that includes an acquisition unit for explanatory text from a language model, an evaluation unit for accuracy, a first generation unit for training sentences, and a second generation unit for training videos, automatically generating specialized training data for remote monitoring by evaluating and improving the language model's accuracy element by element.

Benefits of technology

Automatically generates domain-specific training data, enhancing the language model's accuracy in describing vehicle conditions, reducing the need for manual data creation and improving the model's performance in remote monitoring tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025150916000001_ABST
    Figure 2025150916000001_ABST
Patent Text Reader

Abstract

To facilitate learning of a language model configured to generate a descriptive text that describes the situation of a vehicle based on a video captured by a camera mounted on the vehicle.SOLUTION: A learning data generation system (1) includes: an acquisition unit (11) which acquires a descriptive text from a language model (2) which receives a video captured by a camera mounted on a vehicle to output a descriptive text in a predetermined format that describes the situation of the vehicle; an evaluation unit (12) which evaluates estimation accuracy of the language model for each multiple elements constituting the descriptive text; a first generation unit (13) which generates a learning text to be learned by the language model, based on a result of evaluating the estimation accuracy; and a second generation unit (14) which generates a learning video indicating the contents of the learning text, the learning video being to be learned by the language model together with the learning text, based on the learning text.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a training data generation system, a training data generation program, and a training data generation method. [Background technology]

[0002] In recent years, with the lifting of the ban on "Level 4" autonomous driving, demonstration experiments of autonomous vehicles have been conducted across the country. In autonomous driving, a huge amount of data is collected from the vehicle, and this data is used by the vehicle's AI to recognize events, predict future events, plan trips, and improve (learn) the AI. For example, Non-Patent Document 1 below discloses the current status of demonstration experiments of such autonomous vehicles. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] “Autonomous driving demonstrations are becoming more active both domestically and internationally! What are the challenges in data management?” [online], October 15, 2020, Autonomous Driving Lab, [Retrieved July 7, 2023], Internet<URL:https: / / jidounten-lab.com / u_data-management-1> Summary of the Invention [Problem to be solved by the invention]

[0004] AI (trained models) are generally constructed using large amounts of training data available on the internet. As a result, trained models specialized for specific domains, such as remote monitoring of autonomous driving, have not yet been constructed to the point where they can demonstrate sufficient accuracy (they suffer from problems such as hallucination). These problems can be solved by preparing domain-specific training data. However, there is not much training data available in the world that is specialized for remote monitoring of autonomous driving, and it still needs to be created manually. The amount of training data required for training is enormous, and creating it requires a great deal of time and effort.

[0005] Furthermore, building a trained model requires not only video but also text that explains the content of the video. While language models that generate text from video have existed up until now, building a trained model specialized for remote monitoring of autonomous driving requires text that accurately describes the driving conditions of an autonomous vehicle obtained from the vehicle's onboard camera. However, currently, there is no established technology that can generate text that accurately describes the vehicle's driving conditions. [Means for solving the problem]

[0006] In order to solve the above problem, a training data generation system according to one aspect of the present invention includes: an acquisition unit that acquires explanatory text from a language model that outputs explanatory text in a predetermined format that explains the situation of the vehicle when a captured video image captured by a camera mounted on the vehicle is input; and an evaluation unit that evaluates the estimation accuracy of the language model for each of a plurality of elements that make up the explanatory text. The system includes a first generation unit that generates training sentences for training the language model based on the evaluation results of the estimation accuracy, and a second generation unit that generates training videos showing the content of the training sentences based on the training sentences, for training the language model together with the training sentences.

[0007] In addition, a training data generation program according to another aspect of the present invention causes a computer to execute the following processes: an acquisition process for acquiring explanatory sentences from a language model that outputs explanatory sentences in a predetermined format explaining the status of the vehicle when video captured by a camera mounted on a vehicle is input; an evaluation process for evaluating the estimation accuracy of the language model for each of a plurality of elements constituting the explanatory sentences; a first generation process for generating training sentences for training the language model based on the evaluation results of the estimation accuracy; and a second generation process for generating training video showing the content of the training sentences based on the training sentences, for training the language model together with the training sentences. Note that a computer-readable recording medium having the training data generation program recorded thereon is also within the scope of the present invention.

[0008] Furthermore, a training data generation method according to another aspect of the present invention includes an acquisition step in which a computer acquires explanatory sentences from a language model that outputs explanatory sentences in a predetermined format that explain the situation of the vehicle when video footage captured by a camera mounted on the vehicle is input; an evaluation step in which the computer evaluates the estimation accuracy of the language model for each of a plurality of elements that make up the explanatory sentences; a first generation step in which the computer generates training sentences for training the language model based on the evaluation results of the estimation accuracy; and a second generation step in which the computer generates training video that shows the content of the training sentences, based on the training sentences, to be trained by the language model together with the training sentences. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram illustrating an example of the functional configuration of a training data generation system and a model training system according to an embodiment of the present invention. [Figure 2] 10 is a diagram illustrating a format of explanatory text generated by a language model included in the model learning system according to the embodiment. FIG. [Figure 3] FIG. 2 is a diagram showing an example of a captured video to be input to the language model. [Figure 4]10 is a table showing an example of an evaluation result of the estimation accuracy of a language model by the training data generation system according to the embodiment. [Figure 5] 10A and 10B are diagrams illustrating how a learning sentence is generated by the learning data generation system according to the embodiment. [Figure 6] 10A to 10C are diagrams illustrating how a learning video is generated by the learning data generation system according to the embodiment. [Figure 7] 10 is a table showing the contents of learning data (data set) accumulated by the model learning system according to the embodiment. [Figure 8] 1 is a flowchart illustrating an example of the flow of a training data generation method and a model training method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0010] <Learning Data Generation System 1 (Model Learning System 100)> Hereinafter, a training data generation device according to an embodiment of the present invention will be described in detail.

[0011] [Configuration of Model Learning System 100 (Model Learning System 100)] 1, the training data generation system 1 according to this embodiment constitutes a model training system 100. In addition to the training data generation system 1, the model training system 100 includes a language model 2, a storage unit 3, and a tuning unit 4.

[0012] [Language Model 2] Language model 2 is a trained model constructed by machine learning using a set of video and sentences that explain the content of the video as training data. Therefore, when video captured by a camera mounted on a vehicle is input, language model 2 outputs explanatory sentences in a predetermined format that explain the situation of the vehicle. The captured video may be of the exterior or interior of the vehicle.

[0013] The predetermined format means that the explanatory text contains essential elements. The "elements" include at least the vehicle's driving status. Specifically, as shown in FIG. 2, the "elements" include at least the vehicle's position and the vehicle's driving intention. The "elements" further include at least one of the following: an object shown in the captured video, the object's position, the object's motion, and the vehicle's current motion. The "object's position" may further include a relative position with respect to the vehicle and an absolute position.

[0014] When a single captured video is input, the language model 2 according to this embodiment may output multiple types of explanatory sentences, each of which explains a different type of driving situation. For example, when a captured video in which a car and a bicycle (or a pedestrian) are simultaneously captured, as shown in FIG. 3, is input, the language model 2 generates explanatory sentences about the car and about the bicycle. The language model 2 may be provided in the vehicle or in the training data generation system 1. The language model 2 may also be provided in another device independent of the vehicle and the training data generation system 1.

[0015] [Learning Data Generation System 1] The training data generation system 1 according to this embodiment includes an acquisition unit 11, an evaluation unit 12, a first generation unit 13, and a second generation unit 14. The training data generation system 1 according to this embodiment further includes a selection unit 15, a second language model 16, a second storage unit 17, and a third storage unit 18.

[0016] (Second Language Model 16) The second language model 16 is a trained model different from the language model 2. The second language model is a large language model (LLM). Specifically, the second language model 16 is configured with at least one of "BestScore", "Sentence Bert", and "GPT4". Note that the second language model 16 may be a trained model other than a large language model.

[0017] (Second storage unit 17) The second storage unit 17 stores a correct sentence to be used by the evaluation unit 12, which will be described later. The second storage unit 17 according to this embodiment is configured by, for example, a semiconductor memory, a hard disk, or the like.

[0018] (Third memory section 18) The third storage unit 18 stores template images used by the second generation unit 14, which will be described later. The third storage unit 18 according to this embodiment is configured with, for example, a semiconductor memory, a hard disk, etc. The third storage unit 18 may be configured integrally with the second storage unit 17.

[0019] (Acquisition part 11) The acquisition unit 11 executes an acquisition process. In the acquisition process, the acquisition unit 11 acquires explanatory sentences from a language model. As described above, the language model 2 according to this embodiment outputs multiple types of explanatory sentences when one captured video is input. Therefore, the acquisition unit 11 according to this embodiment acquires multiple types of explanatory sentences from the language model 2.

[0020] (Selection section 15) The selection unit 15 executes a selection process. In the selection process, the selection unit 15 selects, from among multiple types of explanatory sentences, the explanatory sentence whose content will have the greatest impact on the vehicle's future driving. The selection of explanatory sentences may be performed using a trained model, or may be performed by quantifying the content of each explanatory sentence and then based on the numerical values. Furthermore, if the language model 2 generates only one explanatory sentence for each captured video, the training data generation system 1 may not be equipped with the selection unit 15.

[0021] (Evaluation Section 12) The evaluation unit 12 executes an evaluation process. In the evaluation process, the evaluation unit 12 evaluates the estimation accuracy of the language model for each of the multiple elements constituting the explanatory text. As described above, the training data generation system 1 according to this embodiment includes a selection unit 15. Therefore, the evaluation unit 12 according to this embodiment evaluates the estimation accuracy of the language model using the explanatory text selected by the selection unit 15. The evaluation unit 12 according to this embodiment calculates the accuracy rate for each element by comparing the explanatory text with the correct answer sentence using the second language model 16. As a result, as shown in FIG. 4, the language model 2 is evaluated in such a way that, for a certain element (e.g., the position of a vehicle), for a captured image showing an intersection, the language model 2 outputs an explanatory text indicating that the vehicle is at an intersection with a 30% accuracy rate, and for a captured image showing a straight road, the language model 2 outputs an explanatory text indicating that the vehicle is on a straight road with a 100% accuracy rate. Note that the evaluation unit 12 may be configured to check whether the explanatory text is in a natural English format before evaluating the explanatory text.

[0022] (First generation section 13) The first generation unit 13 executes a first generation process. In the first generation process, the first generation unit 13 generates training sentences for training a language model based on the evaluation result of the estimation accuracy. As shown in FIG. 5, the first generation unit 13 according to this embodiment generates training sentences in which elements with low accuracy rates are weighted. That is, the training data generation system 1 generates training data (training video and training sentence pairs: teacher data) that improves the accuracy of elements with low estimation accuracy in the language model. Furthermore, the first generation unit 13 according to this embodiment generates training sentences in the same format as the explanatory sentences. Note that the first generation unit 13 may be configured to generate training data that improves the accuracy of all elements or training data that further improves the accuracy of elements with high estimation accuracy.

[0023] (Second generation unit 14) The second generation unit 14 executes a second generation process. In the second generation process, the second generation unit 14 generates a training video based on the training sentence generated by the first generation unit 13. The training video is a video showing the content of the training sentence, which is used to train a language model together with the training sentence. The second generation unit 14 according to the present embodiment generates the training video using template videos prepared in advance (stored in the third storage unit 18). Each template video has content associated with each element of the explanatory text. For example, the template video includes a video of a road (straight road, crossroads, T-junction, etc.) at the vehicle position indicated by the explanatory text. For example, when generating a training video of a pedestrian crossing a straight road while a vehicle is traveling straight on the straight road, the second generation unit 14 acquires a template video of a straight road from template videos of multiple roads linked to a road map, as shown in FIG. 6. The second generation unit 14 also acquires a template video of a pedestrian crossing a road from template videos of pedestrians. Then, the second generation unit 14 generates a learning video by combining the acquired template video of a pedestrian with the acquired template video of a straight road. In this way, the second generation unit 14 can generate learning videos more easily than in the past (compared to creating videos from scratch). Note that the template video may include something that is not associated with an element (for example, something that is unrelated to the vehicle position indicated by the explanatory text).

[0024] [Storage section 3] The storage unit 3 is capable of storing the learning data (training video and learning sentence pairs: teacher data) generated by the learning data generation system 1. As a result, as the learning data generation system 1 repeatedly generates learning data, the storage unit 3 accumulates multiple pairs of learning sentences in various situations (different elements) and learning videos linked to the learning sentences to be written, as shown in FIG. 7. The storage unit 3 according to this embodiment is configured, for example, with a semiconductor memory, a hard disk, etc. Note that the storage unit 3 may be provided in the learning data generation system 1.

[0025] [Tuning section 4] The tuning unit 4 fine-tunes the language model using the training data stored in the storage unit 3. The tuning unit 4 may be provided in the vehicle or in the training data generation system 1. Alternatively, the tuning unit 4 may be provided in another device independent of the vehicle and the training data generation system 1.

[0026] [Effects of the learning data generation system 1 (model learning system 100)] The training data generation system 1 described above automatically generates training data (training video and training text: training data) specialized for a specific domain, namely remote monitoring, by using a method of evaluating explanatory text in a predetermined format element by element. Therefore, the training data generation system 1 automatically trains a language model 2 using newly generated training data, creating a loop (model training system 100) in which the accuracy of sentence generation automatically improves. This eliminates the need to manually generate vast amounts of training data. In other words, the training data generation system 1 makes it easy to train a language model.

[0027] Furthermore, according to the model learning system 100, the estimation accuracy of the language model (the accuracy of the content of the generated explanatory text) gradually improves.

[0028] <Learning data generation method S1 (model learning method S100)> Hereinafter, a training data generation method according to another embodiment of the present invention will be described in detail.

[0029] [Model learning method S100 flow] 8, the training data generation method S1 according to this embodiment forms part of a model training method S100. In addition to the training data generation method S1, the model training method S100 includes a tuning step S2.

[0030] [Flow of training data generation method S1] The training data generation method S1 includes an acquisition step S11, an evaluation step S12, a first generation step S13, and a second generation step S14. The training data generation method S1 according to this embodiment further includes a selection step S15.

[0031] (Acquisition step S11) In the first acquisition step S11, explanatory sentences are acquired from the language model 2. The explanatory sentences may be acquired using the training data generation system 1 described above, or may be acquired using other systems or devices. As described above, the language model 2 according to this embodiment outputs multiple types of explanatory sentences when a single captured video is input. Therefore, in the acquisition step S11 according to this embodiment, multiple types of explanatory sentences are acquired from the language model 2.

[0032] (Selection step S15) After acquiring the explanatory sentence, the process proceeds to a selection step. In the selection step S15, the explanatory sentence with the greatest impact on the vehicle's future driving is selected from among the multiple types of explanatory sentences. The selection unit 15 according to this embodiment performs the selection using a trained model. The selection of the explanatory sentence may be performed using the training data generation system 1 described above, or may be performed using another system or device. Note that if the language model 2 generates only one explanatory sentence for one captured video, the training data generation method S1 may not include the selection step S15.

[0033] (Evaluation step S12) After the explanatory text is acquired or selected, the process proceeds to an evaluation step. In the evaluation step S12, the estimation accuracy of the language model 2 is evaluated for each of the multiple elements that make up the explanatory text. The estimation accuracy may be evaluated using the training data generation system 1 described above, or may be evaluated using another system or device.

[0034] (First generation step S13) After evaluating the estimation accuracy of the language model 2, the process proceeds to a first generation step S13. In the first generation step S13, training sentences for training the language model are generated based on the evaluation results of the estimation accuracy. The training sentences may be generated using the training data generation system 1 described above, or may be generated using another system or device.

[0035] (Second generation step S14) After the training sentences are generated, the process proceeds to the second generation step S14. In the second generation step S14, training videos are generated based on the training sentences. The training videos may be generated using the training data generation system 1 described above, or may be generated using another system or device.

[0036] (Tuning step S2) After generating the training data (a set of training sentences and training videos), the process proceeds to tuning step S2. In tuning step S2, the training data is used to fine-tune the language model 2. The fine-tuning of the language model 2 may be performed using the model training system 100 described above, or may be performed using another system or device.

[0037] [Effects of learning data generation method S1 and model learning method S100] The training data generation method S1 described above automatically generates training data specialized for a specific domain, remote monitoring, by evaluating explanatory text in a predetermined format element by element. Therefore, the training data generation method S1 can automatically train a language model using newly generated training data, creating a loop in which the accuracy of sentence generation automatically improves. As a result, it is possible to eliminate the need to manually generate a huge amount of training data. In other words, the training data generation method S1 makes it easy to train a language model.

[0038] Furthermore, according to the model learning method S100, the estimation accuracy of the language model (the accuracy of the content of the generated explanatory text) gradually improves.

[0039] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

[0040] For example, the storage unit 3 may be included in the training data generation system 1.

[0041] For example, each unit of the learning data generation system 1 is a program for causing a computer to function as each unit, and can be realized by a learning data generation program for causing a computer to function as each unit. In this case, each unit includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the learning data generation program. Each unit is realized by executing the learning data generation program using the control device and storage device. The training data generation program may be stored non-temporarily on one or more computer-readable recording media. Each unit may or may not have a recording medium. In the latter case, the program may be supplied to each unit via any wired or wireless transmission medium. In addition, some or all of the functions of each unit can be realized by logic circuits. For example, integrated circuits in which logic circuits functioning as each unit are formed are also included in the scope of the present invention. In addition, the functions of each unit can also be realized by, for example, a quantum computer.

[0042] 〔summary〕 A training data generation system according to a first aspect of the present invention includes an acquisition unit that acquires explanatory text from a language model that outputs explanatory text in a predetermined format that explains the status of the vehicle when video captured by a camera mounted on the vehicle is input; an evaluation unit that evaluates the estimation accuracy of the language model for each of a plurality of elements that make up the explanatory text; a first generation unit that generates training text for training the language model based on the evaluation result of the estimation accuracy; and a second generation unit that generates training video showing the content of the training text, based on the training text, for training the language model together with the training text.

[0043] A learning data generation system according to a second aspect of the present invention may be configured in the above-mentioned first aspect, wherein the elements include at least a driving situation of the vehicle.

[0044] A learning data generation system according to aspect 3 of the present invention may be configured in the above-mentioned aspect 2 such that the elements include at least the position of the vehicle and the driving intention of the vehicle, and further include at least one of an object appearing in the captured video, the position of the object, the movement of the object, and the current movement of the vehicle.

[0045] The training data generation system according to aspect 4 of the present invention may be configured such that, in any of aspects 1 to 3 above, the evaluation unit calculates the accuracy rate for each element by comparing the explanatory sentence with a correct answer sentence using a second language model different from the language model.

[0046] A training data generation system according to aspect 5 of the present invention may be configured in the above-mentioned aspect 4 such that the first generation unit generates the training sentences in which the elements with low correct answer rates are weighted.

[0047] A training data generation system according to a sixth aspect of the present invention may be configured in the fifth aspect above, wherein the first generation unit generates the training sentences in the same format as the explanatory sentences.

[0048] A training data generation system according to aspect 7 of the present invention may be configured in any one of aspects 1 to 6 above, wherein the second generation unit generates the training video using a template video prepared in advance.

[0049] The training data generation system according to an eighth aspect of the present invention may be configured in the seventh aspect above, wherein the template video includes a video of a road at the position of the vehicle indicated by the explanatory text.

[0050] A learning data generation system according to a ninth aspect of the present invention may be configured in any one of the first to eighth aspects above, wherein the acquisition unit, when the captured video is input, acquires multiple types of explanatory sentences from a language model that outputs multiple types of explanatory sentences, each of which describes a different type of driving situation, and includes a selection unit that selects, from the multiple types of explanatory sentences, the explanatory sentence with the content that will have the greatest impact on the future driving of the vehicle, and the evaluation unit evaluates the estimation accuracy of the language model using the selected explanatory sentence.

[0051] A training data generation program according to aspect 10 of the present invention is configured to cause a computer to execute the following processes: an acquisition process for acquiring explanatory text from a language model that outputs explanatory text in a predetermined format that explains the situation of the vehicle when video footage captured by a camera mounted on a vehicle is input; an evaluation process for evaluating the estimation accuracy of the language model for each of the multiple elements that make up the explanatory text; a first generation process for generating training text for training the language model based on the evaluation results of the estimation accuracy; and a second generation process for generating training video showing the content of the training text based on the training text, for training the language model together with the training text.

[0052] A training data generation method according to aspect 11 of the present invention is a method including: an acquisition step in which a computer acquires explanatory text from a language model that, when video footage captured by a camera mounted on a vehicle is input, outputs explanatory text in a predetermined format that explains the situation of the vehicle; an evaluation step in which the computer evaluates the estimation accuracy of the language model for each of a plurality of elements that make up the explanatory text; a first generation step in which the computer generates training text for training the language model based on the evaluation result of the estimation accuracy; and a second generation step in which the computer generates training video showing the content of the training text, based on the training text, for training the language model together with the training text. [Explanation of symbols]

[0053] 100 Model Learning Systems 1. Training data generation system 11 Acquisition Department 12 Evaluation Section 13 First generation part 14 Second generator 15 Selection section 16 Second Language Model 17 Second memory section 18 Third Memory 2. Language Model 3 Storage section 4 Tuning section S100 model training method S1 Training data generation method S11 Acquisition step S12 Evaluation step S13 First generation step S14 Second generation step S15 Selection Step S2 Tuning Step

Claims

1. an acquisition unit that, when a captured image captured by a camera mounted on a vehicle is input, acquires the explanatory text from a language model that outputs explanatory text in a predetermined format that explains the situation of the vehicle; an evaluation unit that evaluates the estimation accuracy of the language model for each of a plurality of elements that constitute the explanatory text; a first generation unit that generates training sentences for training the language model based on the evaluation result of the estimation accuracy; a second generation unit that generates, based on the training sentences, training videos that indicate the contents of the training sentences, for training the language model together with the training sentences; Equipped with Training data generation system.

2. The elements include at least the driving status of the vehicle. The training data generation system according to claim 1 .

3. The element is The information includes at least the position of the vehicle and the driving intention of the vehicle, The information further includes at least one of an object appearing in the captured image, a position of the object, a movement of the object, and a current movement of the vehicle. The training data generation system according to claim 2 .

4. the evaluation unit calculates a correct answer rate for each of the elements by comparing the explanatory sentence with a correct answer sentence using a second language model different from the language model. The training data generation system according to claim 1 .

5. the first generation unit generates the learning sentences in which the elements with low correct answer rates are weighted. The training data generation system according to claim 4 .

6. the first generation unit generates the training text in the same format as the explanatory text; The training data generation system according to claim 5 .

7. the second generation unit generates the learning video by using a template video prepared in advance; The training data generation system according to claim 1 .

8. the template image includes an image of a road at the position of the vehicle indicated by the explanatory sentence; The training data generation system according to claim 7 .

9. when the captured video is input, the acquisition unit acquires a plurality of types of explanatory sentences from a language model that outputs a plurality of types of explanatory sentences each describing a different content of a driving situation, a selection unit that selects, from among the plurality of types of explanatory text, the explanatory text having the content that will have the greatest impact on future traveling of the vehicle; the evaluation unit evaluates the estimation accuracy of the language model using the selected explanatory sentence. The training data generation system according to claim 1 .

10. On the computer, an acquisition process for acquiring explanatory text from a language model that outputs explanatory text in a predetermined format that explains the situation of the vehicle when a captured image obtained by a camera mounted on the vehicle is input; an evaluation process for evaluating the estimation accuracy of the language model for each of a plurality of elements constituting the explanatory text; a first generation process for generating training sentences for training the language model based on the evaluation result of the estimation accuracy; a second generation process for generating, based on the training sentences, training videos showing the contents of the training sentences, for training the language model together with the training sentences; Execute Training data generation program.

11. an acquisition step in which, when a captured image captured by a camera mounted on a vehicle is input, a computer acquires the explanatory text from a language model that outputs explanatory text in a predetermined format that explains the situation of the vehicle; an evaluation step in which a computer evaluates estimation accuracy of the language model for each of a plurality of elements constituting the explanatory text; a first generation step in which the computer generates training sentences for training the language model based on the evaluation result of the estimation accuracy; a second generation step in which a computer generates, based on the training sentences, training videos showing the contents of the training sentences, for training the language model together with the training sentences; Including, Training data generation method.

Citation Information

Patent Citations

  • Explanation sentence creation device

    JP2021174172A

  • Sentence generator, program, and method for generating sentence

    JP2022116979A

  • System and program

    JP2025025521A