Training data generation system, training data generation program, and training data generation method

The learning data generation system addresses the lack of accurate training data for autonomous vehicle AI by automating the generation of explanatory text and videos, improving model accuracy for remote monitoring.

JP7850193B2Active Publication Date: 2026-04-22SOFTBANK CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SOFTBANK CORPORATION
Filing Date
2024-03-27
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing AI models for remote monitoring of autonomous vehicles lack sufficient accuracy due to insufficient domain-specific training data, and there is a lack of technology to generate accurate text descriptions of driving conditions from vehicle camera footage.

Method used

A learning data generation system that includes an acquisition unit to acquire explanatory text from a language model, an evaluation unit to assess the text's accuracy, a first generation unit to create training text, and a second generation unit to generate training videos, all aimed at improving the language model's accuracy.

Benefits of technology

Automatically generates domain-specific training data, reducing the effort required to create large amounts of training data and enhancing the language model's accuracy in describing driving conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850193000001
    Figure 0007850193000001
  • Figure 0007850193000002
    Figure 0007850193000002
  • Figure 0007850193000003
    Figure 0007850193000003
Patent Text Reader

Abstract

To facilitate learning of a language model configured to generate a descriptive text that describes the situation of a vehicle based on a video captured by a camera mounted on the vehicle.SOLUTION: A learning data generation system (1) includes: an acquisition unit (11) which acquires a descriptive text from a language model (2) which receives a video captured by a camera mounted on a vehicle to output a descriptive text in a predetermined format that describes the situation of the vehicle; an evaluation unit (12) which evaluates estimation accuracy of the language model for each multiple elements constituting the descriptive text; a first generation unit (13) which generates a learning text to be learned by the language model, based on a result of evaluating the estimation accuracy; and a second generation unit (14) which generates a learning video indicating the contents of the learning text, the learning video being to be learned by the language model together with the learning text, based on the learning text.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning data generation system, a learning data generation program, and a learning data generation method.

Background Art

[0002] In recent years, with the relaxation of restrictions on "Level 4" autonomous driving, demonstration experiments of autonomous vehicles have been conducted across the country. In autonomous driving, a huge amount of data is collected from vehicles and used for event recognition by the vehicle's AI, future prediction, driving planning, improvement (learning) of the AI, etc. For example, Non-Patent Document 1 below discloses the current situation of such demonstration experiments of autonomous vehicles.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] AI (trained models) are generally built using a large amount of training data available on the internet. Therefore, trained models specialized for specific domains, such as remote monitoring of autonomous vehicles, have not yet been developed to achieve sufficient accuracy (they suffer from problems such as hallucination). These problems can be solved by preparing domain-specific training data. However, there is not much training data specifically for remote monitoring of autonomous vehicles available, and it still needs to be created manually. The amount of training data required is enormous, and creating it requires a great deal of time and effort.

[0005] Furthermore, building a pre-trained model requires not only video footage but also text that describes what the video shows. While language models that generate text from video have existed before, building a pre-trained model specifically for remote monitoring of autonomous vehicles requires text that accurately describes the driving conditions of the autonomous vehicle, obtained from the vehicle's onboard camera. However, currently, technology for generating text that accurately describes the driving conditions of a vehicle has not been established. [Means for solving the problem]

[0006] To solve the above problems, a learning data generation system according to one aspect of the present invention includes an acquisition unit that, when captured video footage obtained from a camera mounted on a vehicle is input, acquires an explanatory text from a language model that outputs an explanatory text in a predetermined format that describes the situation of the vehicle, and an evaluation unit that evaluates the estimation accuracy of the language model for each of the multiple elements that constitute the explanatory text. The system includes a first generation unit that generates training text for the language model based on the evaluation results of the estimation accuracy, and a second generation unit that generates training video showing the content of the training text, which is to be trained on the language model together with the training text.

[0007] Furthermore, a learning data generation program according to another aspect of the present invention causes a computer to perform the following processes: an acquisition process to acquire explanatory text from a language model that outputs explanatory text in a predetermined format describing the situation of a vehicle when captured video obtained from a camera mounted on the vehicle is input; an evaluation process to evaluate the estimation accuracy of the language model for each of the multiple elements constituting the explanatory text; a first generation process to generate learning text for the language model to learn based on the evaluation result of the estimation accuracy; and a second generation process to generate learning video showing the content of the learning text, to be learned by the language model together with the learning text. Note that a computer-readable recording medium on which the learning data generation program is recorded also falls within the scope of the present invention.

[0008] Furthermore, a learning data generation method according to another aspect of the present invention includes: an acquisition step in which a computer acquires an explanatory text from a language model that outputs an explanatory text in a predetermined format describing the situation of a vehicle when an image captured by a camera mounted on the vehicle is input; an evaluation step in which the computer evaluates the estimation accuracy of the language model for each of the multiple elements constituting the explanatory text; a first generation step in which the computer generates a learning text for the language model to learn based on the evaluation result of the estimation accuracy; and a second generation step in which the computer generates a learning video showing the content of the learning text for the language model to learn together with the learning text. [Brief explanation of the drawing]

[0009] [Figure 1] This block diagram shows an example of the functional configuration of a learning data generation system and a model learning system according to an embodiment of the present invention. [Figure 2] This figure illustrates the format of explanatory text generated by the language model included in the model learning system according to the same embodiment. [Figure 3] This figure shows an example of a video feed that is input to the language model. [Figure 4]This table shows an example of the evaluation results of the estimation accuracy of the language model by the learning data generation system according to the same embodiment. [Figure 5] This figure illustrates how training texts are generated by the training data generation system according to the same embodiment. [Figure 6] This figure illustrates how to generate training videos using the training data generation system according to the same embodiment. [Figure 7] This table shows the contents of the training data (dataset) accumulated by the model learning system according to the embodiment of this model. [Figure 8] This flowchart shows an example of the flow of a training data generation method and a model training method according to an embodiment of the present invention. [Modes for carrying out the invention]

[0010] <Training Data Generation System 1 (Model Learning System 100)> Hereinafter, a learning data generation device according to one embodiment of the present invention will be described in detail.

[0011] [Configuration of Model Learning System 100 (Model Learning System 100)] As shown in Figure 1, the learning data generation system 1 according to this embodiment constitutes a model learning system 100. The model learning system 100 includes the learning data generation system 1, a language model 2, a storage unit 3, and a tuning unit 4.

[0012] [Language Model 2] Language Model 2 is a trained model constructed using machine learning, with training data consisting of pairs of images and sentences describing the content of those images. Therefore, when Language Model 2 receives video footage captured by a camera mounted on a vehicle, it outputs a descriptive sentence in a predetermined format that describes the situation of the vehicle. The captured video footage may be of the exterior or interior of the vehicle.

[0013] The specified format indicates that the explanatory text contains essential elements. The "elements" include at least the driving situation of the vehicle. Specifically, as shown in Figure 2, the "elements" include at least the position of the vehicle and the driving intention of the vehicle. In addition, the "elements" further include at least any one of the objects shown in the captured video, the position of the object, the movement of the object, and the current movement of the vehicle. Note that the "position of the object" may further include the relative position to the vehicle and the absolute position.

[0014] When one captured video is input, the language model 2 according to this embodiment may output explanatory texts of multiple types with different contents of the driving situation to be explained respectively. For example, when a captured video in which a car and a bicycle (or a pedestrian) are simultaneously captured as shown in Figure 3 is input, the language model 2 generates an explanatory text about the car and an explanatory text about the bicycle. Note that the language model 2 may be provided in the vehicle, or may be provided in the learning data generation system 1. In addition, the language model 2 may be provided in another device independent of the vehicle and the learning data generation system 1.

[0015] 〔Learning Data Generation System 1〕 The learning data generation system 1 according to this embodiment includes an acquisition unit 11, an evaluation unit 12, a first generation unit 13, and a second generation unit 14. The learning data generation system 1 according to this embodiment further includes a selection unit 15, a second language model 16, a second storage unit 17, and a third storage unit 18.

[0016] (Second Language Model 16) The second language model 16 is a trained model different from the language model 2. The second language model is a large language model (LLMs). Specifically, the second language model 16 is composed of at least any one of "BestScore", "Sentence Bert", and "GPT4". Note that the second language model 16 may be a trained model other than the large language model.

[0017] (Second Storage Unit 17) The second storage unit 17 stores the correct sentences used by the evaluation unit 12 described later. The second storage unit 17 according to the present embodiment is composed of, for example, a semiconductor memory, a hard disk, or the like.

[0018] (Third storage unit 18) The third storage unit 18 stores the template video used by the second generation unit 14 described later. The third storage unit 18 according to the present embodiment is composed of, for example, a semiconductor memory, a hard disk, or the like. Note that the third storage unit 18 may be integrally configured with the second storage unit 17.

[0019] (Acquisition unit 11) The acquisition unit 11 executes an acquisition process. In the acquisition process, the acquisition unit 11 acquires an explanatory text from the language model. As described above, when one captured video is input, the language model 2 according to the present embodiment outputs a plurality of types of explanatory texts. Therefore, the acquisition unit 11 according to the present embodiment acquires a plurality of types of explanatory texts from the language model 2.

[0020] (Selection unit 15) The selection unit 15 executes a selection process. In the selection process, the selection unit 15 selects an explanatory text with the most significant impact on the future driving of the vehicle from among the plurality of types of explanatory texts. The selection of the explanatory text may be performed using a learned model, or the content of each explanatory text may be quantified and performed based on each numerical value. Also, when the language model 2 generates only one explanatory text for one captured video, the learning data generation system 1 may not include the selection unit 15.

[0021] (Evaluation unit 12) The evaluation unit 12 performs an evaluation process. In the evaluation process, the evaluation unit 12 evaluates the estimation accuracy of the language model for each of the multiple elements that make up the explanatory text. As described above, the learning data generation system 1 according to this embodiment includes a selection unit 15. Therefore, the evaluation unit 12 according to this embodiment evaluates the estimation accuracy of the language model using the explanatory text selected by the selection unit 15. The evaluation unit 12 according to this embodiment calculates the accuracy rate for each element by comparing the explanatory text with the correct answer text using the second language model 16. As a result, as shown in Figure 4, the language model 2 is evaluated such that, for a certain element (for example, the position of a vehicle), it outputs an explanatory text stating that the vehicle is at an intersection with an accuracy rate of 30% for a video showing an intersection, and outputs an explanatory text stating that the vehicle is on a straight road with an accuracy rate of 100% for a video showing a straight road. Note that the evaluation unit 12 may be configured to check whether the format of the explanatory text is in English first before evaluating the explanatory text.

[0022] (First generation section 13) The first generation unit 13 executes the first generation process. In the first generation process, the first generation unit 13 generates training texts for the language model to learn from, based on the evaluation results of the estimation accuracy. In this embodiment, the first generation unit 13 generates training texts in which elements with low accuracy rates are weighted, as shown in Figure 5. That is, the learning data generation system 1 generates training data (a pair of training video and training text: teacher data) that improves the accuracy of elements with low estimation accuracy in the language model. In addition, the first generation unit 13 in this embodiment generates the training texts in the same format as the explanatory texts. The first generation unit 13 may also be configured to generate training data that improves the accuracy of all elements, or training data that further improves the accuracy of elements with high estimation accuracy.

[0023] (Second generation unit 14) The second generation unit 14 executes a second generation process. In the second generation process, the second generation unit 14 generates a training video based on the training text generated by the first generation unit 13. The training video is a video that shows the content of the training text, which is to be trained on the language model together with the training text. In this embodiment, the second generation unit 14 generates the training video using pre-prepared template videos (stored in the third storage unit 18). Each template video has content that corresponds to each element of the explanatory text. For example, the template video includes a video of a road (straight road, intersection, T-junction, etc.) at the location of a vehicle indicated in the explanatory text. For example, if it is desired to generate a training video of a pedestrian crossing a straight road while a vehicle is moving straight down the road, the second generation unit 14 obtains a template video of a straight road from among multiple road template videos linked to the road map, as shown in Figure 6. The second generation unit 14 also obtains a template video of a pedestrian crossing the road from among the pedestrian template videos. The second generation unit 14 then generates a training video by combining the acquired template video of a pedestrian with the acquired template video of a straight road. In this way, the second generation unit 14 can generate training videos more easily than before (compared to creating them from scratch). Note that the template video may include elements that are not associated with any particular element (for example, elements unrelated to the position of a vehicle as indicated in the explanatory text).

[0024] [Storage section 3] The memory unit 3 is capable of storing the learning data (a pair of learning videos and learning texts: teacher data) generated by the learning data generation system 1. As a result, as the learning data generation system 1 repeatedly generates learning data, the memory unit 3 accumulates multiple pairs of learning texts in various situations (with different elements) and learning videos linked to the learning texts to be written, as shown in Figure 7. In this embodiment, the memory unit 3 is composed of, for example, a semiconductor memory or a hard disk. Note that the memory unit 3 may also be provided by the learning data generation system 1.

[0025] [Tuning Section 4] The tuning unit 4 fine-tunes the language model using the learning data stored in the memory unit 3. The tuning unit 4 may be provided in the vehicle or in the learning data generation system 1. Alternatively, the tuning unit 4 may be provided in another device independent of the vehicle and the learning data generation system 1.

[0026] [Effects of the learning data generation system 1 (model learning system 100)] The learning data generation system 1 described above automatically generates learning data (a pair of learning videos and learning texts: training data) specialized for the specific domain of remote monitoring by evaluating each element of an explanatory text in a predetermined format. Therefore, with the learning data generation system 1, it is possible to create a loop (model learning system 100) in which the language model 2 is automatically trained using the newly generated learning data, and the accuracy of text generation automatically improves. As a result, the effort of manually generating a large amount of learning data can be eliminated. In other words, with the learning data generation system 1, it becomes possible to easily train a language model.

[0027] Furthermore, according to the model learning system 100, the estimation accuracy of the language model (the accuracy of the content of the generated explanatory text) gradually improves.

[0028] <Training data generation method S1 (Model training method S100)> The following describes in detail a learning data generation method according to another embodiment of the present invention.

[0029] [Model Learning Method S100 Flow] As shown in Figure 8, the training data generation method S1 according to this embodiment is part of the model learning method S100. The model learning method S100 includes the training data generation method S1 as well as a tuning step S2.

[0030] [Flow of training data generation method S1] The learning data generation method S1 includes an acquisition step S11, an evaluation step S12, a first generation step S13, and a second generation step S14. The learning data generation method S1 according to this embodiment further includes a selection step S15.

[0031] (Acquisition step S11) In the initial acquisition step S11, explanatory text is obtained from the language model 2. The acquisition of explanatory text may be performed using the learning data generation system 1 described above, or using other systems or devices. As described above, the language model 2 according to this embodiment outputs multiple types of explanatory text when a single captured video is input. Therefore, in the acquisition step S11 according to this embodiment, multiple types of explanatory text are obtained from the language model 2.

[0032] (Selection step S15) After obtaining the explanatory text, the process moves to the selection step. In the selection step S15, the explanatory text that has the greatest impact on the vehicle's future driving is selected from among several types of explanatory texts. In this embodiment, the selection unit 15 makes the selection using a trained model. The selection of explanatory texts may be performed using the training data generation system 1 described above, or it may be performed using other systems or devices. Note that if the language model 2 generates only one explanatory text for each captured video, the training data generation method S1 does not need to include the selection step S15.

[0033] (Evaluation step S12) After obtaining or selecting the explanatory text, the process moves to the evaluation step. In evaluation step S12, the estimation accuracy of the language model 2 is evaluated for each of the multiple elements that make up the explanatory text. The evaluation of estimation accuracy may be performed using the learning data generation system 1 described above, or it may be performed using other systems or devices.

[0034] (First generation step S13) After evaluating the estimation accuracy of language model 2, the process moves to the first generation step. In the first generation step S13, training texts are generated for training the language model based on the evaluation results of the estimation accuracy. The training texts may be generated using the training data generation system 1 described above, or they may be generated using other systems or devices.

[0035] (Second generation step S14) After generating the training text, the process moves to the second generation step. In the second generation step S14, training videos are generated based on the training text. The training videos may be generated using the training data generation system 1 described above, or they may be generated using other systems or devices.

[0036] (Tuning Step S2) After generating the training data (a set of training text and training video), the process moves to tuning step S2. In tuning step S2, the language model 2 is fine-tuned using the training data. The fine-tuning of language model 2 may be performed using the model learning system 100 described above, or it may be performed using other systems or devices.

[0037] [Effects of training data generation method S1 and model training method S100] The training data generation method S1 described above automatically generates training data specific to the domain of remote monitoring by evaluating each element of an explanatory text in a predetermined format. Therefore, with training data generation method S1, it is possible to create a loop in which the language model is automatically trained using the newly generated training data, and the accuracy of text generation automatically improves. As a result, the effort of manually generating a large amount of training data can be eliminated. In other words, training the language model becomes easier with training data generation method S1.

[0038] Furthermore, according to the model learning method S100, the estimation accuracy of the language model (the accuracy of the content of the generated explanatory text) gradually improves.

[0039] The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

[0040] For example, the memory unit 3 may be provided by the learning data generation system 1.

[0041] For example, each part of the learning data generation system 1 is a program that causes the computer to function as each part, and each part can be realized by a learning data generation program that causes the computer to function as each part. In this case, each component includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the above-mentioned learning data generation program. Each component is realized by executing the above-mentioned learning data generation program using this control device and storage device. The above-described learning data generation program may be recorded on one or more computer-readable recording media, not temporary ones. Each unit may or may not have such recording media. In the latter case, the program may be supplied to each unit via any wired or wireless transmission medium. Furthermore, some or all of the functions of each part can be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each part are formed is also included in the scope of this invention. In addition, it is also possible to realize the functions of each part by, for example, a quantum computer.

[0042] 〔summary〕 A learning data generation system according to embodiment 1 of the present invention comprises: an acquisition unit that, upon input of captured video footage obtained from a camera mounted on a vehicle, acquires explanatory text from a language model that outputs explanatory text in a predetermined format describing the situation of the vehicle; an evaluation unit that evaluates the estimation accuracy of the language model for each of the multiple elements constituting the explanatory text; a first generation unit that generates learning text for the language model to learn based on the evaluation result of the estimation accuracy; and a second generation unit that generates learning video showing the content of the learning text, to be learned by the language model together with the learning text, based on the learning text.

[0043] In the learning data generation system according to aspect 2 of the present invention, the elements may be configured such that they include at least the driving conditions of the vehicle, as described in aspect 1 above.

[0044] The learning data generation system according to embodiment 3 of the present invention may be configured such that, in embodiment 2 above, the elements include at least the position of the vehicle and the vehicle's driving intention, and further include at least one of an object shown in the captured video, the position of the object, the movement of the object, and the current movement of the vehicle.

[0045] In the learning data generation system according to aspect 4 of the present invention, in any of aspects 1 to 3 described above, the evaluation unit may be configured to calculate the correct answer rate for each element by comparing the explanatory text with the correct answer text using a second language model different from the language model.

[0046] In the learning data generation system according to aspect 5 of the present invention, the first generation unit may be configured to generate the learning text in which the elements with low correct answer rates are weighted, as described in aspect 4 above.

[0047] In the learning data generation system according to embodiment 6 of the present invention, the first generation unit may be configured to generate the learning text in the same format as the explanatory text, as described in embodiment 5 above.

[0048] In the learning data generation system according to embodiment 7 of the present invention, in any of embodiments 1 to 6 described above, the second generation unit may be configured to generate the learning video using a pre-prepared template video.

[0049] In the learning data generation system according to embodiment 8 of the present invention, in embodiment 7 described above, the template video may be configured to include a video of the road at the location of the vehicle indicated in the explanatory text.

[0050] A learning data generation system according to aspect 9 of the present invention may be configured such that, in any one of aspects 1 to 8 described above, the acquisition unit acquires a plurality of types of explanatory texts from a language model that outputs a plurality of types of explanatory texts, each with different content describing driving conditions, when the captured video is input, and selects the explanatory text from the plurality of types of explanatory texts that has the greatest impact on the future driving of the vehicle, and the evaluation unit evaluates the estimation accuracy of the language model using the selected explanatory text.

[0051] A learning data generation program according to embodiment 10 of the present invention is configured to cause a computer to execute the following: an acquisition process to acquire an explanatory text from a language model that outputs an explanatory text in a predetermined format describing the situation of a vehicle when an image captured by a camera mounted on a vehicle is input to the computer; an evaluation process to evaluate the estimation accuracy of the language model for each of the multiple elements constituting the explanatory text; a first generation process to generate a learning text for the language model to learn based on the evaluation result of the estimation accuracy; and a second generation process to generate a learning video showing the content of the learning text, which is to be learned by the language model together with the learning text, based on the learning text.

[0052] A learning data generation method according to aspect 11 of the present invention is a method comprising: an acquisition step in which a computer acquires an explanatory text from a language model that outputs an explanatory text in a predetermined format describing the situation of a vehicle when an image captured by a camera mounted on the vehicle is input; an evaluation step in which the computer evaluates the estimation accuracy of the language model for each of a plurality of elements constituting the explanatory text; a first generation step in which the computer generates a learning text for the language model to learn based on the evaluation result of the estimation accuracy; and a second generation step in which the computer generates a learning video showing the content of the learning text for the language model to learn together with the learning text. [Explanation of Symbols]

[0053] 100 Model Learning Systems 1. Training Data Generation System 11 Acquisition Department 12 Evaluation Department 13 First generation part 14 Second generation part 15 Selection Section 16. Second Language Model 17 Second memory section 18 Third Memory 2 language models 3 Storage section 4. Tuning Section S100 Model Learning Method S1 Method for generating training data S11 Acquisition Steps S12 Evaluation Steps S13 First Generation Step S14 Second generation step S15 Selection Step S2 Tuning Step

Claims

1. When video footage captured by a camera mounted on the vehicle is input, an acquisition unit acquires the explanatory text from a language model that outputs an explanatory text in a predetermined format that describes the status of the vehicle. An evaluation unit evaluates the estimation accuracy of the language model for each of the multiple elements that constitute the explanatory text, A first generation unit generates training texts for training the language model based on the evaluation results of the estimation accuracy, A second generation unit generates a learning video showing the content of the learning text, which is to be trained on the language model together with the learning text, based on the learning text. Equipped with, A system for generating training data.

2. The aforementioned elements include at least the driving conditions of the vehicle, The learning data generation system according to claim 1.

3. The aforementioned element is This includes at least the location of the vehicle and the intention of the vehicle to travel, The definition further includes at least one of the following: an object visible in the captured video, the position of the object, the movement of the object, and the current state of the vehicle. The learning data generation system according to claim 2.

4. The evaluation unit calculates the correct answer rate for each element by comparing the explanatory text with the correct answer text using a second language model different from the language model. The learning data generation system according to claim 1.

5. The first generation unit generates the training text in which the elements with low correct answer rates are weighted. The learning data generation system according to claim 4.

6. The first generation unit generates the learning text in the same format as the explanatory text. The learning data generation system according to claim 5.

7. The second generation unit generates the learning video using a pre-prepared template video. The learning data generation system according to claim 1.

8. The template video includes a video of the road at the location of the vehicle as indicated in the explanatory text. The learning data generation system according to claim 7.

9. When the acquired unit receives the captured video as input, it acquires multiple types of explanatory texts from a language model that outputs multiple types of explanatory texts, each of which describes different driving conditions. The system includes a selection unit that selects the explanatory document from among several types of explanatory documents that has the greatest impact on the future driving of the vehicle, The evaluation unit evaluates the estimation accuracy of the language model using the selected explanatory text. A learning data generation system according to any one of claims 1 to 8.

10. On the computer, When video footage captured by a camera mounted on the vehicle is input, an acquisition process is performed to obtain the explanatory text from a language model that outputs an explanatory text in a predetermined format that describes the status of the vehicle. For each of the multiple elements that constitute the explanatory text, an evaluation process is performed to evaluate the estimation accuracy of the language model, Based on the evaluation results of the estimation accuracy, a first generation process generates training texts for training the language model, A second generation process generates a learning video showing the content of the learning text, based on the learning text, to be trained on the language model together with the learning text. To execute A program for generating training data.

11. The computer, upon receiving video footage captured by a camera mounted on the vehicle, receives an input from a language model that outputs a descriptive document in a predetermined format describing the vehicle's condition. This includes an acquisition step of obtaining the descriptive document from the language model. An evaluation step in which the computer evaluates the estimation accuracy of the language model for each of the multiple elements that constitute the explanatory text, A first generation step in which the computer generates training text for the language model to be trained based on the evaluation result of the estimation accuracy, A second generation step in which the computer generates a learning video showing the content of the learning text, based on the learning text, to be used to train the language model along with the learning text; including, Method for generating training data.

Citation Information

Patent Citations

  • Explanation sentence creation device

    JP2021174172A

  • Sentence generator, program, and method for generating sentence

    JP2022116979A

  • System and program

    JP2025025521A