Evaluation system, evaluation program, and evaluation method

The evaluation system addresses the inadequacy of conventional methods by using a dual-language model approach to accurately assess sentence accuracy and relevance for autonomous driving, enhancing the evaluation of language models' safety information extraction capabilities.

JP2025162818AActive Publication Date: 2025-10-28SOFTBANK CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024066260
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-16
Publication Date
2025-10-28
Estimated Expiration
2044-04-16

AI Technical Summary

Technical Problem

Conventional evaluation methods for language models in autonomous driving lack optimization for remote monitoring, failing to accurately distinguish between sentences with similar word usage or context, which is crucial for evaluating their ability to extract important information for autonomous vehicle safety.

Method used

An evaluation system and method that utilizes a first language model and a second language model trained on a traffic-related perspective, incorporating word-based and context-based evaluation, along with a human-perspective evaluation to assess the accuracy and relevance of generated sentences for autonomous driving scenarios.

Benefits of technology

The system accurately evaluates the performance of language models by distinguishing between sentences with different meanings, ensuring they can effectively extract critical safety-related information, reducing the need for manual verification and lowering evaluation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025162818000001_ABST
    Figure 2025162818000001_ABST
Patent Text Reader

Abstract

To enable accurate and easy evaluation of a first language model that outputs a sentence describing the content of a traffic-related image when the image is input.SOLUTION: An evaluation system (100) evaluates the performance of a first language model and includes: a first evaluation unit configured to output a first evaluation value indicating the accuracy of a sentence based on a predetermined evaluation index; a second evaluation unit configured to use a second language model constructed through machine learning based on a traffic-related viewpoint and output a second evaluation value indicating the accuracy of a sentence; and a third evaluation unit configured to output an evaluation result of the first language model based on the first and second evaluation values. The second evaluation unit inputs, into the second language model, a correct answer sentence indicating the content that should be mainly interpreted from an image, together with a sentence and outputs, based on a first index derived from the content of the correct answer sentence and a second index derived from the content of the sentence, the second evaluation value indicating the degree of separation between the first and second indices, calculated by the second language model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an evaluation system, an evaluation program, and an evaluation method. [Background technology]

[0002] In recent years, with the lifting of the ban on "Level 4" autonomous driving, demonstration experiments of autonomous vehicles have been conducted across the country. In autonomous driving, a huge amount of data is collected from the vehicle, and this data is used by the vehicle's AI to recognize events, predict future events, plan trips, and improve (learn) the AI. For example, Non-Patent Document 1 below discloses the current status of demonstration experiments of such autonomous vehicles. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] “Autonomous driving demonstrations are becoming more active both domestically and internationally! What are the challenges in data management?” [online], October 15, 2020, Autonomous Driving Lab, [Retrieved July 7, 2023], Internet<URL:https: / / jidounten-lab.com / u_data-management-1> Summary of the Invention [Problem to be solved by the invention]

[0004] Conventionally, evaluation of language models (multimodal AI) that output sentences describing the content of traffic-related images when they are input has been carried out by comparing the sentences generated by the language model with the correct answer according to a predetermined evaluation index (word-based evaluation indexes such as BLEU and BERTscore, or context-based evaluation indexes such as sentenceBERT). However, none of the evaluation indexes used so far have been optimized from the perspective of remote monitoring of autonomous driving. [Means for solving the problem]

[0005] In order to solve the above problem, one embodiment of the present invention provides an evaluation system that evaluates the performance of a first language model that, when a traffic-related image is input, outputs a sentence describing the content of the image, and includes: a first evaluation unit that outputs a first evaluation value indicating the accuracy of the sentence based on a predetermined evaluation index; a second evaluation unit that uses a second language model different from the first language model, the second language model being constructed by machine learning based on a traffic-related perspective, and outputs a second evaluation value indicating the accuracy of the sentence; and a third evaluation unit that outputs an evaluation result of the first language model based on the first evaluation value and the second evaluation value.The second evaluation unit inputs a correct answer sentence, which is a sentence that indicates the content that should most be read from the image, and the sentence into the second language model, and outputs a second evaluation value that indicates the degree of deviation between the first index and the second index based on the first index based on the content of the correct answer sentence and the second index based on the content of the sentence, calculated by the second language model.

[0006] Another aspect of the present invention provides an evaluation program for evaluating the performance of a first language model that, when a traffic-related image is input, outputs a sentence describing the content of the image. The evaluation program causes a computer to execute a first evaluation process that outputs a first evaluation value indicating the accuracy of the sentence based on a predetermined evaluation index, a second evaluation process that uses a second language model different from the first language model and constructed by machine learning based on a traffic-related perspective to output a second evaluation value indicating the accuracy of the sentence, and a third evaluation process that outputs an evaluation result of the first language model based on the first evaluation value and the second evaluation value. In the second evaluation process, the computer inputs a correct answer sentence, which is a sentence that indicates the content that should be most read from the image, and the sentence into the second language model, and outputs a second evaluation value that indicates the degree of deviation between the first index and the second index based on the first index and the second index calculated by the second language model. Note that a computer-readable recording medium having a training data generation program recorded thereon is also within the scope of the present invention.

[0007] Another aspect of the present invention provides an evaluation method for evaluating the performance of a first language model that, when a traffic-related image is input, outputs a sentence describing the content of the image, the evaluation method including: a first evaluation step in which a computer outputs a first evaluation value indicating the accuracy of the sentence based on a predetermined evaluation index; a second evaluation step in which the computer uses a second language model different from the first language model, the second language model being constructed by machine learning based on a traffic-related perspective, to output a second evaluation value indicating the accuracy of the sentence; and a third evaluation step in which the computer outputs an evaluation result of the first language model based on the first evaluation value and the second evaluation value. In the second evaluation step, the computer inputs a correct answer sentence, which is a sentence indicating the content that should most be read from the image, and the sentence to the second language model, and outputs a second evaluation value that indicates the degree of deviation between the first index and the second index based on the first index based on the content of the correct answer sentence and the second index calculated by the second language model. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram showing an example of a functional configuration of an evaluation system according to an embodiment of the present invention. [Figure 2] 10A and 10B are diagrams illustrating a modified example of a method for outputting a second evaluation value performed by a second evaluation unit included in the system. [Figure 3] 10A and 10B are diagrams illustrating a modified example of a method for outputting a second evaluation value performed by a second evaluation unit included in the system. [Figure 4] 10A and 10B are diagrams illustrating a modified example of a method for outputting a second evaluation value performed by a second evaluation unit included in the system. [Figure 5] 1 is a flowchart showing an example of the flow of an evaluation method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] <Rating System 100> The evaluation system 100 according to an embodiment of the present invention will be described in detail below.

[0010] [Evaluation system target] The evaluation system 100 is a system for evaluating the performance of a first language model M1. The first language model M1 to be evaluated is a trained model constructed by machine learning using a set of traffic-related images and sentences describing the contents of the images as training data. The traffic-related images are images captured by a camera mounted on an autonomous vehicle. The images may be images of the exterior or interior of the vehicle. The images may be videos or still images. When a new traffic-related image is input, the first language model M1 outputs sentences describing the contents of the image.

[0011] [Evaluation system configuration] The evaluation system 100 evaluates the performance of the first language model M1 based on the sentences output by the first language model M1. As shown in FIG. 1, the evaluation system 100 includes a first evaluation unit 1, a second evaluation unit 2, and a third evaluation unit 3.

[0012] [First Evaluation Section 1] The first evaluation unit 1 outputs a first evaluation value indicating the accuracy of a sentence based on a predetermined evaluation index. The first evaluation unit 1 according to this embodiment includes a word evaluation unit 11, a context evaluation unit 12, and a calculation unit 13.

[0013] (Word Evaluation Section 11) The word evaluation unit 11 outputs a first numerical value based on the number of words included in a sentence that match with words included in a correct sentence corresponding to the sentence. The correct sentence is a sentence that explains the content of a traffic-related image input to the first language model M1 and is generated by a means other than the first language model M1 (for example, a person). In other words, the correct sentence is a sentence that indicates the content that should be most easily read from the image. The correct sentence may be stored in a storage unit (not shown), or may be generated each time evaluation is performed. The word evaluation unit 11 according to this embodiment outputs the first numerical value using an evaluation index such as "BLEU" or "BERTscore." The first numerical value output by the word evaluation unit 11 according to this embodiment is a number greater than 0 and equal to or less than 1.

[0014] (Context Evaluator 12) The context evaluation unit 12 outputs a second numerical value based on the similarity between the context of the sentence and the context of the correct sentence corresponding to the sentence. The word evaluation unit 11 according to this embodiment outputs the second numerical value using an evaluation index such as "sentenceBERT." The second numerical value output by the context evaluation unit 12 according to this embodiment is a numerical value greater than 0 and equal to or less than 1.

[0015] (Calculation unit 13) The calculation unit 13 calculates a first evaluation value based on the first numerical value and the second numerical value. The first evaluation value is a statistical value of the first numerical value and the second numerical value. "Statistical value" includes an additive value, a weighted additive value, an average value, a weighted average value, etc. The additive value is the sum of the first evaluation value and the second evaluation value. The average value is the average of the first evaluation value and the second evaluation value. The weighted average value is the average of values ​​obtained by multiplying at least one of the first evaluation value and the second evaluation value by a weighting coefficient. The first evaluation value is a numerical value greater than 0 and less than or equal to 1. In other words, the closer the first evaluation value is to 1, the higher the accuracy of the sentence is evaluated by the first evaluation unit 1.

[0016] [Second Evaluation Section 2] The second evaluation unit 2 uses a second language model M2 to output a second evaluation value indicating the accuracy of the sentence. The second language model M2 is a language model different from the first language model M1 and is a trained model constructed by machine learning based on a traffic-related perspective. The second language model M2 used by the second evaluation unit 2 according to this embodiment is a large language model (LLM). Specifically, the second language model M2 is configured using at least one of "BestScore," "Sentence Bert," and "GPT4." By using the second language model M2 configured in this manner, the accuracy of the sentence (the performance of the first language model M1) can be evaluated from a human perspective. That is, sentences that cannot be distinguished using word-based evaluation metrics because they use almost the same words (e.g., sentences with reverse causal relationships, such as "The car stopped suddenly because the truck went straight" and "The truck went straight because the car stopped suddenly") can be distinguished as sentences with different content. In addition, sentences that could not be distinguished using context-based evaluation indices (for example, sentences with almost the same context, such as "The car stopped suddenly because the truck went straight" and "The car stopped slowly because the truck went straight") can also be distinguished as sentences with different contents. Note that the second language model M2 may be a trained model other than a large-scale language model. The second evaluation unit 2 includes a human-perspective evaluation unit 21 and a second calculation unit 22.

[0017] (Human Perspective Evaluation Division 21) The human-perspective evaluation unit 21 inputs a sentence and a correct answer sentence corresponding to the sentence to the second language model M2. The human-perspective evaluation unit 21 includes a first-person perspective evaluation unit 211 and a second-person perspective evaluation unit 212. The first-person perspective evaluation unit 211 inputs the correct answer sentence to the second language model M2. The second-person perspective evaluation unit 212 inputs the sentence to the second language model M2. The second language model M2 to which the sentence and the correct answer sentence have been input calculates a first index and a second index. The first index is an index based on the content of the correct answer sentence. The first index calculated by the second language model M2 according to this embodiment is a risk level indicating the degree of risk of the event indicated by the correct answer sentence. The second index is an index based on the content of the sentence. The second index calculated by the second language model M2 according to this embodiment is a risk level indicating the degree of risk of the event indicated by the sentence. The first index and second index output by the second language model M2 according to this embodiment are numerical values ​​greater than 0 and equal to or less than 1. Then, the first-person perspective evaluation unit 211 acquires the first index calculated by the second language model M2. Also, the second-person perspective evaluation unit 212 acquires the second index calculated by the second language model M2.

[0018] (Second calculation unit 22) The second calculation unit 22 calculates a numerical value indicating the degree of deviation between the first index and the second index based on the first index and the second index calculated by the second language model M2. Specifically, the second calculation unit 22 calculates the numerical value by substituting the first index and the second index into the following formula (1). The second calculation unit 22 then outputs the numerical value calculated by the second language model M2 as a second evaluation value. The second evaluation value output by the second calculation unit 22 is a numerical value greater than 0 and equal to or less than 1. In other words, the closer the second evaluation value is to 1, the smaller the deviation between the content of the sentence and the correct sentence (the higher the accuracy of the sentence), which is evaluated by the second evaluation unit 2. The value indicating the degree of deviation = 1.0 - |First index - Second index|··(1)

[0019] 2, the second evaluation unit 2 may be configured to input, before inputting a sentence into the second language model M2, one or more pairs of example sentences describing traffic-related events and risk levels indicating the degree of risk of the events described in the example sentences into the second language model M2. In this way, by performing short-shot learning, the second language model M2 can calculate a more accurate first index. As a result, the second evaluation unit 2 can output a more accurate second evaluation value.

[0020] Furthermore, as shown in FIG. 3, the second evaluation unit 2 may be configured to input, to the second language model M2, in addition to the correct sentence and the sentence, an instruction (prompt) to evaluate the image from the perspective of traffic safety. The instruction is an instruction to evaluate from the perspective of a remote monitor of the autonomous vehicle. In this case, the second evaluation unit 2 outputs, as the second evaluation value, a numerical value indicating the degree of deviation between the first index and the second index based on the first index and the second index, which the second language model M2 obtains based on the instruction. In this way, the second language model M2 can calculate a more accurate first index. As a result, the second evaluation unit 2 can output a more accurate second evaluation value.

[0021] The second evaluation unit 2 may also be configured to output the second evaluation value using Retrieval Augmented Generation (RAG). In this case, the evaluation system 100 further includes a database 4, as shown in FIG. 4. The database 4 stores multiple past sentences with different contents, each of which describes a traffic-related image from the past. The past sentences include sentences previously output by the first language model M1, sentences describing traffic accidents recorded by the police, and sentences about traffic accidents reported in the past. When inputting a sentence into the second language model M2, the second evaluation unit 2 extracts past sentences similar in content to the sentence from the multiple past sentences stored in the database 4. The second evaluation unit 2 then inputs the extracted past sentences together with the sentence into the second language model M2. This allows the second language model M2 to calculate a more accurate first index. As a result, the second evaluation unit 2 can output a more accurate second evaluation value.

[0022] [Third Evaluation Section 3] The third evaluation unit 3 outputs an evaluation result of the first language model M1 based on the first evaluation value and the second evaluation value. The third evaluation unit 3 according to this embodiment outputs the product of the first evaluation value and the second evaluation value as the evaluation result. As described above, the first evaluation value and the second evaluation value are each a numerical value greater than 0 and equal to or less than 1. Therefore, the evaluation result, which is the product of the first evaluation value and the second evaluation value, is also a numerical value greater than 0 and equal to or less than 1, and the closer to 1 the value is, the higher the performance of the first language model M1.

[0023] [Modification of the evaluation system 100] The first language model M1 may be constructed to output a formatted sentence. A "formatted sentence" refers to a sentence that includes essential elements. The "elements" include at least the vehicle's driving situation (vehicle position, vehicle driving intention, objects, object positions, object movements, current vehicle movements, etc.). Furthermore, when one image is input, the first language model M1 may output multiple types of sentences each describing a different driving situation. In this case, the first evaluation unit 1 and the second evaluation unit 2 may be configured to acquire multiple types of sentences from the first language model M1 and select, from the multiple types of explanatory sentences, the sentence with the content that will have the greatest impact on the future driving of the vehicle.

[0024] 1 illustrates an example of an evaluation system 100 that does not include the first language model M1 and the second language model M2. However, the evaluation system 100 may include at least one of the first language model M1 and the second language model M2.

[0025] 1 illustrates an example in which sentences are directly supplied from the first language model M1 to the first evaluation unit 1 and the second evaluation unit 2. However, the evaluation system 100 may be configured to store the sentences output by the first language model M1 in a storage unit (not shown). The first evaluation unit 1 and the second evaluation unit 2 may be configured to indirectly acquire the sentences from the storage unit.

[0026] [Effects of the Evaluation System 100] For example, the sentence "The car stopped suddenly because the truck was going straight ahead" and the sentence "The truck went straight ahead because the car stopped suddenly" have reversed causal relationships, and would be considered completely different sentences from a human perspective. However, conventional word-based evaluation metrics could not distinguish between the two sentences because they use almost the same words. Furthermore, the sentences "The car stopped suddenly because the truck was going straight ahead" and "The car stopped slowly because the truck was going straight ahead" differ in the way the car was stopped, and would be considered different safety-related sentences from a human perspective. However, conventional context-based evaluation metrics could not distinguish between the two sentences because the contexts are almost the same. In other words, conventional methods have low accuracy in evaluating whether a language model can correctly document situations when a high-risk event occurs, which is the most important requirement for automating remote monitoring. For this reason, humans still had to verify whether the sentences generated by a language model extract important information for remote monitoring. However, manual review of sentences generated by a language model and evaluation of a language model require significant effort and cost. However, in the evaluation system 100 described above, the third evaluation unit 3 outputs the evaluation result of the first language model M1 by taking into account not only the first evaluation value indicating the structural accuracy of the sentence based on conventional evaluation indexes, but also the second evaluation value indicating the accuracy of the sentence evaluated by the second language model M2 from the perspective of remote monitoring. Therefore, the evaluation system 100 makes it possible to accurately and easily evaluate the first language model M1 (confirm whether the sentence generated by the first language model M1 can extract information important for remote monitoring).

[0027] <Evaluation Method S100> The evaluation method S100 according to another embodiment of the present invention will be described in detail below.

[0028] [Evaluation method S100 flow] The evaluation method S100 is a method for evaluating the performance of the first language model M1. As shown in Fig. 5, the evaluation method S100 includes a first evaluation step S1, a second evaluation step S2, and a third evaluation step S3.

[0029] [First evaluation step S1] In the first evaluation step S1, the computer outputs a first evaluation value that indicates the accuracy of the text based on a predetermined evaluation index. The calculation of the first evaluation value may be performed using the first evaluation unit 1 of the evaluation system described above, or may be performed using other means.

[0030] [Second evaluation step S2] In the second evaluation step S2, the computer uses the second language model M2 to output a second evaluation value indicating the accuracy of the sentence. In the second evaluation step S2, the computer inputs the correct sentence and the sentence into the second language model M2. Also, in the second evaluation step S2, the computer outputs a numerical value indicating the degree of deviation between the first index and the second index based on the first index and the second index calculated by the second language model M2 as a second evaluation value. The calculation of the second evaluation value may be performed using the second evaluation unit 2 of the evaluation system or by other means.

[0031] [Third evaluation step S3] In the third evaluation step S3, the computer outputs an evaluation result of the first language model M1 based on the first evaluation value and the second evaluation value. The evaluation result may be output by the third evaluation unit 3 of the evaluation system or by other means.

[0032] [Effects of evaluation method S100] In the third evaluation step S3, the evaluation method S100 described above outputs an evaluation result of the first language model M1 by taking into account not only a first evaluation value indicating the structural accuracy of the sentence based on conventional evaluation indices, but also a second evaluation value indicating the accuracy of the sentence evaluated by the second language model M2 from the perspective of remote monitoring. Therefore, similar to the evaluation system 100, the evaluation method S100 can accurately and easily evaluate the first language model M1 (confirm whether the sentence generated by the first language model M1 can extract information important for remote monitoring).

[0033] <Modification> The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

[0034] For example, each unit of evaluation system 100 can be realized by a program that causes a computer to function as each unit, which is a training data generation program that causes a computer to function as each unit. In this case, each unit includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the training data generation program. Each unit is realized by executing each process (first evaluation process, second evaluation process, third evaluation process) of the training data generation program using this control device and storage device. Note that evaluation system 100 may be configured so that one computer realizes one of the units, or so that one computer realizes two or more of the units.

[0035] The training data generation program may be stored in one or more computer-readable storage media, rather than being stored temporarily. Each unit may or may not have a storage medium. In the latter case, the program may be supplied to each unit via any wired or wireless transmission medium.

[0036] In addition, some or all of the functions of each unit can be realized by logic circuits. For example, integrated circuits in which logic circuits functioning as each unit are formed are also included in the scope of the present invention. In addition, the functions of each unit can also be realized by, for example, a quantum computer.

[0037] 〔summary〕 The evaluation system according to a first aspect of the present invention is an evaluation system for evaluating the performance of a first language model that, when a traffic-related image is input, outputs a sentence describing the content of the image, and includes: a first evaluation unit that outputs a first evaluation value indicating the accuracy of the sentence based on a predetermined evaluation index; a second evaluation unit that uses a second language model different from the first language model, the second language model being constructed by machine learning based on a traffic-related perspective, and outputs a second evaluation value indicating the accuracy of the sentence; and a third evaluation unit that outputs an evaluation result of the first language model based on the first evaluation value and the second evaluation value. The second evaluation unit inputs a correct answer sentence, which is a sentence indicating the content that should most be read from the image, and the sentence into the second language model, and outputs a second evaluation value that indicates the degree of deviation between the first index and the second index based on the first index and the second index calculated by the second language model.

[0038] An evaluation system according to aspect 2 of the present invention may be configured in the above aspect 1 such that the first evaluation unit includes a word evaluation unit that outputs a first numerical value based on the number of words included in the sentence that match multiple words included in the correct sentence, a context evaluation unit that outputs a second numerical value based on the similarity between the context of the sentence and the context of the correct sentence, and a calculation unit that calculates the first evaluation value based on the first numerical value and the second numerical value.

[0039] An evaluation system according to aspect 3 of the present invention may be configured such that, in aspect 2 above, the first evaluation value and the second evaluation value are each numerical values ​​greater than 0 and less than or equal to 1, the first evaluation value is a statistical value of the first numerical value and the second numerical value, and the third evaluation unit outputs the product of the first evaluation value and the second evaluation value as the evaluation result.

[0040] The evaluation system according to aspect 4 of the present invention may be configured such that, in any of aspects 1 to 3 above, the first index is a risk level indicating the degree of risk of the event indicated by the correct sentence, and the second index is a risk level indicating the degree of risk of the event indicated by the sentence.

[0041] An evaluation system according to a fifth aspect of the present invention may be configured such that, in the above-mentioned fourth aspect, the second language model is a large-scale language model, and the second evaluation unit, before inputting the sentence into the second language model, inputs into the second language model one or more pairs of example sentences explaining traffic-related events and risk levels indicating the degree of risk of the events indicated by the example sentences.

[0042] The evaluation system according to aspect 6 of the present invention may be configured in any of aspects 1 to 5 above, to include a database that stores past sentences with different contents, each past sentence explaining the content of a past traffic-related image, and when inputting the sentence into the second language model, the second evaluation unit extracts past sentences from the past sentences that have similar content to the sentence, and inputs the extracted past sentence into the second language model together with the sentence.

[0043] The evaluation system of aspect 7 of the present invention may be configured such that, in any of aspects 1 to 6 above, the second evaluation unit inputs to the second language model, in addition to the correct sentence and the sentence, an instruction to evaluate the image from the perspective of traffic safety, and the second language model outputs, as the second evaluation value, a numerical value indicating the degree of deviation between the first index and the second index calculated based on the instruction.

[0044] An evaluation system according to aspect 8 of the present invention may be configured in the above-mentioned aspect 7 such that the traffic-related images are images obtained by capturing images using a camera mounted on an autonomous vehicle, and the instruction is to evaluate the autonomous vehicle from the perspective of a remote monitor.

[0045] An evaluation program according to a ninth aspect of the present invention is an evaluation program for evaluating the performance of a first language model that, when a traffic-related image is input, outputs a sentence describing the content of the image, and causes a computer to execute a first evaluation process that outputs a first evaluation value indicating the accuracy of the sentence based on a predetermined evaluation index; a second evaluation process that uses a second language model different from the first language model, the second language model being constructed by machine learning based on a traffic-related perspective, to output a second evaluation value indicating the accuracy of the sentence; and a third evaluation process that outputs an evaluation result of the first language model based on the first evaluation value and the second evaluation value. In the second evaluation process, the computer inputs a correct answer sentence, which is a sentence indicating the content that should most be read from the image, and the sentence into the second language model, and outputs a numerical value indicating the degree of deviation between the first index and the second index based on the first index calculated by the second language model based on the content of the correct answer sentence and the second index calculated by the second language model.

[0046] An evaluation method according to aspect 10 of the present invention is a method for evaluating the performance of a first language model that, when a traffic-related image is input, outputs a sentence describing the content of the image, the method comprising: a first evaluation step in which a computer outputs a first evaluation value indicating the accuracy of the sentence based on a predetermined evaluation index; a second evaluation step in which the computer uses a second language model different from the first language model, the second language model being constructed by machine learning based on a traffic-related perspective, to output a second evaluation value indicating the accuracy of the sentence; and a third evaluation step in which the computer outputs an evaluation result of the first language model based on the first evaluation value and the second evaluation value. In the second evaluation step, the computer inputs a correct answer sentence, which is a sentence indicating the content that should most be read from the image, and the sentence to the second language model, and outputs a numerical value indicating the degree of deviation between the first index and the second index based on a first index based on the content of the correct answer sentence and a second index based on the content of the sentence, calculated by the second language model. [Explanation of symbols]

[0047] 100 rating system 1. First Evaluation Department 11 Word Evaluation Section 12 Context Evaluator 13 Calculation section 2. Second Evaluation Department 21-person perspective evaluation department 22 Second calculation section 3. Third Evaluation Department 4 Database M1 First Language Model M2 Second Language Model S100 Evaluation Method S11 First evaluation step S12 Second evaluation step S13 Third evaluation step

Claims

1. 1. An evaluation system for evaluating performance of a first language model that, when a traffic-related image is input, outputs a sentence describing the content of the image, the system comprising: a first evaluation unit that outputs a first evaluation value that indicates the accuracy of the sentence based on a predetermined evaluation index; a second evaluation unit that uses a second language model different from the first language model, the second language model being constructed by machine learning based on a transportation-related perspective, and outputs a second evaluation value that indicates accuracy of the sentence; a third evaluation unit that outputs an evaluation result of the first language model based on the first evaluation value and the second evaluation value; Equipped with The second evaluation unit a correct answer sentence that indicates the content that should be most read from the image and the sentence are input to the second language model; outputting, as the second evaluation value, a numerical value indicating a degree of deviation between the first index and the second index calculated by the second language model based on the content of the correct sentence and the second index based on the content of the sentence; Rating system.

2. The first evaluation unit a word evaluation unit that outputs a first numerical value based on the number of words that match the words included in the correct sentence among the words included in the sentence; a context evaluation unit that outputs a second numerical value based on the similarity between the context of the sentence and the context of the correct answer sentence; a calculation unit that calculates the first evaluation value based on the first numerical value and the second numerical value; Including, The evaluation system of claim 1 .

3. the first evaluation value and the second evaluation value are each a numerical value greater than 0 and equal to or less than 1, the first evaluation value is a statistical value of the first numerical value and the second numerical value, the third evaluation unit outputs the product of the first evaluation value and the second evaluation value as the evaluation result. The evaluation system according to claim 2 .

4. the first index is a risk level indicating the degree of risk of the event indicated by the correct answer sentence, The second indicator is a risk level indicating the degree of risk of the event indicated by the sentence. The evaluation system of claim 1 .

5. the second language model is a large-scale language model; the second evaluation unit, before inputting the sentence into the second language model, inputs into the second language model one or more pairs of an example sentence explaining a traffic-related event and a risk level indicating a degree of risk of the event indicated by the example sentence; The evaluation system according to claim 4 .

6. A database is provided that stores past sentences with different contents that explain the contents of past traffic-related images, The second evaluation unit When the sentence is input to the second language model, past sentences having content similar to the sentence are extracted from the plurality of past sentences; inputting the extracted past sentences together with the sentence into the second language model; The evaluation system of claim 1 .

7. The second evaluation unit inputting, into the second language model, the correct sentence and the sentence, as well as an instruction to evaluate the image from the viewpoint of traffic safety; the second language model outputs, as the second evaluation value, a numerical value indicating a degree of deviation between the first index and the second index calculated based on the instruction; The evaluation system of claim 1 .

8. the traffic-related images are images obtained by capturing images using a camera mounted on an autonomous vehicle; The instruction is an instruction to evaluate from the perspective of a remote monitor of the autonomous vehicle. The evaluation system according to claim 7 .

9. 1. An evaluation program for evaluating performance of a first language model that, when an image related to traffic is input, outputs a sentence describing the content of the image, the evaluation program comprising: On the computer, a first evaluation process for outputting a first evaluation value indicating the accuracy of the sentence based on a predetermined evaluation index; a second evaluation process that uses a second language model different from the first language model, the second language model being constructed by machine learning based on a transportation-related perspective, and outputs a second evaluation value that indicates the accuracy of the sentence; a third evaluation process for outputting an evaluation result of the first language model based on the first evaluation value and the second evaluation value; Execute In the second evaluation process, a correct answer sentence that indicates the content that should be most read from the image and the sentence are input to the second language model; outputting, as the second evaluation value, a numerical value indicating the degree of deviation between the first index and the second index calculated by the second language model based on the content of the correct sentence and the second index based on the content of the sentence; Evaluation program.

10. 1. A method for evaluating performance of a first language model that, when a traffic-related image is input, outputs a sentence describing the content of the image, the method comprising: a first evaluation step in which a computer outputs a first evaluation value indicating the accuracy of the sentence based on a predetermined evaluation index; a second evaluation step in which the computer uses a second language model different from the first language model, the second language model being constructed by machine learning based on a transportation-related perspective, and outputs a second evaluation value indicating accuracy of the sentence; a third evaluation step in which the computer outputs an evaluation result of the first language model based on the first evaluation value and the second evaluation value; Including, In the second evaluation step, a computer inputs a correct answer sentence, which is a sentence indicating the content that should be most read from the image, and the sentence into the second language model; the computer outputs, as the second evaluation value, a numerical value indicating a degree of deviation between the first index and the second index calculated by the second language model based on the content of the correct sentence and the second index based on the content of the sentence; Evaluation method.

Citation Information

Patent Citations

  • Traffic scene description method and system based on coding and decoding network

    CN112911338A

  • Image description information generation method and apparatus, and electronic device

    US20210042579A1

  • Automatically evaluating caption quality of rich media using context learning

    US20210064879A1

  • Image paragraph generator

    US20230394855A1