Rag database construction system, construction method, and construction program

WO2026204904A1PCT designated stage Publication Date: 2026-10-01DENSO CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/011418
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-03-23
Publication Date
2026-10-01

Smart Images

  • Figure JP2026011418_01102026_PF_FP_ABST
    Figure JP2026011418_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A RAG database construction system according to the present invention is provided with a vehicle database, a text generation unit, and a registration unit. The text generation unit identifies a plurality of features of a vehicle data set stored in the vehicle database and combines the plurality of identified features to generate a scene context representing the interpretation of a scene. The registration unit registers the scene context generated by the text generation unit, with an index corresponding to the plurality of features added thereto, in the RAG database.
Need to check novelty before this filing date? Find Prior Art

Description

RAG database construction system, construction method, and construction program Cross-reference to Related Applications

[0001] This international application claims the benefit of Japanese Patent Application No. 2025-050643 filed with the Japan Patent Office on March 25, 2025, the entire disclosure of which is incorporated herein by reference into this international application.

[0002] The present disclosure relates to a technology for constructing a Retrieval-Augmented Generation (RAG) database from data collected from vehicles.

[0003] The in-vehicle device and server described in Patent Document 1 collect vehicle data related to vehicle position information, vehicle operating conditions, and the like. The collected vehicle data is used to provide services such as vehicle diagnosis, maintenance, and driving assistance.

[0004] Japanese Unexamined Patent Publication No. 2023-84379

[0005] RAG is a method that combines an information retrieval function with a large language model (LLM). It retrieves and acquires information related to a question input by a user, and generates an appropriate answer based on the acquired information. Using such RAG, it is conceivable to generate an answer to a user's question based on vehicle data collected by the in-vehicle device or server described above. However, as a result of detailed studies by the inventor, it has been found that RAG is a method designed for text retrieval and is not optimized for retrieving vehicle data.

[0006] According to one aspect of the present disclosure, it is desirable to be able to construct a database suitable for RAG from data collected from vehicles.

[0007] One aspect of the RAG database construction system described herein comprises a vehicle database, a text generation unit, and a registration unit. The vehicle database stores vehicle datasets collected from vehicles and associated with scenes. The text generation unit identifies multiple features of the vehicle datasets stored in the vehicle database and combines the identified features to generate a scene context that represents the interpretation of the scene. The registration unit assigns an index to the scene context generated by the text generation unit according to the multiple features and registers it in the Retrieval-Augmented Generation (RAG) database.

[0008] In the RAG database construction system disclosed herein, scene contexts representing interpretations of scenes associated with vehicle datasets are generated from vehicle datasets. These scene contexts are then indexed according to the multiple features they contain and registered in the RAG database. Therefore, a database suitable for RAG can be constructed from vehicle datasets.

[0009] Another aspect of the present disclosure of a method for constructing a RAG database involves identifying multiple features of a vehicle dataset stored in a vehicle database, which is collected from vehicles and associated with scenes; generating a scene context that represents the interpretation of a scene based on the identified multiple features; assigning an index to the generated scene context according to the multiple features; and registering it in the RAG database.

[0010] The above method for constructing a RAG database will produce the same effect as the above-mentioned RAG database construction system.

[0011] A further aspect of the present disclosure, a program for constructing a database for RAG, causes a processing unit to identify multiple features of a vehicle dataset stored in a vehicle database, which is collected from a vehicle and associated with a scene; generate a scene context representing an interpretation of the scene based on the identified multiple features; and register the generated scene context in the database for RAG, with an index corresponding to the multiple features.

[0012] According to the above program, the same effect as the RAG database construction system described above will be achieved.

[0013] Figure 2A is a block diagram showing the configuration of the RAG construction system according to this embodiment. Figure 2B is a block diagram showing the functional configuration of the in-vehicle device according to this embodiment. Figure 2B is a block diagram showing the configuration of the center device according to this embodiment. Figure 2B is a block diagram showing the functional configuration of the center device according to this embodiment. Figure 3D is a diagram showing the flow from construction to use of the RAG DB according to this embodiment. Figure 4D is a diagram illustrating the scene context generation process according to this embodiment. Figure 5D is a diagram illustrating the overview of the RAG DB construction process according to this embodiment. Figure 6D is a diagram illustrating the scene context analysis process in the RAG DB construction process according to this embodiment. Figure 7D is a diagram illustrating the classification process in the RAG DB construction process according to this embodiment. Figure 8D is a diagram illustrating the hierarchical structuring process in the RAG DB construction process according to this embodiment. Figure 8A is a diagram showing the generalized feature generation process according to this embodiment. Figure 8B is a diagram showing clusters registered in the RAG DB before update according to this embodiment. Figure 8C is a diagram showing the input data according to this embodiment. Figure 8D is a diagram showing clusters registered in the RAG DB after update according to this embodiment. Figure 9A is a diagram showing the process of generating specialized features in the RAG DB according to this embodiment. Figure 9B is a diagram showing the generalized features registered in the RAG DB before update according to this embodiment. Figure 9C is a diagram showing the input data according to this embodiment. Figure 9D is a diagram showing specialized features registered in the updated RAG DB according to this embodiment, distinct from generalized features. Figure 10A is a diagram showing the classification and replacement process of generalized features in the RAG DB according to this embodiment. Figure 10B is a diagram showing specialized features registered in the updated RAG DB according to this embodiment, distinct from generalized features. This is a diagram showing an overview of the construction of the generalized RAG DB by the central device and the construction of the specialized RAG DB by the in-vehicle device according to this embodiment. This is a flowchart showing the scene context generation process according to this embodiment. This is a flowchart showing the basic process of registration to the RAG DB according to this embodiment.This flowchart shows the generalization feature generation process to be registered in the RAG database according to this embodiment. This flowchart shows the specialized feature generation process to be registered in the RAG database according to this embodiment. This flowchart shows the classification process for specialized features to be registered in the RAG database according to this embodiment. This flowchart shows the replacement process for specialized features to be registered in the RAG database according to this embodiment.

[0014] (Embodiment) <1. Configuration> <1-1. System Configuration> Referring to Figures 1 to 3, the Retrieval-Augmented Generation (RAG) database (hereinafter referred to as DB) construction system 1 according to this embodiment will be described. The RAG DB construction system 1 comprises a plurality of in-vehicle devices 2, a plurality of RAG DBs 5, and a central device 3.

[0015] The on-board device 2 is mounted on the vehicle 50 and connected to the RAG DB 5. The RAG DB 5 is a search database for RAG. The on-board device 2 collects vehicle data sets related to the driver of the vehicle 50 and constructs the RAG DB 5 based on the collected vehicle data sets. The on-board device 2 can also communicate data with the central device 3 via a wide-area wireless communication network NW.

[0016] The central device 3 communicates data with multiple in-vehicle devices 2 via a wide-area wireless communication network NW. The central device 3 collects vehicle data sets from the multiple in-vehicle devices 2 and constructs the RAG DB 33c based on the collected vehicle data sets.

[0017] As shown in Figure 2A, the in-vehicle device 2 is, for example, an electronic control unit (hereinafter referred to as ECU). The in-vehicle device 2 comprises a control unit 11, a Controller Area Network (hereinafter referred to as CAN) communication unit 12, a storage unit 13, and a communication unit 14.

[0018] The control unit 11 has a microcomputer equipped with a CPU 21, ROM 22, and RAM 23. The various functions of the control unit 11 are realized by the CPU 21 executing a program stored in a non-transitional physical recording medium. In this embodiment, the ROM 22 corresponds to the non-transitional physical recording medium that stores the program. Furthermore, the execution of this program executes a method corresponding to the program. Note that some or all of the functions realized by the CPU 21 may be realized by one or more hardware such as integrated circuits. The control unit 11 may have one microcomputer or two or more microcomputers.

[0019] The control unit 11 collects vehicle data sets based on the collection condition data received from the central device 3. The control unit 11 then stores the collected vehicle data sets in the vehicle DB 131 and uploads them to the central device 3. The collection condition data includes collection data items indicating the items of data to be collected, collection start conditions for initiating data collection, data sampling conditions, video data trimming conditions, and upload start conditions for the central device 3.

[0020] The CAN communication unit 12 is connected to multiple ECUs via communication lines and transmits and receives data with the multiple ECUs according to the CAN communication protocol. In this embodiment, ECUs 16, 17, and 18 are connected to the CAN communication unit 12. ECU 16 is an engine ECU that controls the engine. ECU 17 is a steering ECU that controls the steering. ECU 18 is a suspension ECU that controls the suspension. In another embodiment, one or two ECUs may be connected to the CAN communication unit 12, or four or more ECUs may be connected. The control targets of the ECUs connected to the CAN communication unit 12 are not particularly limited. In addition, sensors and / or actuators may be connected to the CAN communication unit 12 instead of ECUs, or in addition to ECUs.

[0021] The storage unit 13 is a storage device that stores various types of data. For example, the storage unit 13 may be a hard disk, optical disk, flash memory, or storage device. The storage unit 13 stores the vehicle data set collected by the in-vehicle device 2. The storage unit 13 also has a vehicle database 131 and a vehicle knowledge database 132.

[0022] The vehicle database 131 stores the vehicle dataset collected by the control unit 11. The vehicle dataset contains multiple types of vehicle data. One vehicle dataset is a group of vehicle data corresponding to one scene. A scene includes ambient brightness, vehicle behavior, driver operation, road conditions, application execution status, etc.

[0023] Multiple types of vehicle data are data that differs in nature or characteristics from one another. Multiple types of vehicle data include vehicle data, video data, and application operation data received from the CAN communication unit 12. The data received from the CAN communication unit 12 includes data related to the behavior of the vehicle 50. The data related to the behavior of the vehicle 50 includes vehicle speed, brake on / off, wiper on / off, light on / off, etc. Video data includes video data of the outside of the vehicle, video data of the inside of the vehicle, etc. Application operation data includes navigation operation data, etc.

[0024] The vehicle database 131 stores vehicle datasets linked to the driver ID, collection date and time, and scene ID. The driver ID is an identifier assigned to the driver of vehicle 50, and can identify an individual. If multiple drivers use vehicle 50, each driver is assigned an ID. For example, the control unit 11 performs facial recognition based on image data to identify the driver. The scene ID is an identifier that identifies a scene.

[0025] The vehicle knowledge DB 132 stores interpretation rules for multiple features included in the vehicle dataset. That is, the vehicle knowledge DB 132 stores rules on how to combine and interpret multiple features. In another embodiment, the storage unit 13 does not have to be built into the in-vehicle device 2, but may be an external storage device connected to the in-vehicle device 2 by wire or wireless connection.

[0026] The communication unit 14 communicates wirelessly with the center device 3 via a wide-area wireless communication network NW.

[0027] As shown in Figure 3, the central device 3 comprises a control unit 31, a communication unit 32, and a storage unit 33.

[0028] The control unit 31 has a microcomputer comprising a CPU 41, a ROM 42, and a RAM 43. The various functions of the control unit 31 are realized by the CPU 41 executing a program stored in a non-transitional physical recording medium. In this embodiment, the ROM 42 corresponds to the non-transitional physical recording medium that stores the program. Furthermore, the execution of this program executes a method corresponding to the program. Note that some or all of the functions realized by the CPU 41 may be realized by one or more hardware such as integrated circuits. The control unit 31 may have one microcomputer or may have multiple microcomputers.

[0029] The communication unit 32 performs wireless communication with multiple in-vehicle devices 2 via a wide-area wireless communication network NW.

[0030] The memory unit 33 is a storage device that stores various types of data. For example, the memory unit 33 may be a hard disk, optical disk, flash memory, or storage device. The memory unit 33 includes a collection conditions DB 33a, a collection data DB 33b, a RAG DB 33c, and a vehicle knowledge DB 33d.

[0031] The collection conditions DB33a stores the collection conditions for vehicle data.

[0032] The collected data DB 33b stores vehicle datasets collected from multiple in-vehicle devices 2 by the control unit 31. The collected data DB 33b stores vehicle datasets linked to the vehicle ID, driver ID', collection date and time, and scene ID. The vehicle ID is the identifier for each vehicle 50. The driver ID' is an identifier that allows multiple drivers to be distinguished from each other, but does not allow for the identification of an individual. For example, the driver ID' is a hashed value of the driver ID.

[0033] The RAG database 33c is a search database for RAG and includes the generalized RAG database 331 and the specialized RAG database 332. The vehicle knowledge database 33d, like the vehicle knowledge database 132, stores interpretation rules for multiple features included in the vehicle dataset.

[0034] In this embodiment, the in-vehicle device 2 corresponds to the in-vehicle device of this disclosure, and the center device 3 corresponds to the external device of this disclosure. In addition, in this embodiment, the storage unit 13 and the collected data DB 33b correspond to the vehicle database of this disclosure.

[0035] <1-2. Functional Configuration of the In-Vehicle Device> As shown in Figure 2B, the control unit 11 of the in-vehicle device 2 has the functions of a text generation unit 211, a registration unit 212, and a utilization unit 213, which are executed by the CPU 21. As shown in Figure 5, the text generation unit 211 and the registration unit 212 construct the RAG DB 5, and the utilization unit 213 utilizes the constructed RAG DB 5.

[0036] The text generation unit 211 identifies the feature quantities of each of the multiple types of vehicle data contained in the vehicle dataset stored in the vehicle DB 131. Then, the text generation unit 211 combines the identified feature quantities to generate a scene context. The scene context represents the interpretation of the scene associated with the vehicle dataset. The text generation unit 211 generates a scene context from the vehicle dataset for each scene. Furthermore, the identified feature quantities are specialized feature quantities related to individual drivers, and the generated scene context is a specialized scene context related to individual drivers.

[0037] The registration unit 212 assigns an index for document retrieval to the scene context generated by the text generation unit 211, according to multiple feature quantities. Then, the registration unit 212 vectorizes the indexed scene context and registers it in the RAG DB5.

[0038] When a user query is input to the in-vehicle device 2, the user unit 213 converts the query into a searchable format and searches the RAG DB 5. The user unit 213 then generates and presents answers from documents highly relevant to the query. The user is the driver, the owner of the vehicle 50, the maintenance company for the vehicle 50, etc. In another embodiment, the in-vehicle device 2 does not need to have the functions of the user unit 213. The user unit 213 may be provided in the user's terminal device. The terminal device may be a smartphone, tablet device, PC, wearable device, etc.

[0039] <1-3. Functional Configuration of the Center Device> As shown in Figure 4, the center device 3 comprises a data collection unit 101, a collected data and product storage unit 102, a collected data management and analysis unit 103, a machine learning unit 104, a collection condition management unit 105, a Continuous Integration (CI) / Continuous Delivery (CD) unit 106, a User Interface (UI) unit 107, authentication and authorization units 108, 109, 110, 111, 112, a collection condition file transmission unit 113, an Artificial Intelligence (AI) model file transmission unit 114, a text generation unit 151, a registration unit 152, and a utilization unit 153.

[0040] The data acquisition unit 101 collects vehicle data sets from multiple in-vehicle devices 2.

[0041] The collected data and product storage unit 102 stores the vehicle dataset collected by the data collection unit 101 and the products generated by the collected data management and analysis unit 103, the machine learning unit 104, and the collection condition management unit 105.

[0042] The collected data management and analysis unit 103 includes a preprocessing unit 121 and a collection status confirmation unit 122. The preprocessing unit 121 performs preprocessing on the collected vehicle dataset to make it easier to use in machine learning (for example, changing the resolution, changing to grayscale, etc.). The collection status confirmation unit 122 confirms the amount of data in the collected vehicle dataset.

[0043] The machine learning unit 104 includes an annotation unit 123, a model training unit 124, and a model evaluation unit 125. The annotation unit 123 performs annotation to add information (e.g., labels) for machine learning to the data preprocessed by the preprocessing unit 121. The model training unit 124 trains a machine learning model using the data generated by the annotation unit 123. The machine learning model is, for example, a model that takes image data captured by an in-vehicle camera as input data, determines an object in the image, and outputs a determination result. The model evaluation unit 125 evaluates the accuracy of the machine learning model based on the training result obtained by the model training unit 124.

[0044] The collection condition management unit 105 manages collection conditions used for collecting data to be used for training a machine learning model. The collection condition management unit 105 includes a collection condition generation unit 126. The collection condition generation unit 126 generates collection condition data indicating collection conditions for collecting data for improving the accuracy of the machine learning model based on the evaluation result obtained by the model evaluation unit 125. The collection condition data may be automatically generated by an AI model, or may be manually generated by an engineer.

[0045] The CI / CD unit 106 includes a collection condition building and distribution unit 127, an AI logic building and distribution unit 128, and a distribution status management unit 129. The collection condition building and distribution unit 127 converts the collection condition data generated by the collection condition generation unit 126 or an engineer into a format recognizable by the in-vehicle device 2, and distributes the converted data to the in-vehicle device 2. The AI logic building and distribution unit 128 converts the machine learning model (i.e., AI logic) generated by the machine learning unit 104 into a format executable by the in-vehicle device 2, and distributes the converted model to the in-vehicle device 2. The distribution status management unit 129 checks the distribution status of the collection condition data and the AI logic.

[0046] The UI unit 107 is a device for exchanging information between data engineers, AI engineers, software engineers, test engineers and operators and the center device 3. A data engineer is an engineer who works with data used for machine learning. An AI engineer is an engineer engaged in the generation of machine learning models. A software engineer is an engineer that works with applications generated by combining a plurality of machine learning models. A test engineer is an engineer that evaluates machine learning models and applications. The operator manages the collection condition data and the distribution status of AI logic.

[0047] The authentication / authorization unit 108 performs authentication and authorization on access to the center device 3 by a data engineer. For example, the authentication / authorization unit 108 performs authentication and authorization such that the data engineer can only access the collected data management / analysis unit 103.

[0048] The authentication / authorization unit 109 performs authentication and authorization on access to the center device 3 by an AI engineer. For example, the authentication / authorization unit 109 performs authentication and authorization such that the AI engineer can only access the machine learning unit 104.

[0049] The authentication / authorization unit 110 performs authentication and authorization on access to the center device 3 by a software engineer. For example, the authentication / authorization unit 110 performs authentication and authorization such that the software engineer can only access the machine learning unit 104.

[0050] The authentication / authorization unit 111 performs authentication and authorization on access to the center device 3 by a test engineer. For example, the authentication / authorization unit 111 performs authentication and authorization such that the test engineer can only access the collection condition management unit 105.

[0051] The authentication / authorization unit 112 performs authentication and authorization on access to the center device 3 by an operator. For example, the authentication / authorization unit 112 performs authentication and authorization such that the operator can only access the CI / CD unit 106.

[0052] The collection condition file transmission unit 113 transmits the collection condition file generated by the collection condition build and distribution unit 127 to multiple in-vehicle devices 2.

[0053] The AI ​​model file transmission unit 114 transmits the AI ​​model file generated by the AI ​​logic build and distribution unit 128 to multiple in-vehicle devices 2.

[0054] The text generation unit 151, like the text generation unit 211, identifies multiple features from the vehicle dataset stored in the collected data DB 33b and combines the identified features to generate a scene context.

[0055] More specifically, the text generation unit 151 identifies multiple generalization features from a vehicle dataset that corresponds to the generalization criterion and generates a generalized scene context by combining these multiple generalization features. In addition, the text generation unit 151 identifies multiple specialized features from a vehicle dataset that corresponds to the specialization criterion and generates a specialized scene context by combining these multiple specialized features.

[0056] Vehicle datasets that meet the generalization criteria correspond to multiple in-vehicle devices 2 or multiple drivers. Vehicle datasets that meet the specialization criteria correspond to one driver or a driver with specific attributes. Specific attributes include being under a certain age, being above a certain age, or having less than a certain number of years of driving experience.

[0057] The registration unit 152 indexes the generalized scene context generated by the text generation unit 211, vectorizes it, and registers it in the generalized RAG DB 331. The registration unit 152 also indexes the specialized scene context generated by the text generation unit 211, vectorizes it, and registers it in the specialized RAG DB 332.

[0058] When a user query is input via the in-vehicle device 2 or terminal device, the user unit 153 determines whether the query relates to generalized features or specialized features. If the query relates to generalized features, the user unit 153 extracts and reconstructs the portion of the query that relates to generalized features and searches the generalized RAG DB 331. The user unit 153 then generates an answer related to the generalized features and outputs the generated answer to the in-vehicle device 2 or terminal device.

[0059] Furthermore, if the query is related to a specialized feature, the utilization unit 153 extracts and reconstructs the portion of the query that is related to the specialized feature and searches the specialized RAG DB 332. The utilization unit 153 then generates an answer regarding the specialized feature and outputs the generated answer to the in-vehicle device 2 or the user's terminal device.

[0060] In another embodiment, at least one of the following may be removed from the center device 3: data acquisition unit 101, acquired data / product storage unit 102, acquired data management / analysis unit 103, machine learning unit 104, acquisition condition management unit 105, CI / CD unit 106, UI unit 107, authentication / authorization units 108, 109, 110, 111, 112, acquisition condition file transmission unit 113, AI model file transmission unit 114, and text generation unit 151, registration unit 152, and utilization unit 153.

[0061] <2. Processing> <2-1. Scene Context Generation Processing> The scene context generation processing performed by the text generation unit 151 will be explained with reference to the flowchart in Figure 12. The text generation unit 151 repeatedly performs this processing at predetermined intervals. The text generation unit 211 also performs the same processing as the text generation unit 151.

[0062] In S10, the text generation unit 151 receives a vehicle dataset associated with the scene ID. Specifically, as shown in Figure 6, the text generation unit 151 receives motion image data, CAN data, application operation data, etc., associated with the same scene ID.

[0063] Next, in S20, the text generation unit 151 identifies the feature quantities of each of the multiple types of vehicle data included in the vehicle dataset. As shown in Figure 6, for example, the feature quantities identified from the video data include brightness, the number of people or objects such as signs, the position of people or objects such as signs, the size of people or objects such as signs, the movement of objects, and the type of background (specifically, urban areas, highways, suburbs, etc.). For example, the feature quantities identified from the CAN data include whether the vehicle speed is low, medium, or high, whether it is accelerating, decelerating, or constant speed, the wiper speed, and the timing of vehicle operation.

[0064] Next, in S30, the text generation unit 151 obtains interpretation rules from the vehicle knowledge DB 132 and combines the multiple feature quantities identified in S20 based on the interpretation rules.

[0065] Next, in S40, the text generation unit 151 generates a scene context from a combination of multiple features using a language model (for example, a large-scale language model). Based on the interpretation rules of the vehicle knowledge DB 132, the text generation unit 151 identifies features corresponding to the cause of the scene and features corresponding to the result of the scene, and generates a scene context that expresses the relationship between the cause of the scene and the result of the scene. One example of a scene context is, "In the suburbs, it was bright enough and the vehicle speed was low, so the wipers were stopped." Another example of a scene context is, "It was raining lightly, but the vehicle speed increased, so the wiper speed was changed to medium." Yet another example of a scene context is, "The acceleration and deceleration were intense, and the driver was tired because pedestrians were being overlooked in the video." The text generation unit 151 outputs the generated scene context and the multiple features included in the scene context to the registration unit 152.

[0066] <2-2. Basic Processing for Registration to the RAG Database> Next, referring to the flowchart in Figure 13, the basic processing for registration to the RAG database 33c performed by the registration unit 152 will be explained. The registration unit 152 performs this processing when no data has been registered in the RAG database 33c. The registration unit 212 performs the same processing as the registration unit 152.

[0067] In S100, the registration unit 152 receives the scene context and multiple feature quantities generated by the text generation unit 151. As shown in Figure 7A, in the following steps S110 to S170, the registration unit 152 performs scene context analysis, classification using a machine learning model, hierarchical structuring, and vectorization on the received scene context.

[0068] In S110, the registration unit 152 uses a language model to extract the cause and effect parts from the scene context. For example, as shown in Figure 7B, from scene context A: "In the suburbs, it was bright enough and the vehicle speed had slowed down, so the wipers were stopped," the result part: "The wipers were stopped" is extracted. The registration unit 152 also extracts the cause parts: "In the suburbs," "It was bright," and "The vehicle speed had slowed down."

[0069] Next, in S120, the registration unit 152 sets the feature quantity corresponding to the result portion as the top event. For example, for scene context A, the registration unit 152 sets "wiper operation occurred → stopped" as the top event.

[0070] Next, in S130, the registration unit 152 determines which feature corresponds to the cause part and associates the value of the feature corresponding to the cause part. For example, for scene context A, the registration unit 152 determines that the feature corresponding to the cause: "in the suburbs" is "location" and associates its value "suburbs". It also determines that the feature corresponding to the cause: "bright" is "brightness" and associates its value "126". Furthermore, it determines that the feature corresponding to the cause: "vehicle speed also became slow" is "vehicle speed" and associates its value "30 km / h".

[0071] Next, in S140, the registration unit 152 classifies multiple features using a machine learning method. For example, as shown in Figure 7C, the registration unit 152 uses a random forest to classify whether each feature corresponds to a top event or not. The registration unit 152 determines that three features, "brightness: 126", "location: suburban", and "vehicle speed: 30 km / h", correspond to top events for scene context A.

[0072] Next, in S150, the registration unit 152 calculates the importance of each feature corresponding to the top event. Specifically, the registration unit 152 calculates the importance of each feature using the library function of the random forest.

[0073] Next, in S160, the registration unit 152 arranges multiple feature quantities corresponding to the top event in a hierarchical structure in order of importance. For example, as shown in Figure 7D, the feature quantity with high importance, "vehicle speed: 30 km / h or less," is placed below the top event. Further below that, the feature quantities with low importance, "brightness: 126 or more" and "location: suburban," are placed. The registration unit 152 then expresses the determined hierarchical structure in text format. In this embodiment, one hierarchical structure with the top event as the top layer is called a cluster, and the second layer and below are called subclusters.

[0074] Next, in S170, as shown in Figure 7E, the registration unit 152 assigns indices to the hierarchically structured text according to its features. For example, the top-level feature is assigned index "1", the second-level feature is assigned index "1-1", and the third-level feature is assigned indices "1-1-1" and "1-1-2". The registration unit 152 then vectorizes the hierarchically structured text using an existing vector model, etc., and registers it in the generalized RAG DB 331 and the specialized RAG DB 332, respectively. The vectorized text is represented, for example, in JavaScript Object Notation format.

[0075] <2-3. Generating Generalized Features> Next, the generalized feature generation process executed by the registration unit 152 will be explained with reference to the flowchart in Figure 14. This process is one of the update processes for the generalized RAG DB 331, and the registration unit 152 executes this process when data is registered in the generalized RAG DB 331 and a scene context is generated by the text generation unit 151.

[0076] In S200, the registration unit 152 receives one or more scene contexts generated by the text generation unit 151.

[0077] Next, in S210, the registration unit 152 executes the processes in S100 to S170 to hierarchically vectorize the scene context. As shown in Figure 8A, in the following processes S220 to S290, the registration unit 152 updates the existing cluster by averaging the features of the input data and the existing cluster when the similarity between the vectorized input data and the existing cluster registered in the generalized RAG DB 331 is high.

[0078] In S220, the registration unit 152 calculates the similarity between the input data and existing clusters registered in the generalized RAG DB 331. If there are multiple existing clusters, the similarity between the input data and each existing cluster is calculated. Specifically, the registration unit 152 calculates the cosine similarity between the input data vector and the existing cluster vector. For example, the registration unit 152 calculates the cosine similarity between the input data vector shown in Figure 8C and the existing cluster vector shown in Figure 8B. Note that upper layer features contribute more to the similarity than lower layer features. For example, if the second layer features of the input data and the existing cluster are different, or if the difference in the values ​​of the second layer features is large, the similarity will be calculated to be low. If the second layer features of the input data and the existing cluster are the same and the feature values ​​are close, the similarity will be calculated to be high even if the third layer features are different, or if the difference in the values ​​of the third layer features is large. If the first layer (i.e., the top event) is different, the similarity will be calculated to be very low. In another embodiment, the registration unit 152 may calculate, instead of cosine similarity, the distance between the two vectors (e.g., Euclidean distance, Manhattan distance, etc.) or the correlation coefficient (e.g., Person correlation coefficient, etc.) as a measure of similarity.

[0079] Next, in S230, the registration unit 152 determines whether any of the similarities calculated in S220 are greater than the similarity threshold Sth. If the registration unit 152 determines that any of the similarities are greater than the set similarity threshold Sth, it proceeds to the process in S240. If it determines that all similarities are less than or equal to the similarity threshold Sth, it proceeds to the process in S260.

[0080] In S240, the registration unit 152 calculates the average value of the input data and the features of the existing cluster and updates the existing cluster. If there are multiple existing clusters corresponding to similarity exceeding the similarity threshold Sth, the input data and each of the corresponding existing clusters are averaged and each of the corresponding existing clusters is updated. Specifically, the registration unit 152 averages the values ​​of the features at the same position in the same hierarchical level of the input data and the existing cluster (i.e., features assigned the same index). For example, the registration unit 152 averages the "vehicle speed: 30 km / h or less" of the second layer of the input data shown in Figure 8C with the "vehicle speed: 40 km / h or less" of the second layer of the existing cluster shown in Figure 8B. Then, as shown in Figure 8D, the registration unit 152 updates the second layer of the existing cluster to "vehicle speed: 35 km / h or less". The left "brightness: 126 or more" of the third layer of the input data shown in Figure 8C is the same as the left "brightness" of the third layer of the existing cluster shown in Figure 8B. Therefore, as shown in Figure 8D, the left side of the third layer of the updated existing cluster is the same as the left side of the third layer of the pre-update existing cluster.

[0081] Next, in S250, the registration unit 152 adds a new element (in other words, a subcluster) to the existing cluster's hierarchical structure if the input data's hierarchical structure includes a new element that does not exist in the existing cluster's hierarchical structure. The right side of the third layer of the new cluster shown in Figure 8C does not exist in the existing cluster shown in Figure 8B. Therefore, as shown in Figure 8D, the registration unit 152 updates the existing cluster by adding "Location: Suburbs" to the right side of the third layer of the existing cluster.

[0082] In S260, the registration unit 152 registers the input data as a new cluster in the generalized RAG DB 331 because the input data is not similar to an existing cluster.

[0083] Next, in S270, the registration unit 152 determines whether or not there are any unprocessed scene contexts among the scene contexts received in S200. That is, the registration unit 152 determines whether or not there are any scene contexts for which the processing in S210 to S260 has not been executed. If the registration unit 152 determines that there are any unprocessed scene contexts, it returns to the processing in S210 and repeats the processing in S210 to S260. If the registration unit 152 determines that there are no unprocessed scene contexts, it proceeds to the processing in S280.

[0084] Next, in S280, the registration unit 152 determines whether or not clusters corresponding to only a small number of drivers exist in the generalized RAG DB 331. If the registration unit 152 determines that such clusters do not exist, it terminates this process. If the registration unit 152 determines that such clusters exist, it proceeds to the process in S290.

[0085] Next, in S290, the registration unit 152 deletes clusters corresponding to only a small number of drivers from the generalized RAG DB 331. Specifically, the registration unit 152 deletes clusters whose top layer consists of individual vehicle datasets for drivers with fewer than a certain number of drivers. The registration unit 152 also deletes clusters whose lower layers consist of individual vehicle datasets for drivers with fewer than a certain number of drivers. As the processing in S210 to S260 is repeated, the existing clusters become generalized clusters composed of generalized features corresponding to the average trends of a large number of drivers. On the other hand, clusters corresponding to only a small number of drivers are generated, so these clusters are deleted.

[0086] <2-4. Specialized Feature Generation Process> Next, the specialized feature generation process executed by the registration unit 152 will be explained with reference to the flowchart in Figure 15. This process is one of the update processes for the specialized RAG DB 332. The registration unit 152 executes this process when data has been registered in the specialized RAG DB 332 and a scene context has been generated by the text generation unit 151.

[0087] In S300, the registration unit 152 receives one or more scene contexts generated by the text generation unit 151.

[0088] Next, in S310, the registration unit 152 executes the processes in S100 to S170 to hierarchically vectorize the scene context. As shown in Figure 9A, in the following processes S320 to S370, if the similarity between the vectorized input data and an existing cluster registered in the specialized RAG DB 332 is high, the registration unit 152 averages the features of the input data and the features of the existing cluster to update the existing cluster.

[0089] In S320, the registration unit 152 calculates the generalization similarity and the specialized similarity. The generalization similarity is the similarity between the input data and each of the generalization clusters registered in the generalization RAG DB 331. The specialized similarity is the similarity between the input data and each of the specialized clusters registered in the specialized RAG DB 332. The registration unit 152 calculates the generalization similarity and the specialized similarity in the same way as in the processing in S220.

[0090] Next, in S330, the registration unit 152 determines whether any of the generalized similarities calculated in S320 are greater than the similarity threshold Sth. For example, if the generalized cluster shown in Figure 9B is registered and the input data shown in Figure 9C is input, the calculated generalized similarities are less than or equal to the similarity threshold Sth. If the registration unit 152 determines that any of the generalized similarities are greater than the similarity threshold Sth, it proceeds to the process in S340. If the registration unit 152 determines that all generalized similarities are less than or equal to the similarity threshold Sth, it proceeds to the process in S350.

[0091] In S340, the registration unit 152 decides not to register the new cluster in the specialized RAG DB 332 and terminates this process.

[0092] In S350, the registration unit 152 determines whether any of the specialized similarity values ​​calculated in S320 are greater than the similarity threshold Sth. For example, if the specialized cluster shown in Figure 9D is registered and the input data shown in Figure 9C is input, the calculated specialized similarity is greater than the similarity threshold Sth. If the registration unit 152 determines that any of the specialized similarity values ​​are greater than the similarity threshold Sth, it proceeds to the process in S360. If the registration unit 152 determines that all specialized similarity values ​​are less than or equal to the similarity threshold Sth, it proceeds to the process in S380. Note that the similarity threshold Sth used in S220, S330, and S350 may all be the same value or may be different.

[0093] In S360, the registration unit 152, similar to S240, calculates the average value of the feature quantities at the same hierarchical level and position in the input data and the specialized cluster, and updates the specialized cluster.

[0094] Next, in S370, the registration unit 152, similar to S250, adds a new element to the hierarchical structure of the specialized cluster if the hierarchical structure of the input data contains a new element that does not exist in the hierarchical structure of the specialized cluster.

[0095] In S380, the registration unit 152 newly registers the input data as a new specialized cluster in the specialized RAG DB 332. As a result, multiple specialized clusters composed of specialized features are registered in the specialized RAG DB 332.

[0096] The registration unit 212 also updates the RAG DB5 by executing the processes in S300-S320 and S350-S380. In the process in S320, the registration unit 212 calculates only the specialization similarity.

[0097] <2-5. Classification Processing of Specialized Features> Next, the classification processing of specialized features performed by the registration unit 152 will be explained with reference to the flowchart in Figure 16. This processing is one of the update processing for the specialized RAG DB 332. The registration unit 152 executes this processing when data is registered in the specialized RAG DB 332 and a scene context is generated by the text generation unit 151. The registration unit 212 also executes this processing in the same way as the registration unit 152.

[0098] In S400, the registration unit 152 receives one or more scene contexts generated by the text generation unit 151.

[0099] Next, in S410, the registration unit 152 executes the processes in S100 to S170 to hierarchically vectorize the scene context. As shown in Figure 10A, in the following processes S420 to S450, the registration unit 152 classifies the specialized features into long-term features and short-term features based on their similarity to the specialized features already registered in the specialized RAG DB 332 and the input history of the specialized features.

[0100] In S420, the registration unit 152 checks the input history of the corresponding cluster registered in the specialized RAG DB 332. The corresponding cluster is a specialized cluster that is the same cluster as the vectorized input data and has the same hierarchical structure as the new cluster. A cluster that is the same as the input data is a cluster whose upper layer is composed of the same features as the input data. The input history is the number of times the corresponding cluster has been input from the text generation unit 151 to the registration unit 152 during a predetermined period.

[0101] Next, in S430, the registration unit 152 determines, based on the input history confirmed in S420, whether or not the specialized features included in the corresponding cluster have been input at a high frequency recently. That is, the registration unit 152 determines whether or not the specialized features included in the corresponding cluster have been input at a predetermined frequency exceeding a predetermined frequency over a predetermined period. If the registration unit 152 determines that the specialized features included in the corresponding cluster have been input at a high frequency recently, it proceeds to the process in S440. If the registration unit 152 determines that the specialized features included in the corresponding cluster have not been input at a high frequency recently, it proceeds to the process in S450.

[0102] In S450, as shown in Figure 10B, the registration unit 152 assigns an index of long-term features to the corresponding clusters. The registration unit 152 indexes the corresponding clusters that have been input frequently in recent times, considering them to be long-term features that do not change much due to the individual driver's personality or habits.

[0103] In S460, as shown in Figure 10B, the registration unit 152 assigns an index of short-term features to the corresponding clusters. The registration unit 152 indexes corresponding clusters that have not been input frequently recently, considering them to be short-term features with little correlation to the driver's individual personality or habits. In other words, the registration unit 152 indexes corresponding clusters, considering them to be features that occur only when specific conditions coincide.

[0104] <2-6. Specialized Feature Replacement Process> Next, the specialized feature replacement process performed by the registration unit 152 will be explained with reference to the flowchart in Figure 17. This process is one of the update processes for the specialized RAG DB 332. The registration unit 152 executes this process when data is registered in the specialized RAG DB 332 and a scene context is generated by the text generation unit 151. The registration unit 212 also executes this process in the same way as the registration unit 152.

[0105] In S500, the registration unit 152 receives one or more scene contexts generated by the text generation unit 151.

[0106] Next, in S510, the registration unit 152 executes the processes from S100 to S170 to create a hierarchical structure and vectorize the scene context. As shown in Figure 10A, the specialized features are replaced based on the similarity of the specialized features registered in the specialized RAG DB 332 and the input history of the specialized features.

[0107] Next, in S520, the registration unit 152 checks the input history of similar clusters registered in the specialized RAG DB 332. Similar clusters are specialized clusters that are the same cluster as the vectorized input data and have a similar hierarchical structure to the input data. Similar hierarchical structures have the same upper-level structure but different lower-level structures (i.e., the number of subclusters that make up the lower level).

[0108] Next, in S530, the registration unit 152 determines, based on the input history confirmed in S520, whether or not similar clusters have been entered frequently recently. If the registration unit 152 determines that similar clusters have been entered frequently recently, it proceeds to the process in S540. If the registration unit 152 determines that similar clusters have not been entered frequently recently, it proceeds to the process in S550.

[0109] In S540, the registration unit 152 assumes that the specialized features included in the similar cluster have changed and replaces the specialized features of the similar cluster with the specialized features of the input data.

[0110] Next, in S550, the registration unit 152 calculates the similarity between the input data and the generalized clusters registered in the generalized RAG DB 331, similar to S220.

[0111] Next, in S560, the registration unit 152 determines whether the similarity calculated in S550 is greater than the similarity threshold Sth. If the registration unit 152 determines that the similarity is greater than the similarity threshold Sth, it proceeds to the process in S570. If the registration unit 152 determines that the similarity is less than or equal to the similarity threshold Sth, it terminates this process. The similarity threshold Sth may be the same value as the similarity threshold Sth used in S230, or it may be different.

[0112] In S570, the registration unit 152 assumes that the specialized features have changed into generalized features and deletes similar clusters from the specialized RAG DB 332.

[0113] As shown in Figure 11, the central device 3 does not have a specialized RAG DB 332, and may only construct a generalized RAG DB 331. That is, the central device 3 may construct a RAG DB 33c in which generalized features are registered from a vehicle dataset corresponding to multiple drivers, and the in-vehicle device 2 may construct a RAG DB 5 in which specialized features are registered from a vehicle dataset corresponding to an individual driver. Alternatively, the central device 3 may construct a generalized RAG DB 331 and / or a specialized RAG DB 332, and the in-vehicle device 2 may not construct a RAG DB. Alternatively, the in-vehicle device 2 may construct a RAG DB 5, and the central device 3 may not construct a RAG DB.

[0114] <3. Effects> The embodiment described in detail above provides the following effects.

[0115] (1) The text generation units 151 and 211 generate scene contexts from the vehicle dataset that represent the interpretation of scenes associated with the vehicle dataset. Then, the registration units 152 and 212 assign indices to the scene contexts according to the multiple features contained in the scene contexts and register them in the RAG DB 33c and 5. Thus, it is possible to construct RAG-compatible RAG DB 33c and 5 from the vehicle dataset.

[0116] (2) The text generation units 151 and 211 can generate a scene context by combining the respective feature quantities of multiple types of vehicle data.

[0117] (3) The text generation units 151 and 211 can generate an appropriate scene context by using the interpretation rules stored in the vehicle knowledge databases 33d and 132.

[0118] (4) The text generation units 151 and 211 can generate a scene context by appropriately combining multiple feature quantities according to interpretation rules.

[0119] (5) The text generation unit 151 can generate a scene context that has a causal relationship.

[0120] (6) The registration units 152 and 212 assign an index to the scene context according to the clustering result and register it in the RAG DB 33c and 5. This makes it possible to construct an RAG DB 33c and 5 that can be searched efficiently.

[0121] (7) The registration units 152 and 212 cluster the scene context based on the feature quantities corresponding to the results. This makes it possible to construct RAG DBs 33c and 5 that can be efficiently searched.

[0122] (8) The registration units 152 and 212 determine the hierarchical structure of the cluster based on the feature quantities corresponding to the causes. This makes it possible to construct RAG databases 33c and 5 that can be searched more efficiently.

[0123] (9) The registration units 152 and 212 place the feature quantities with higher importance higher in the hierarchical structure. This makes it possible to construct RAG databases 33c and 5 that can be searched more efficiently.

[0124] (10) The registration unit 152 generates a generalized cluster composed of generalized features from a vehicle dataset corresponding to multiple drivers and registers it in the generalized RAG DB 331. This makes it possible to construct a generalized RAG DB 331 that is suitable for multiple drivers.

[0125] (11) The registration unit 152 generates specialized clusters composed of specialized features from the vehicle dataset that meets the specialized criteria and registers them in the specialized RAG DB 332. This makes it possible to construct a RAG DB that is tailored to the individual preferences of the driver.

[0126] (12) If the query received from the user is related to generalized features, the user unit 153 searches the generalized RAG DB 331, and if the query is related to specialized features, it searches the specialized RAG DB 332. This allows the user unit 153 to search the RAG DB 33c more efficiently.

[0127] (13) If a vehicle dataset conforming to the specialization criteria corresponds to one driver, the text generation units 151, 211 and the registration units 152, 212 can construct a specialized RAG DB 332 that is adapted to the individual driver.

[0128] (14) If a vehicle dataset conforming to the specialization criteria corresponds to a driver with specific attributes, the text generation units 151, 211 and the registration units 152, 212 can construct a specialized RAG DB 332 conforming to the specific attributes.

[0129] (15) The registration unit 152 assigns a short-term feature index to specialized features whose degree of change is greater than the standard degree, and assigns a long-term feature index to specialized features whose degree of change is less than the standard degree. As a result, specialized features corresponding to individual characteristics that occur under specific conditions are distinguished from specialized features corresponding to individual personality and habits, and registered in the specialized RAG DB 332. This makes it possible to construct a specialized RAG DB 332 that can be searched more efficiently.

[0130] (16) When the specialized feature quantities registered in the specialized RAG DB 332 approach the generalized feature quantities registered in the generalized RAG DB 331, the registration unit 152 deletes those specialized feature quantities from the specialized RAG DB 332. This allows the specialized RAG DB 332 to be adapted to changes in the driver's preferences.

[0131] (17) When the central device 3 constructs only the generalized RAG DB 331 and the in-vehicle device 2 constructs the RAG DB 5 in which specialized features are registered, the privacy of the individual driver can be protected.

[0132] (Other Embodiments) Although embodiments of the present disclosure have been described above, the present disclosure is not limited to the embodiments described above and can be implemented in various modified forms.

[0133] (a) The control units 11, 31 and their methods described herein may be implemented by a dedicated computer provided by configuring a processor and memory programmed to perform one or more functions embodied by a computer program. Alternatively, the control units 11, 31 and their methods described herein may be implemented by a dedicated computer provided by configuring a processor by one or more dedicated hardware logic circuits. Alternatively, the control units 11, 31 and their methods described herein may be implemented by one or more dedicated computers configured by a combination of a processor and memory programmed to perform one or more functions and a processor configured by one or more hardware logic circuits. Furthermore, the computer program may be stored as instructions executed by the computer on a computer-readable non-transitional substantial recording medium. The methods for implementing the functions of each part included in the control units 11, 31 do not necessarily need to include software, and all of their functions may be implemented using one or more hardware components.

[0134] (b) Multiple functions of one component in the above embodiment may be realized by multiple components, or one function of one component may be realized by multiple components. Also, multiple functions of multiple components may be realized by one component, or one function realized by multiple components may be realized by one component. Furthermore, some of the configurations of the above embodiment may be omitted. Furthermore, at least some of the configurations of the above embodiment may be added to or replaced with the configurations of other above embodiments. [Technical Concept Disclosed in This Specification] [Item 1] A database construction system for RAG, comprising: a vehicle database (131, 33b) storing vehicle datasets collected from a vehicle (50) and associated with a scene; a text generation unit (211, 151) configured to identify a plurality of features of the vehicle dataset stored in the vehicle database and to generate a scene context representing an interpretation of the scene by combining the identified plurality of features; and a registration unit (212, 152) configured to assign an index corresponding to the plurality of features to the scene context generated by the text generation unit and register it in a database for Retrieval-Augmented Generation (RAG) (5, 33c). [Item 2] The RAG database construction system according to Item 1, wherein the vehicle dataset includes a plurality of different types of vehicle data, and the text generation unit (211, 151) is configured to identify the features of each of the plurality of types of vehicle data. [Item 3] The RAG database construction system according to Item 1 or 2, wherein the text generation unit (211, 151) is configured to generate the scene context based on the interpretation rules input from the vehicle knowledge database (132, 33d) which stores the interpretation rules for the plurality of features, and the vehicle dataset input from the vehicle database. [Item 4] The RAG database construction system according to Item 3, wherein the text generation unit (211, 151) is configured to combine the plurality of features according to the interpretation rules.[Item 5] The RAG database construction system according to Item 4, wherein the text generation unit (211, 151) is configured to identify a feature corresponding to the cause of the scene and a feature corresponding to the result of the scene from among the plurality of feature quantities, and to generate the scene context in which the relationship between the cause and the result is expressed. [Item 6] The RAG database construction system according to any one of Items 1 to 5, wherein the registration unit (212, 152) is configured to cluster the scene context generated by the text generation unit (211, 151), assign an index to the scene context according to the result of the clustering, and register it in the RAG database (5, 33c). [Item 7] The RAG database construction system according to Item 6, wherein the registration unit (212, 152) is configured to identify a feature corresponding to the result of the scene from among the plurality of feature quantities of the scene context, and to cluster the scene context based on the feature corresponding to the result of the scene. [Item 8] The RAG database construction system according to Item 6 or 7, wherein the registration unit (212, 152) is configured to identify a feature corresponding to the cause of the scene from the plurality of feature quantities of the scene context, and to determine the hierarchical structure of the cluster based on the feature corresponding to the cause of the scene. [Item 9] The RAG database construction system according to Item 8, wherein the registration unit (212, 152) is configured to calculate the importance of the feature corresponding to the cause of the scene, and to place the feature with higher importance higher in the hierarchical structure.[Item 10] The RAG database construction system according to any one of items 1 to 9, wherein the vehicle database (33b) stores multiple vehicle datasets corresponding to multiple drivers in a manner that distinguishes each of the multiple drivers, the RAG database (33c) includes a generalized RAG database (331), the text generation unit (151) is configured to generate multiple scene contexts from the multiple vehicle datasets, and the registration unit (152) is configured to cluster the multiple scene contexts, exclude scene contexts belonging to clusters that satisfy the exclusion condition from the multiple scene contexts, and register them in the generalized RAG database (331). [Item 11] The RAG database construction system according to item 10, wherein the exclusion condition is that the top layer of the cluster hierarchical structure consists of the vehicle datasets of each of the multiple drivers that are less than a certain number. [Item 12] The RAG database construction system according to Item 10, wherein the outlier condition is that the lower layers of the cluster's hierarchical structure consist of the vehicle datasets of drivers fewer than a certain number of drivers among the multiple drivers. [Item 13] The RAG database construction system according to any one of Items 10 to 12, wherein the RAG database (33c) further includes a specialized RAG database (332), the text generation unit (151) is configured to identify multiple features of a specialized dataset which is the vehicle dataset corresponding to the specialization criterion, and to combine the multiple features of the specialized dataset to generate a specialized scene context, and the registration unit (152) is configured to register the specialized scene context in the specialized RAG database with an index corresponding to the multiple features of the specialized dataset when the similarity between the multiple features of the specialized dataset and the multiple features registered in the generalized RAG database (331) is less than a certain value.[Item 14] The RAG database construction system according to Item 13, further comprising a user unit (153) configured to receive queries from a user and to search the generalized RAG database or the specialized RAG database based on the received queries. [Item 15] The RAG database construction system according to Item 13 or 14, wherein the specialized dataset corresponds to a single driver. [Item 16] The RAG database construction system according to Item 13 or 14, wherein the specialized dataset corresponds to a driver with specific attributes. [Item 17] The text generation unit (151) is configured to repeatedly acquire the specialized dataset and generate the specialized scene context from the specialized dataset each time the specialized dataset is acquired, and the registration unit (152) is configured to (i) update the specialized RAG database (332) each time the specialized scene context is generated by the text generation unit, and (ii) assign a first index to the feature quantities among the plurality of feature quantities registered in the specialized RAG database whose degree of change is greater than a reference degree, and / or assign a second index to the feature quantities whose degree of change is less than the reference degree, as described in any one of items 13 to 16. [Item 18] The RAG database construction system according to Item 17, wherein the registration unit (152) is configured to delete from the specialized RAG database (332) any feature from the specialized RAG database whose similarity to any of the feature from the generalized RAG database (331) exceeds a threshold value, as the specialized RAG database (332) is repeatedly updated.[Item 19] The vehicle database (131, 33b) includes a first vehicle database (131) and a second vehicle database (33b), the text generation unit (211, 151) includes a first text generation unit (211) and a second text generation unit (151), the registration unit (212, 152) includes a first registration unit (212) and a second registration unit (152), the vehicle (50) is equipped with an on-board device (2) having the first text generation unit, the first registration unit and the specialized RAG database (5), and an external device (3) having the second text generation unit, the second registration unit and the generalized RAG database (331), the first vehicle database stores the vehicle dataset corresponding to the driver of the vehicle, A RAG database construction system according to any one of items 13 to 18, wherein the first text generation unit is configured to generate the scene context from the vehicle dataset stored in the first vehicle database, the first registration unit is configured to register the scene context generated by the first text generation unit in the specialized RAG database, the second vehicle database stores multiple vehicle datasets corresponding to multiple drivers in a manner that distinguishes each of the multiple drivers, the second text generation unit is configured to generate multiple scene contexts from the multiple vehicle datasets stored in the second vehicle database, and the second registration unit is configured to register the multiple scene contexts generated by the second text generation unit in the generalized RAG database.[Item 20] The RAG database construction system according to Item 19, wherein the external device (3) is configured to repeatedly execute a cycle comprising: (i) machine learning a predetermined application using the plurality of vehicle datasets stored in the second vehicle database (33b); (ii) evaluating the learning results; (iii) collecting the plurality of vehicle datasets based on the evaluation of the learning results; and (iv) storing the collected plurality of vehicle datasets in the second vehicle database. [Item 21] A method for constructing a RAG database, comprising: identifying a plurality of features of a vehicle dataset stored in a vehicle database (131, 33b) that has been collected from a vehicle (50) and associated with a scene; generating a scene context representing an interpretation of the scene based on the identified plurality of features; and registering the generated scene context in a RAG database (5, 33c) with an index corresponding to the plurality of features. [Item 22] A RAG database construction program that causes a processing unit (11, 31) to perform the following actions: identify a plurality of features of a vehicle dataset stored in a vehicle database (131, 33b) which is collected from a vehicle (50) and associated with a scene; generate a scene context representing the interpretation of the scene based on the identified plurality of features; and register the generated scene context in a RAG database (5, 33c) with an index corresponding to the plurality of features.

Claims

1. A RAG database construction system comprising: a vehicle database (131, 33b) storing vehicle datasets collected from vehicles (50) and associated with scenes; a text generation unit (211, 151) configured to identify multiple feature quantities of the vehicle dataset stored in the vehicle database and combine the identified multiple feature quantities to generate a scene context representing the interpretation of the scene; and a registration unit (212, 152) configured to assign an index corresponding to the multiple feature quantities to the scene context generated by the text generation unit and register it in a Retrieval-Augmented Generation (RAG) database (5, 33c).

2. The RAG database construction system according to claim 1, wherein the vehicle dataset includes multiple types of vehicle data that are different from each other, and the text generation unit (211, 151) is configured to identify the feature quantities of each of the multiple types of vehicle data.

3. The RAG database construction system according to claim 1 or 2, wherein the text generation unit (211, 151) is configured to generate the scene context based on the interpretation rules input from the vehicle knowledge database (132, 33d) which stores the interpretation rules for the plurality of features, and the vehicle dataset input from the vehicle database.

4. The RAG database construction system according to claim 3, wherein the text generation unit (211, 151) is configured to combine the plurality of feature quantities in accordance with the interpretation rules.

5. The RAG database construction system according to claim 4, wherein the text generation unit (211, 151) is configured to identify a feature corresponding to the cause of the scene and a feature corresponding to the result of the scene from among the plurality of feature quantities, and to generate the scene context in which the relationship between the cause of the scene and the result of the scene is expressed.

6. The RAG database construction system according to claim 1 or 2, wherein the registration unit (212, 152) is configured to cluster the scene contexts generated by the text generation unit (211, 151), assign an index to the scene contexts according to the clustering results, and register them in the RAG database (5, 33c).

7. The RAG database construction system according to claim 6, wherein the registration unit (212, 152) is configured to identify a feature corresponding to the result of the scene from the plurality of feature quantities of the scene context, and to cluster the scene context based on the feature corresponding to the result of the scene.

8. The RAG database construction system according to claim 6, wherein the registration unit (212, 152) is configured to identify a feature corresponding to the cause of the scene from the plurality of feature quantities of the scene context, and to determine the hierarchical structure of the cluster based on the feature corresponding to the cause of the scene.

9. The registration unit (212, 152) is configured to calculate the importance of feature quantities corresponding to the causes of the scenes, and to place feature quantities with higher importance higher in the hierarchical structure, as described in claim 8.

10. The RAG database construction system according to claim 1 or 2, wherein the vehicle database (33b) stores multiple vehicle datasets corresponding to multiple drivers in a manner that distinguishes each of the multiple drivers, the RAG database (33c) includes a generalized RAG database (331), the text generation unit (151) is configured to generate multiple scene contexts from the multiple vehicle datasets, and the registration unit (152) is configured to cluster the multiple scene contexts, exclude scene contexts belonging to clusters that satisfy exclusion conditions from the multiple scene contexts, and register them in the generalized RAG database (331).

11. The RAG database construction system according to claim 10, wherein the outlier condition is that the top layer of the cluster's hierarchical structure consists of the vehicle datasets of drivers fewer than a certain number of drivers among the plurality of drivers.

12. The RAG database construction system according to claim 10, wherein the outlier condition is that the lower layers of the cluster's hierarchical structure consist of the vehicle datasets of drivers fewer than a certain number of drivers among the plurality of drivers.

13. The RAG database construction system according to claim 10, wherein the RAG database (33c) further includes a specialized RAG database (332), the text generation unit (151) is configured to identify a plurality of features of a specialized dataset which is the vehicle dataset that corresponds to the specialization criteria, and to generate a specialized scene context by combining the plurality of features of the specialized dataset, and the registration unit (152) is configured to register the specialized scene context in the specialized RAG database with an index corresponding to the plurality of features of the specialized dataset when the similarity between the plurality of features of the specialized dataset and the plurality of features registered in the generalized RAG database (331) is less than a criterion value.

14. The RAG database construction system according to claim 13, further comprising a user unit (153) configured to receive queries from users and to search the generalized RAG database or the specialized RAG database based on the received queries.

15. The RAG database construction system according to claim 13, wherein the specialized dataset corresponds to one driver.

16. The RAG database construction system according to claim 13, wherein the specialized dataset corresponds to drivers with specific attributes.

17. The RAG database construction system according to claim 13, wherein the text generation unit (151) is configured to repeatedly acquire the specialized dataset and generate the specialized scene context from the specialized dataset each time the specialized dataset is acquired, and the registration unit (152) is configured to (i) update the specialized RAG database (332) each time the specialized scene context is generated by the text generation unit, and (ii) assign a first index to the feature quantities among the plurality of feature quantities registered in the specialized RAG database whose degree of change is greater than a reference degree, and / or assign a second index to the feature quantities whose degree of change is less than the reference degree.

18. The RAG database construction system according to claim 17, wherein the registration unit (152) is configured to delete from the specialized RAG database (332) any feature from the specialized RAG database whose similarity to any of the feature from the generalized RAG database (331) exceeds a threshold value, as the specialized RAG database (332) is repeatedly updated.

19. The vehicle database (131, 33b) includes a first vehicle database (131) and a second vehicle database (33b), the text generation unit (211, 151) includes a first text generation unit (211) and a second text generation unit (151), the registration unit (212, 152) includes a first registration unit (212) and a second registration unit (152), the vehicle (50) is equipped with an on-board device (2) having the first text generation unit, the first registration unit and the specialized RAG database (5), and an external device (3) having the second text generation unit, the second registration unit and the generalized RAG database (331), the first vehicle database stores the vehicle dataset corresponding to the driver of the vehicle, The RAG database construction system according to claim 13, wherein the first text generation unit is configured to generate the scene context from the vehicle dataset stored in the first vehicle database, the first registration unit is configured to register the scene context generated by the first text generation unit in the specialized RAG database, the second vehicle database stores a plurality of vehicle datasets corresponding to a plurality of drivers in a manner that distinguishes each of the plurality of drivers, the second text generation unit is configured to generate a plurality of scene contexts from the plurality of vehicle datasets stored in the second vehicle database, and the second registration unit is configured to register the plurality of scene contexts generated by the second text generation unit in the generalized RAG database.

20. The RAG database construction system according to claim 19, wherein the external device (3) is configured to repeatedly execute a cycle comprising: (i) machine learning a predetermined application using the plurality of vehicle datasets stored in the second vehicle database (33b); (ii) evaluating the learning results; (iii) collecting the plurality of vehicle datasets based on the evaluation of the learning results; and (iv) storing the collected plurality of vehicle datasets in the second vehicle database.

21. A method for constructing a RAG database, comprising: identifying multiple features of a vehicle dataset stored in a vehicle database (131, 33b) that is collected from a vehicle (50) and associated with a scene; generating a scene context representing the interpretation of the scene based on the identified multiple features; and registering the generated scene context in a RAG database (5, 33c) with an index corresponding to the multiple features.

22. A RAG database construction program that causes a processing unit (11, 31) to perform the following actions: identify a plurality of features of a vehicle dataset stored in a vehicle database (131, 33b) which is collected from a vehicle (50) and associated with a scene; generate a scene context representing the interpretation of the scene based on the identified plurality of features; and register the generated scene context in a RAG database (5, 33c) with an index corresponding to the plurality of features.