Large-model-based multi-mode intelligent question and answer robot in driving test field and development method of large-model-based multi-mode intelligent question and answer robot

Through a multimodal intelligent question-and-answer robot based on large pre-training models and driving test knowledge base, the interactive limitations of traditional robots are solved, and the intelligence and personalization of driving test assistance systems are realized, and learning efficiency and test pass rate are improved.

CN120336459APending Publication Date: 2025-07-18GUIYANG SHIJIHENGTONG TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510339337.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional multimodal intelligent question and answer robots have interactive limitations in the driving test field, which is difficult to provide intelligent and personalized assistance, and cannot effectively improve candidates' learning efficiency and test pass rate.

Method used

Large pre-trained models such as GPT-3 are adopted, and field customized training is carried out in combination with driving test knowledge base to realize multimodal input and output technology, have adaptive learning ability, and optimize the model through multimodal data fusion and user feedback.

Benefits of technology

It has achieved the improvement of the intelligent level of the driving test assistance system, provided a richer and more natural interactive experience, improved the accuracy of question and answer and user satisfaction, and is suitable for a variety of driving test scenarios.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a large-model-based multi-modal intelligent question and answer robot in the field of driving tests and a development method thereof. The method comprises the following steps: selecting and customizing a large pre-training model, constructing a driving test knowledge base, realizing a multi-modal input and output technology, performing question and answer by using multi-modal data, and performing self-adaptive learning ability. Through multi-modal interaction, examinees can obtain information and review knowledge more visually and more conveniently, by means of a large pre-training model and field customization training, the robot can accurately answer questions concerned by the examinees, and valuable guidance is provided. Self performance can be continuously optimized according to user feedback; the robot can be continuously improved, and better services are provided for users. The method is suitable for various driving test scenes, such as theoretical tests and simulated driving. The multifunctionality and flexibility of the robot enable the robot to be widely applied to various aspects of the field of driving tests.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a multi-modal intelligent question-answering robot in the field of driving tests based on a large model and its development method. Background Art

[0002] In the development process of artificial intelligence technology, multi-modal intelligent question-answering robots have emerged as the times require. It can communicate with users through various interaction methods (such as voice, text, images, etc.), providing a richer and more natural interaction experience. However, traditional multi-modal intelligent question-answering robots have certain limitations when dealing with complex problems and specific requirements in the field of driving tests. In recent years, multi-modal intelligent question-answering technology based on large pre-trained models has gradually attracted attention. Large pre-trained models have powerful language understanding and generation capabilities, can handle various complex problems, and can combine multi-modal data for more accurate question answering.

[0003] The present invention relates to such a multi-modal intelligent question-answering robot, which focuses on the field of driving tests and aims to provide a more intelligent and personalized auxiliary tool for candidates. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-modal intelligent question-answering robot in the field of driving tests based on a large model and its development method to achieve a richer and more natural interaction experience and improve the intelligence level of the driving test assistance system. We hope that such a robot can help candidates prepare for the driving test more effectively and improve their learning efficiency and passing rate of the exam.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] A development method for a multi-modal intelligent question-answering robot in the field of driving tests based on a large model, comprising the following steps:

[0007] (1) Select a mature large pre-trained model as the base model, and apply the large pre-trained model to the field of driving tests through transfer learning technology;

[0008] (2) Collect text materials, laws and regulations, and exam outlines related to driving tests, organize and structure the knowledge base to ensure the accuracy of information for easy query and answer by the model;

[0009] (3) Combine the driving test knowledge base to perform domain customization training on the base model so that it can handle complex problems in the field of driving tests;

[0010] (4) Implement multi-modal input and output technology, specifically including:

[0011] Develop a voice recognition and synthesis module to support voice interaction;

[0012] Develop image recognition and processing functions so that the robot can understand and generate image information;

[0013] Integrate text input / output to ensure that the robot can handle text queries and provide text answers;

[0014] (5) Utilize multimodal data for more accurate question answering, specifically including:

[0015] Design a multimodal data processing flow to ensure that text, voice, and image information can be effectively combined during the question answering process;

[0016] Implement a multimodal data fusion algorithm so that the robot can provide the best answer by integrating multimodal information when answering questions;

[0017] (6) Adaptive learning ability, specifically including:

[0018] Collect user interaction data, including user questions, robot answers, and user feedback; Use machine learning algorithms to continuously optimize the robot model to improve the accuracy of answers and user satisfaction;

[0019] Furthermore, the large pre-trained model selects GPT-3.

[0020] Furthermore, applying the large pre-trained model to the driving test field is achieved by fine-tuning on the pre-trained model and retraining the model using the driving test field dataset.

[0021] Furthermore, step 3 is implemented by machine learning methods including supervised learning or reinforcement learning.

[0022] The present invention also provides a multimodal intelligent question answering robot constructed by using the above multimodal intelligent question answering robot development method.

[0023] The present invention has the following beneficial effects:

[0024] 1. Achieve a richer and more natural interaction experience and improve the intelligence level of the driving test assistance system; Through multimodal interaction, candidates can obtain information and review knowledge more intuitively and conveniently.

[0025] 2. Handle complex problems in the driving test field and improve the accuracy of question answering and user experience; With the help of large pre-trained models and domain-specific customized training, the robot can accurately answer questions that candidates care about and provide valuable guidance.

[0026] 3. Have the ability of adaptive learning and can continuously optimize its own performance according to user feedback; This enables the robot to continuously improve and provide better services for users.

[0027] 4. Applicable to various driving test scenarios, such as theory tests, simulated driving, etc.; the versatility and flexibility of the robot enable it to be widely used in all aspects of the driving test field. Detailed implementation manners

[0028] The following further describes the detailed implementation manners of the present invention. It should be noted here that the description of these implementation manners is used to help understand the present invention, but does not constitute a limitation to the present invention. In addition, the technical features involved in the various implementation manners of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0029] A method for developing a multi-modal intelligent Q&A robot in the driving test field based on a large model includes the following steps:

[0030] 1. Selection of a large pre-trained model and customization of the model

[0031] Select a large pre-trained model as the base model, such as GPT-3. GPT-3 is a large language model developed by OpenAI and has the ability to generate coherent and in-depth text.

[0032] Customization of the model: Apply the large pre-trained model to the driving test field through transfer learning techniques. This can be achieved by fine-tuning on the pre-trained model and retraining the model using the driving test field dataset.

[0033] 2. Construction of a knowledge base in the driving test field

[0034] Collect text materials, laws and regulations, and examination outlines related to driving tests, organize and structure the knowledge base to ensure the accuracy of information for easy query and answer by the model; this knowledge base should contain all knowledge points related to driving tests, such as traffic rules, driving skills, vehicle structure, etc. These knowledges can be obtained by collecting driving test textbooks, past years' real exam questions, expert interviews, etc. The construction of the knowledge base is the basis for domain-specific customization training.

[0035] 3. Perform domain-specific customization training on the base model.

[0036] We combine the knowledge base in the driving test field with the large pre-trained model, and through specific training data and strategies, enable the model to better understand and answer driving test-related questions. This step can be achieved through machine learning methods such as supervised learning and reinforcement learning. Combine the knowledge base in the driving test field to perform domain-specific customization training on the base model so that it can handle complex problems in the driving test field. This means that we need to create a dataset containing driving test-related questions and then use this dataset to train the base model so that it can understand and answer specific questions in the driving test field.

[0037] 4. Implement multi-modal input and output technologies.

[0038] To enable the robot to interact with users more naturally and richly, we need to support multiple interaction methods. Therefore, we adopt multi-modal input and output technologies to achieve various interaction methods with users, such as voice, text, and image.

[0039] For example, users can ask questions by voice, and the robot can answer by text or voice; users can also input images, such as uploading a picture of a traffic sign, and the robot can recognize the picture and give relevant explanations and tips.

[0040] Specifically, the following work or steps need to be carried out:

[0041] a. Develop speech recognition and synthesis modules to support voice interaction;

[0042] b. Develop image recognition and processing functions so that the robot can understand and generate image information;

[0043] c. Integrate text input / output to ensure that the robot can handle text queries and provide text answers;

[0044] 5. Use multi-modal data for more accurate question answering.

[0045] When processing users' questions, the robot not only relies on text information but also can utilize information in other modalities such as images and voices to obtain more accurate answers. For example, when a user uploads a picture of a driving scene, the robot can not only recognize the elements in the picture but also generate an accurate answer by combining the picture and the user's question. Specifically, the following work or steps need to be carried out:

[0046] a. Design a multi-modal data processing flow to ensure that text, voice, and image information can be effectively combined during the question answering process.

[0047] b. Implement a multi-modal data fusion algorithm so that when answering questions, the robot can comprehensively use multi-modal information to provide the best answer.

[0048] 6. Have the ability of adaptive learning.

[0049] It can continuously optimize its own performance according to user feedback, which means that the robot can learn from interactions with users and continuously improve its question answering strategies and answer quality. This can be achieved by collecting user feedback, analyzing interaction data, etc.

[0050] Collect user interaction data, including users' questions, the robot's answers, and users' feedback, and use machine learning algorithms, such as reinforcement learning, to continuously optimize the robot model to improve the accuracy of answers and user satisfaction.

[0051] Through the above technical solution, we have realized a multi-modal intelligent Q&A robot in the field of driving tests based on large models. Such a robot can provide a rich and natural interaction experience, accurately answer questions related to driving tests, and help candidates prepare for driving tests more effectively.

[0052] For example, by collecting user feedback, analyzing interaction data, etc., the robot can continuously adjust its answer style and answer accuracy to better meet user needs.

[0053] The present invention also provides a multi-modal intelligent Q&A robot constructed by using the above multi-modal intelligent Q&A robot development method.

[0054] The above has described the embodiments of the present invention in detail, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principles and spirit of the present invention, various changes, modifications, substitutions, and variations made to these embodiments still fall within the protection scope of the present invention.

Claims

1. A development method for a multi-modal intelligent Q&A robot in the driving test field based on a large model, characterized in that, It includes the following steps: (1) Select a mature large pre-trained model as the base model, and apply the large pre-trained model to the driving test field through transfer learning technology; (2) Collect text materials, laws and regulations, and examination outlines related to driving tests, organize and structure the knowledge base to ensure the accuracy of information for easy model query and answer; (3) Combine the driving test knowledge base to perform domain customization training on the base model so that it can handle complex problems in the driving test field; (4) Implement multi-modal input and output technologies, specifically including: Develop speech recognition and synthesis modules to support voice interaction; Develop image recognition and processing functions so that the robot can understand and generate image information; Integrate text input / output to ensure that the robot can handle text queries and provide text answers; (5) Use multi-modal data for more accurate question answering, specifically including: Design a multi-modal data processing process to ensure that text, voice, and image information can be effectively combined during the question answering process; Implement a multi-modal data fusion algorithm so that the robot can provide the best answer by integrating multi-modal information when answering questions; (6) Adaptive learning ability, specifically including; Collect user interaction data, including user questions, robot answers, and user feedback; use machine learning algorithms to continuously optimize the robot model to improve the accuracy of answers and user satisfaction.

2. The method for developing a multi-modal intelligent question-answering robot in the driving test field based on a large model according to claim 1, characterized in that: The large pre-trained model selected is GPT-3.

3. The method for developing a multi-modal intelligent Q&A robot according to claim 1, wherein: Applying the large pre-trained model to the driving test field is achieved by fine-tuning on the pre-trained model and retraining the model using the driving test field dataset.

4. The multimodal intelligent Q&A robot development method according to claim 1, wherein: Step 3 is implemented by machine learning methods including supervised learning or reinforcement learning.

5. A multimodal intelligent question-answering robot, characterized in that: It is constructed by adopting the multi-modal intelligent question answering robot development method described in any one of claims 1-4.