A method for generating training samples for a teaching language model for Subject 3
By using multi-source data fusion and consistency verification methods, training samples with high semantics and high consistency are generated, which solves the problems of insufficient semantic expression and lack of temporal context in the teaching language model of Subject 3, and achieves personalized and standardized teaching effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YIXIAN INTELLIGENCE
- Filing Date
- 2026-01-21
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies in the teaching language model for Subject 3 suffer from insufficient semantic expression of multi-source data, lack of temporal context, poor sample consistency, and limited scenario adaptability. This results in inaccurate and inconsistent generation of teaching instructions, making it difficult to meet the standardization, scalability, and personalization needs of the driver training industry.
By constructing a unified semantic expression mechanism for multi-source heterogeneous data, integrating vehicle status, road environment, and trainee operation timing features, high semantic and high consistency training samples are generated. Key operation history construction and timing enhancement strategies are introduced, multi-source teaching annotations are integrated, and consistency verification and conflict resolution are performed to achieve the engineering and configurability of the sample generation process.
It improves the accuracy of the language model's understanding of driving behavior and the consistency of teaching, adapts to the personalized needs of different driving test subjects and student characteristics, reduces the reliance on manual annotation and experience-based rule design, and improves the flexibility and practicality of the teaching system.
Smart Images

Figure CN122132831A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of natural language processing and driver training, specifically to a method for generating training samples for a driving test language model. Background Technology
[0002] In traditional driving training for the third stage of the driving test, instruction typically relies on a human instructor providing real-time guidance from the passenger seat. The instructor provides verbal instruction based on the vehicle's driving status, the student's actions, and the road or practice area environment, including corrective actions, key test points, and safety reminders. While this method can help students master basic driving skills to some extent, it is highly dependent on the instructor's individual experience, limiting the consistency and scalability of instruction.
[0003] With the development of IoT and vehicle sensing technologies, modern driver training vehicles are typically equipped with various sensors, such as GPS, IMU, and onboard control units, capable of collecting structured numerical information in real time, including vehicle position, attitude, speed, gear position, and braking status. Based on this, some driving instruction systems have begun to introduce automated teaching algorithms, which use pre-configured rules to assess vehicle status and trigger corresponding teaching prompts.
[0004] In recent years, with the development of pre-trained language models, the industry has begun to explore the application of large language models in education and teaching scenarios. In the field of driving instruction, there have been attempts to use structured vehicle state data as model input, hoping to directly generate teaching instructions or explanatory text from the language model. However, language models naturally use natural language as their input and output form, and their ability to understand contextual relationships, behavioral causality, and scene semantics depends on whether the input data has sufficient semantic information. Simply inputting multi-dimensional numerical sensor data directly into the model is often difficult for the model to understand effectively, easily leading to inaccurate or untargeted teaching instructions. Therefore, in existing technologies, how to transform multi-source heterogeneous data such as vehicle state, road information, and student operating behavior into semantic representations that language models can understand has become a key issue restricting the practical application of language models in driving test instruction scenarios.
[0005] The closest existing technical solutions to this invention mainly fall into two categories. The first category consists of rule-based or state machine-based automatic teaching algorithms. These algorithms typically use vehicle sensor data and pre-configured road or site location information as input, directly outputting teaching prompts or evaluation results through manually set thresholds, rules, or state transition logic. Their advantages include simplicity and strong interpretability, but the teaching content is highly dependent on rule design, making it difficult to cover complex or continuously changing driving scenarios, and they lack the contextual semantic expression capabilities required by language models. The second category directly uses structured vehicle state data as input to train or drive a language model to generate teaching instructions. In this approach, data such as vehicle speed, gear, and steering angle are typically input into the model in numerical or simple text fields, and the model generates corresponding teaching outputs based on the current input. However, due to the lack of a unified description of historical operational behavior, road and site semantics, and driving context, this type of solution struggles to establish stable contextual relationships, resulting in low consistency and accuracy in model output results.
[0006] While the aforementioned existing technical solutions introduce automation or intelligent methods to some extent, they still fail to address issues such as insufficient semantic understanding of multi-source states, lack of temporal context, and difficulty in ensuring sample consistency in the context of driving test preparation. Consequently, they struggle to support high-quality, scalable training of teaching models. Specifically, existing technologies exhibit the following main technical shortcomings when applying language models to driver training: 1. The direct input of multi-source vehicle status data into the language model in structured numerical form results in insufficient semantic expressive power. Existing solutions typically input multi-source status data such as vehicle speed, gear position, steering angle, and pose into the language model in numerical or simple field form. Since the language model primarily uses natural language for understanding, purely numerical or weakly semantically structured data struggles to express the contextual meaning and causal relationships behind driving behavior. This leads to the model's inability to accurately understand the current driving state, consequently affecting the accuracy and rationality of generated instruction commands.
[0007] 2. Lack of a unified description mechanism for road environment, enclosed venues, and subject context. In existing technologies, road information, practice venue information, and subject progress often exist in a scattered or implicit manner, failing to form a unified semantic expression. Without a clear scene context, language models struggle to distinguish between different practice stages of Subject 3, different road types, or different venue constraints, easily generating teaching content that does not match the current teaching scenario.
[0008] 3. Ineffective modeling of the temporal context of learner operations. Driving instruction has a distinct temporal characteristic; the vehicle's state at a single moment is insufficient to reflect the learner's true operational intentions. Existing solutions typically focus only on the current state input, lacking a systematic model of the learner's key operational history. This results in the language model's inability to understand the causal relationships between consecutive operations, thus affecting the temporal coherence and relevance of instructional commands.
[0009] 4. The single or inconsistent source of training sample annotations affects the quality of model training. Some existing solutions rely on a single source to generate teaching annotations, such as relying solely on rule-based algorithms or manual annotation, lacking multi-source annotation fusion and consistency verification mechanisms. As the sample size increases, annotation noise or semantic inconsistencies are easily introduced, reducing the stability and generalization ability of language model training.
[0010] 5. The sample generation process lacks engineering constraints, making it difficult to adapt to the needs of different learners and subjects. In existing technologies, sample generation often adopts fixed rules or uniform formats, making it difficult to flexibly adjust the key operation selection strategy and contextual information range according to different learner characteristics or different subject requirements, thus limiting the application effect of language models in personalized teaching scenarios.
[0011] In summary, existing technologies for generating training samples for driving test language models suffer from multiple technical bottlenecks, including insufficient semantic expression, lack of temporal context, poor sample consistency, and limited scenario adaptability. These bottlenecks make it difficult for language models to generate accurate and stable teaching instructions, failing to meet the driving training industry's demands for standardized, large-scale, and personalized teaching. There is an urgent need for a training sample generation solution that can integrate multi-source data, enhance semantic and temporal features, and ensure sample quality. Summary of the Invention
[0012] This invention aims to overcome the shortcomings of existing technologies in generating training samples for driving test language models, such as insufficient semantic representation of multi-source data, lack of unified description of road and subject contexts, missing temporal context, poor sample consistency, and limited scenario adaptability. It provides a method for generating training samples for driving test language models, integrating multi-source vehicle states, road and site information, and student operation temporal characteristics to generate highly semantic and consistent training samples. This method is particularly suitable for language model training scenarios in intelligent driving instruction systems. By constructing a unified semantic representation mechanism for multi-source heterogeneous data, vehicle states, road environments, closed site information, and subject progress are transformed into unified language state descriptions understandable by the language model. Key operation history construction and temporal enhancement strategies are introduced to capture the causal relationships between consecutive student operations. Multi-source instructional annotations are integrated, and consistency verification and conflict resolution mechanisms are established, achieving engineering and configurability of the sample generation process. This results in highly semantic, highly consistent training samples with temporal context, improving the accuracy of the language model's understanding of driving behavior, instructional consistency, and scenario generalization ability, adapting to the personalized needs of different driving test instruction items and student characteristics.
[0013] To achieve the above objectives, the technical solution of this invention is: a method for generating training samples for a teaching language model for Subject 3, comprising the following steps: Multi-source data acquisition and state fusion are performed to acquire static input data and dynamic input data. The static input data includes data describing the driving teaching environment, and the dynamic input data includes data describing the vehicle's operating status and the student's operating behavior. The two data are then fused together, and the current driving teaching scenario type is determined by combining the vehicle status, road environment, closed site information, and subject description. Semantic abstraction and language state description generation: Based on preset semantic abstraction rules, feature extraction and rule judgment are performed on the fused state data to generate a unified language state description with clear semantics; The key operation history is constructed and time-series enhanced. It filters key student operations related to the current subject and teaching objectives, records them in chronological order and associates them with corresponding semantic states to form a key operation history sequence. Post-processing strategies can be configured according to needs. Sample input construction and multi-source annotation fusion: The language state description, key operation history sequence and subject description information are combined into a unified format sample input, multi-source teaching annotations are fused and consistency verification and conflict resolution are performed to generate the final teaching annotations; Complete sample output: Integrate sample input with final teaching annotations to output a complete sample for training the language model for Subject 3 teaching.
[0014] Furthermore, the static input data includes structured information of the driving school's closed course or road, subject description information, teaching rules, and key operation configuration parameters; the dynamic input data includes vehicle component status data, vehicle position information, and real-time student operation behavior data; the structured information of the driving school's closed course or road includes road type, road direction, intersection information, number of lanes, and site constraints related to the Subject 3 exam; the subject description information is used to indicate the current Subject 3 sub-item stage and practice objectives of the teaching or training; the teaching rules and key operation configuration parameters are used to indicate the key operation categories that need to be focused on under different subjects and different student conditions.
[0015] Furthermore, the vehicle component status data includes vehicle speed, gear position, braking status, and steering status; the vehicle position and posture information includes position, heading angle, and motion trend; and the student's real-time operation behavior data is used to determine whether key operations such as gear shifting, braking, and turn signal operation have occurred.
[0016] Furthermore, the scenario types include scenarios related to driving in a straight line, turning, passing through intersections, deceleration control, and gear shifting, which are all related to the teaching of driving test subject three.
[0017] Furthermore, the semantic abstraction and rule judgment specifically include extracting the following state features with pedagogical significance: Current driving behavior trends, including acceleration, deceleration, and stable driving; Does the current operation meet the teaching requirements of Subject 3? Teaching focus points that may exist in the current scenario include maintaining stable vehicle speed while driving in a straight line and the timing of turn signal operation when turning.
[0018] Furthermore, the unified language state description can explicitly express the following information: The current driving scenario of the vehicle; The student's current operational status and its rationality; Contextual information relevant to the current teaching objectives of Subject 3.
[0019] Furthermore, the key operations include gear shifting, braking, turn signal operation, steering wheel steering, throttle control, and other operations related to the teaching objectives of Subject 3.
[0020] Furthermore, the key operation history record is a sequence of key operations that are filtered and recorded in chronological order and associated with corresponding semantic states.
[0021] Furthermore, the configurable post-processing strategy includes adjusting key operation types, historical record lengths, or recording rules based on student characteristics and the needs of the Subject 3 teaching project to achieve personalized sample generation. Student characteristics include student learning progress, operational proficiency, and common error-prone features.
[0022] Furthermore, the sample input construction involves combining the language state description generated at the current moment, the key operation history sequence, and the subject description information to form a sample input in a unified format; the multi-source teaching annotation includes: Teaching annotations generated based on the driving test (subject 3) teaching algorithm deployed on real vehicles; Simulation teaching annotations generated based on the distribution of real test data; Manually annotated teaching instructions or correction results; The consistency verification and conflict resolution process involves analyzing the consistency of instructional annotations from different sources. When conflicts exist, they are resolved using preset rules or priority strategies to generate the final instructional annotation results for training.
[0023] Compared with the prior art, the present invention has the following significant advantages: By using a unified semantic abstraction and linguistic expression mechanism based on multi-source data, the purely numerical vehicle status, road environment, and subject information are transformed into natural language descriptions with clear scene meanings and causal relationships. This solves the problem of insufficient semantic expression in existing technologies, enabling the language model to accurately understand the core elements of driving teaching scenarios and significantly improving the pertinence and rationality of teaching instruction generation.
[0024] By constructing key operation histories and employing temporal enhancement strategies, the system records the relationships and semantic states of trainees' continuous operations, overcoming the shortcomings of existing technologies that lack temporal modeling. Based on complete operational temporal logic, the language model can understand the cause and effect of driving behavior, generating coherent and realistic teaching content that helps trainees establish a clear logical chain of driving behavior.
[0025] By integrating real-vehicle algorithms, simulation data, and manual annotations into a multi-source annotation mechanism, coupled with consistency verification and conflict resolution strategies, the noise and semantic inconsistency issues inherent in single-source annotations are effectively avoided, generating highly consistent and reliable training samples. This design significantly improves the stability and generalization ability of language model training, providing high-quality data support for large-scale, standardized teaching.
[0026] The sample generation process supports engineering configuration, allowing adjustments to parameters such as key operation screening rules and historical record length based on different driving test (Part 3) teaching items and student learning characteristics, overcoming the limitations of fixed formats in existing technologies. This feature enables training samples to adapt to the learning pace and ability levels of different students, while also being compatible with diverse training venues and road scenarios, significantly improving the flexibility and practicality of the teaching system.
[0027] This invention achieves automated and standardized generation of driving instruction samples through technological means, reducing reliance on manual annotation and experience-based rule design, and lowering labor costs and teaching barriers in the driver training industry. Simultaneously, its supporting language model can output standardized and unified teaching content, promoting the transformation of the driver training industry from experience-based teaching to standardized and competency-based teaching, significantly improving overall training quality and efficiency. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings used in the description of the embodiment or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this embodiment. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a flowchart of the training sample generation method of the present invention. Detailed Implementation
[0030] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this embodiment. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this embodiment as detailed in the appended claims.
[0031] like Figure 1 As shown, this embodiment provides a method for generating training samples for a teaching language model for Subject 3, including the following steps: Multi-source data acquisition and state fusion are performed to acquire static and dynamic input data. The static input data includes data describing the driving teaching environment, while the dynamic input data includes data describing the vehicle's operating status and the student's operational behavior. These two data sources are then fused together, combining vehicle status, road environment, closed-course information, and subject descriptions to determine the current driving teaching scenario type. The static input data provides basic structural information about the driving environment, including road type, intersections, and the number of lanes, helping the language model make accurate teaching judgments in different scenarios. The dynamic input data provides necessary context for the language model by providing real-time feedback on vehicle status and student operational behavior, ensuring that teaching instructions are consistent with the current actual driving situation.
[0032] By integrating static and dynamic data, language models can be provided with more comprehensive and real-time information, enabling the generation of teaching content to better reflect changes in the actual driving environment, thereby improving the personalization and accuracy of teaching.
[0033] Semantic abstraction and language state description generation: Based on preset semantic abstraction rules, feature extraction and rule judgment are performed on the fused state data to generate a unified language state description with clear semantics. This step transforms the original numerical data into language expressions with practical teaching significance by defining a set of preset semantic abstraction rules, ensuring that the language model can extract behavioral patterns, driving trends and other meaningful features from driving data, enhancing the language model's comprehension ability, and making the teaching content output by the model more targeted and reasonable.
[0034] The system constructs and enhances the history of key operations, filtering key student operations relevant to the current subject and teaching objectives. These operations are recorded chronologically and associated with corresponding semantic states to form a key operation history sequence. Post-processing strategies can be configured as needed. By modeling the temporal sequence of student operations, the system ensures that it understands the continuous actions and causal relationships of students, thereby improving the coherence and accuracy of generated teaching content. Configurable post-processing strategies allow the sample generation process to be flexibly adjusted according to the learning needs of different students, enhancing the effectiveness of personalized teaching.
[0035] The sample input construction and multi-source annotation fusion process combines the language state description, key operation history sequence, and subject description information into a unified format sample input. It then integrates multi-source teaching annotations, performs consistency checks and conflict resolution, and generates the final teaching annotations. This step merges multiple annotation sources from real-vehicle teaching algorithms, simulation data, and manual annotations, and ensures data consistency and reliability through consistency analysis. This not only enhances sample diversity but also avoids annotation conflicts or data noise, ensuring high-quality training data.
[0036] The complete sample output integrates the sample input and final teaching annotations to produce a complete sample for training the language model for Subject 3 teaching. By integrating language state descriptions, operation history, and teaching annotation information, this step ensures that the generated training samples meet the needs of the language model in practical applications and support efficient model training.
[0037] As one implementation method, the static input data in this embodiment includes structured information of the driving school's closed course or road, subject description information, teaching rules, and key operation configuration parameters; the dynamic input data includes vehicle component status data, vehicle position information, and real-time student operation behavior data; the structured information of the driving school's closed course or road includes road type, road direction, intersection information, number of lanes, and site constraints related to the Subject 3 exam; the subject description information is used to indicate the current Subject 3 sub-item stage and practice objectives of the teaching or training; the teaching rules and key operation configuration parameters are used to indicate the key operation categories that need to be focused on under different subjects and different student conditions.
[0038] The "structured information of driving school closed sites or roads" in the static input data provides clear road environment features, helping the language model understand the specific requirements and constraints of the driving site, such as intersections and the number of lanes.
[0039] The "subject description information" precisely indicates the current teaching objectives and training stage, ensuring that the language model can generate guidance content that aligns with the specific stage. This data fusion ensures that the language model can fully understand the current driving scenario, including the geographical environment, teaching objectives, and site limitations, thereby making the generated teaching content more accurate and adaptable.
[0040] As one implementation method, the vehicle component status data in this embodiment includes vehicle speed, gear position, braking status, and steering status; the vehicle pose information includes position, heading angle, and motion trend; the student's real-time operation behavior data is used to determine whether key operations such as gear shifting, braking, and turn signal operation have occurred. Vehicle component status data, such as vehicle speed, gear position, and braking status, provides real-time feedback on vehicle operation, helping the language model determine whether the current driving behavior conforms to teaching standards. Vehicle pose information (including position and heading angle) helps the model determine the vehicle's precise position and direction of travel on the road, thereby providing better and more appropriate teaching feedback. The student's real-time operation behavior data enables the model to identify the student's operations at each stage, thus providing targeted teaching guidance.
[0041] The introduction of real-time dynamic data makes feedback in the teaching process more immediate and personalized, providing accurate corrective suggestions and guidance based on the student's driving behavior, avoiding generic prompts, and improving the pertinence and accuracy of teaching effectiveness.
[0042] As one implementation method, the scenario types described in this embodiment include scenarios related to driving test part three, such as straight driving, turning, intersection passage, deceleration control, and gear shifting. Defining scenario types helps the language model accurately identify different driving situations, such as straight driving and turning. Each scenario involves different teaching objectives and operational requirements.
[0043] This scenario classification ensures that the language model provides more targeted instructional feedback based on the current scenario, helping learners to operate accurately in specific situations.
[0044] As one implementation method, the semantic abstraction and rule judgment described in this embodiment specifically includes extracting the following state features with pedagogical significance: Current driving behavior trends, including acceleration, deceleration, and stable driving; Does the current operation meet the teaching requirements of Subject 3? Potential teaching focuses in the current context.
[0045] The current driving behavior trend (such as acceleration, deceleration, and stable driving) helps the language model determine whether the learner should perform certain specific actions (such as acceleration or deceleration) in the current driving state. Determining whether the current action meets the requirements of the driving test (Part 3) helps the model understand whether the learner's actions conform to driving standards, thus providing improvement suggestions. Potential teaching focuses within the current scenario further ensure that the language model responds to specific challenges in the scenario.
[0046] By extracting and analyzing these key features, language models can provide accurate evaluations of learners' actions and offer timely guidance for improvement, thereby enhancing teaching effectiveness.
[0047] As one implementation method, the unified language state description described in this embodiment can clearly express the following information: The current driving scenario of the vehicle; The student's current operational status and its rationality; Contextual information relevant to the current teaching objectives of Subject 3.
[0048] The Unified Language State Description is a comprehensive abstraction of the current driving scenario, ensuring that the language model understands the causal relationships between various driving behaviors, thereby generating more context-appropriate teaching feedback.
[0049] A consistent language helps the system clearly convey the full picture of driving behavior, ensuring that the generated instructional feedback not only aligns with the current operation but also takes into account other important factors in the context.
[0050] As one implementation method, the key operations described in this embodiment include gear shifting, braking, turn signal operation, steering wheel steering, throttle control, and other operations related to the teaching objectives of Subject 3. By screening key operations, the system can ensure that every piece of teaching feedback is related to the examination requirements of Subject 3 and can provide feedback for each specific operation.
[0051] Precise operational control ensures that instructional feedback is directly related to the student's driving behavior, helping the student focus on improving the most critical operational steps.
[0052] As one implementation method, the key operation history record in this embodiment is to record the filtered key operations in chronological order and associate them with the corresponding semantic states to form a key operation history sequence.
[0053] The key operation history record enables the system to track the student's continuous operations, ensuring causal reasoning of driving behavior and enhancing the temporality and coherence of the teaching content.
[0054] By establishing a connection between operational history and semantic state, language models can understand the relationships between learners' actions, thereby generating more coherent and logically clear instructional feedback.
[0055] As one implementation method, the configurable post-processing strategy described in this embodiment includes adjusting the key operation type, historical record length, or recording rules according to student characteristics and the needs of the Subject 3 teaching project, so as to achieve personalized sample generation.
[0056] Configurable post-processing strategies enable the teaching process to be customized according to learners' characteristics (such as learning speed, responsiveness, etc.), thereby improving the effectiveness of personalized teaching.
[0057] This flexible configuration allows the teaching system to be optimized based on different learners, further enhancing the system's adaptability and flexibility.
[0058] As one implementation method, the sample input construction in this embodiment combines the language state description generated at the current moment, the key operation history sequence, and the subject description information to form a sample input in a unified format; the multi-source teaching annotation includes: Teaching annotations generated based on the driving test (subject 3) teaching algorithm deployed on real vehicles; Simulation teaching annotations generated based on the distribution of real test data; Manually annotated teaching instructions or correction results; The integration of multi-source teaching annotations ensures that the annotations obtained from different data sources (real vehicles, simulations, and manual annotations) can provide comprehensive and accurate teaching feedback.
[0059] Multi-source annotation fusion improves the diversity and reliability of training data, thereby enhancing the generalization ability and accuracy of language models.
[0060] The consistency verification and conflict resolution process involves analyzing the consistency of instructional annotations from different sources. When conflicts exist, they are resolved using preset rules or priority strategies to generate the final instructional annotation results for training.
[0061] Consistency checks and conflict resolution ensure that labeled data from different sources are free of conflicts, and that conflicting data is properly handled through rules or priority strategies to generate consistent and accurate labeling results. This mechanism guarantees the consistency and high quality of labeled data, thereby improving the stability and reliability of model training.
[0062] Example 2: Method for Generating Training Samples for Subject 3 Teaching Language Model The training sample generation algorithm of the present invention accepts multi-source heterogeneous input data, which includes two main categories: static data and dynamic data.
[0063] Static input data is used to describe information about a relatively stable driving training environment, including but not limited to: Structured information about the driving school's closed grounds or roads, such as road type, road direction, intersection information, number of lanes, and site constraints related to the driving test (subject 3); Subject description information indicates the current teaching or training stage of subject three sub-item and the training objectives; Teaching rules and key operation configuration parameters are used to indicate the key operation categories that need to be focused on under different subjects and different student conditions.
[0064] Dynamic input data is used to describe the vehicle's operating status and the trainee's operational behavior in real time, including but not limited to: Vehicle component status data, such as vehicle speed, gear, braking status, steering status, etc. Vehicle pose information, such as position, heading angle, and motion trend; Real-time operational behavior data of trainees is used to determine whether critical operations such as gear shifting, braking, and turn signal operation have occurred.
[0065] The input data mentioned above can come from vehicle sensor systems, vehicle control systems, or simulation systems, and is input into the sample generation algorithm of this invention in the form of structured data.
[0066] The core of this invention lies in a language state generation algorithm, which is used to convert multi-source heterogeneous structured vehicle states, road or site information and teaching context into a unified language state description that can be understood by the language model, thereby solving the problem that the language model has difficulty in directly understanding the contextual relationship of driving scene based on pure numerical states.
[0067] The process mainly includes the following steps: Multi-source state fusion and scenario determination: This involves fusing static and dynamic input data to determine the current driving instruction scenario type based on the vehicle's current state, road or site information, and the subject description. For example, it differentiates between various teaching scenarios such as straight driving, turning, intersection passage, and deceleration control.
[0068] Semantic state abstraction and rule judgment: Based on preset semantic abstraction rules, the fused vehicle state is analyzed to extract state features with pedagogical significance, including but not limited to: The current driving behavior trend (e.g., acceleration, deceleration, stable driving); whether the current operation meets the requirements of the subject; potential teaching focus points in the current scenario.
[0069] The above abstraction process is not a simple data splicing, but rather a process of transforming numerical states into intermediate semantic states with clear semantic meanings through rule judgment, state combination, and scene mapping.
[0070] Unified Language State Description Generation: After obtaining the intermediate semantic state, a language state generation algorithm is used to transform the intermediate semantic state into a state description in natural language form. This language state description can explicitly express: The current driving scenario of the vehicle; The student's current operational status and its rationality; Contextual information relevant to the current teaching objectives of Subject 3.
[0071] The generated language state descriptions serve as the core input to the language model training samples, enabling the language model to understand driving teaching scenarios and their causal relationships based on natural language.
[0072] Construction and temporal enhancement of critical operation history: In order to enhance the language model's ability to understand the temporal relationship of driving behavior, this invention further introduces a critical operation history mechanism.
[0073] Key operation screening: From the trainee's continuous operation behavior, only key operations that are related to the current subject and teaching objectives are screened, such as shifting gears, braking, and turn signal operation, to avoid irrelevant operations from interfering with the understanding of the model.
[0074] Key operation history record: The selected key operations are recorded in chronological order and associated with the corresponding semantic states to form a key operation history sequence.
[0075] Configurable post-processing algorithm: Configure the recording strategy for key operation history according to student characteristics and different subject requirements, such as adjusting the key operation type, history length or recording rules to achieve personalized sample generation.
[0076] The key operation history, as a contextual component of the language state description, together with the current language state description, constitutes a complete model input, thereby improving the language model's ability to understand continuous driving behavior.
[0077] Training sample generation and annotation fusion: After completing the language state description and key operation history construction, this invention further generates complete samples for language model training.
[0078] Sample input construction: Combine the language state description generated at the current moment, key operation history, and subject description information to form a sample input in a unified format.
[0079] Teaching annotation generation: Sample output annotation sources include: teaching annotations generated by the driving test 3 teaching algorithm based on real vehicle deployment; simulation teaching annotations generated based on the distribution of real test data; and teaching instructions or correction results annotated manually.
[0080] Multi-source annotation consistency verification and conflict resolution: Consistency analysis is performed on teaching annotations from different sources. When conflicts exist, they are resolved through preset rules or priority strategies to generate the final teaching annotation results for training.
[0081] Sample generation process description: Combining the above steps, the training sample generation process of this invention includes the following steps: Acquire multi-source static and dynamic data; fuse multi-source states and determine the current driving teaching scenario; perform semantic abstraction of vehicle states and generate language state descriptions; filter and construct key operation histories to form temporal contexts; generate sample inputs and fuse multi-source teaching annotations; output complete samples for language model training.
[0082] Through the above process, this invention achieves the automated generation of training samples that can be understood by the language model from structured vehicle states.
[0083] This invention abstracts roads into a continuous and high-density sequence of road points. By matching the vehicle's current position with these abstract road points, it determines the vehicle's road location, lane information, and road type in real time. Based on this, and combined with the vehicle's current mechanical state (including but not limited to speed, gear, steering, and braking), the system can continuously infer the rationality of the current driving behavior during vehicle operation and instantly generate teaching prompts that match the current driving state. This achieves continuous, dynamic, real-time teaching, avoiding the problems of delayed teaching triggering or disconnection from the actual vehicle state in traditional solutions.
[0084] Because this invention performs interpolation and structured abstraction on the original road points, the road model possesses higher spatial continuity and semantic expressive power. The system can accurately determine whether a vehicle is within the road area, whether it is crossing the lane lines, whether it is deviating from the lane center, and whether it is driving in the wrong direction, based on road geometric features and lane structure. Simultaneously, by introducing contextual information such as the preceding and following road morphology and drivable distance into the abstract points, the teaching judgment is no longer based on a single point state but on reasoning based on the characteristics of a road segment, thereby significantly improving the accuracy and stability of road driving instruction and road condition judgment.
[0085] This invention introduces a vehicle driving trend reasoning mechanism into the teaching process. By analyzing the vehicle's driving direction and predicted trajectory, and combining this with environmental information such as other vehicles, obstacles, and electronic fences, potential safety risks can be identified in advance. Based on this, the system can naturally transform safety-related risk scenarios into teaching content, such as right-of-way prompts when vehicles meet, safe distance instruction when driving side-by-side, and guidance on slowing down and avoiding obstacles. This creates a unified reasoning link between safety control logic and teaching logic, thereby enhancing learners' understanding of safe driving behavior while ensuring driving safety.
[0086] This invention does not rely on manually preset fixed teaching locations. Instead, it dynamically generates teaching content and practice guidance based on abstract road features and the current state of the vehicle, enabling the teaching system to adapt to different road structures and training grounds. By abstracting road characteristics and combining them with teaching needs for reasoning, the system can construct teaching scenarios in real time during actual driving, breaking through the limitations of traditional fixed sites and routes. This improves the adaptability and flexibility of the driving teaching system in complex and changing road environments.
[0087] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for generating training samples for a teaching language model for Subject 3, characterized in that: Includes the following steps: Multi-source data acquisition and state fusion are performed to acquire static input data and dynamic input data. The static input data includes data describing the driving teaching environment, and the dynamic input data includes data describing the vehicle's operating status and the student's operating behavior. The two data are then fused together, and the current driving teaching scenario type is determined by combining the vehicle status, road environment, closed site information, and subject description. Semantic abstraction and language state description generation: Based on preset semantic abstraction rules, feature extraction and rule judgment are performed on the fused state data to generate a unified language state description with clear semantics; The key operation history is constructed and time-series enhanced. It filters key student operations related to the current subject and teaching objectives, records them in chronological order and associates them with corresponding semantic states to form a key operation history sequence. Post-processing strategies can be configured according to needs. Sample input construction and multi-source annotation fusion: The language state description, key operation history sequence and subject description information are combined into a unified format sample input, multi-source teaching annotations are fused and consistency verification and conflict resolution are performed to generate the final teaching annotations; Complete sample output: Integrate sample input with final teaching annotations to output a complete sample for training the language model for Subject 3 teaching.
2. The method according to claim 1, characterized in that: The static input data includes structured information of the driving school's closed training ground or roads, subject description information, teaching rules, and key operation configuration parameters; the dynamic input data includes vehicle component status data, vehicle position information, and real-time student operation behavior data; the structured information of the driving school's closed training ground or roads includes road type, road direction, intersection information, number of lanes, and site constraints related to the Subject 3 exam; the subject description information is used to indicate the current Subject 3 sub-item stage and practice objectives of the teaching or training; the teaching rules and key operation configuration parameters are used to indicate the key operation categories that need to be focused on under different subjects and different student conditions.
3. The method according to claim 2, characterized in that: The vehicle component status data includes vehicle speed, gear position, braking status, and steering status; the vehicle position and posture information includes position, heading angle, and motion trend; the student's real-time operation behavior data is used to determine whether key operations such as gear shifting, braking, and turn signal operation have occurred.
4. The method according to claim 1, characterized in that: The scenario types include scenarios related to driving in a straight line, turning, passing through intersections, deceleration control, and gear shifting, etc.
5. The method according to claim 1, characterized in that: The semantic abstraction and rule-based judgment specifically include extracting the following state features that have pedagogical significance: Current driving behavior trends, including acceleration, deceleration, and stable driving; Does the current operation meet the teaching requirements of Subject 3? Potential teaching focuses in the current context.
6. The method according to claim 1, characterized in that: The unified language state description can explicitly express the following information: The current driving scenario of the vehicle; The student's current operational status and its rationality; Contextual information relevant to the current teaching objectives of Subject 3.
7. The method according to claim 1, characterized in that: The key operations include gear shifting, braking, turn signal operation, steering wheel steering, throttle control, and other operations related to the teaching objectives of Subject 3.
8. The method according to claim 1, characterized in that: The key operation history record is a sequence of key operations that are filtered and recorded in chronological order and associated with corresponding semantic states.
9. The method according to claim 1, characterized in that: The configurable post-processing strategy includes adjusting key operation types, historical record lengths, or recording rules based on student characteristics and the needs of the Subject 3 teaching project to achieve personalized sample generation.
10. The method according to any one of claims 1 to 9, characterized in that: The sample input construction involves combining the language state description generated at the current moment, the key operation history sequence, and the subject description information to form a sample input in a unified format. The multi-source instructional annotations include: Teaching annotations generated based on the driving test (subject 3) teaching algorithm deployed on real vehicles; Simulation teaching annotations generated based on the distribution of real test data; Manually annotated teaching instructions or correction results; The consistency verification and conflict resolution process involves analyzing the consistency of instructional annotations from different sources. When conflicts exist, they are resolved using preset rules or priority strategies to generate the final instructional annotation results for training.