Rail transit AR training and operation optimization method and system based on AI intelligent driving
By collecting multimodal data through head-mounted AR devices to build personalized portraits, and combining AI algorithms and high-precision 3D modeling, personalized AR training is provided to solve the problem of insufficient adaptation of existing AR training systems in rail transit, and realize efficient and flexible training content generation and interactive response.
Patent Information
- Application Number
- CN202510793044.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing AR training system lacks personalized adaptation capabilities in the rail transit field, has low efficiency and poor flexibility in training resource construction, insufficient environmental perception accuracy, fragmented interactive experience, single feedback mechanism, insufficient intelligent support, and weak ability to handle complex problems.
Through head-mounted AR devices, multimodal data of students is collected, personalized portraits are built based on machine learning, AR training resources are matched with AI algorithms, and centimeter-level 3D modeling is performed using the improved YOLOv8 model and SLAM technology. Virtual information display and remote expert assistance are provided, and training content is dynamically updated to meet student needs.
It achieves precise customization of personalized training content, improves training efficiency and immersion, solves the homogeneity limitations of traditional training, and significantly improves training effects and interactive response capabilities in complex scenarios.
Smart Images

Figure CN120656351A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of rail transit training and operation assistance technology, and specifically to a rail transit AR training and operation optimization method and system based on AI intelligent drive. Background Art
[0002] With the rapid development of augmented reality (AR) technology and its deep integration with artificial intelligence (AI), it has shown great potential in vocational skills training, particularly in training and operational guidance in complex industrial sectors such as rail transit. Existing AR training systems typically simulate actual operational processes by constructing virtual equipment models or overlaying information on the real environment. This allows trainees to engage in immersive learning and practice in this virtual-realistic environment, effectively improving operational skills, safety awareness, and emergency response capabilities.
[0003] However, although the existing AR training system has improved the intuitiveness and interactivity of training to a certain extent, it still has several significant defects that need to be addressed when applied to large-scale, high-standard rail transit professional training and operation optimization: lack of personalized adaptation capabilities; low efficiency and poor flexibility in building training resources; insufficient environmental perception accuracy and fragmented interactive experience; single feedback mechanism and lack of closed-loop optimization; insufficient intelligent support and weak ability to handle complex problems.
[0004] Therefore, improvements are needed. Summary of the Invention
[0005] To solve the above technical problems, this application provides a rail transit AR training and operation optimization method and system based on AI intelligent drive.
[0006] The first object of the invention of this application is achieved through the following technical solutions:
[0007] An AI-driven rail transit AR training and operation optimization method includes the following steps:
[0008] S1: When multimodal data of students is collected through head-mounted AR devices, a personalized portrait of the students is constructed based on machine learning algorithms;
[0009] S2: When a drag operation signal is received, the AI algorithm is used to match candidate models from the preset 3D model library and interactive elements from the preset interaction database, and several AR training resources are integrated and stored.
[0010] S3: Generate customized training content based on the personalized portrait and the corresponding AR training resources, and display the customized training content in AR form overlaid on the trainee's real field of view through the head-mounted AR device;
[0011] S4: Based on the improved YOLOv8 model combined with SLAM technology, centimeter-level 3D modeling and spatial positioning of the trainee's rail transit environment are performed;
[0012] S5: When receiving a student's voice or gesture interaction request, combining the 3D modeling and spatial positioning results, providing virtual information display, intelligent question and answer, and remote expert assistance;
[0013] S6: Continuously collect trainees' training data during the training process, wherein the training data includes operation data, interaction data and multimodal data, evaluate the learning status in real time based on the training data, dynamically update the personalized portrait and trigger the adjustment of the training content.
[0014] In a preferred embodiment, the step S1: when the multimodal data of the student is collected through the head-mounted AR device, constructing a personalized portrait of the student based on a machine learning algorithm includes:
[0015] S11: The multimodal data includes eye movement trajectory data, operation behavior data, voice interaction data and physiological signal data;
[0016] S12: Analyze the eye movement trajectory data to identify the trainee's visual focus pattern and the duration and frequency of their focus shifts on AR interface elements;
[0017] S13: Analyze the time series of the operation behavior data, and calculate the time intervals, error rates, and operation fluency indicators between key operation steps;
[0018] S14: Analyze the content semantics of the voice interaction data to identify the types of questions raised by the students, areas of concern, and potential knowledge blind spots;
[0019] S15: Analyze the physiological signal data to assess the student's real-time cognitive load status, where the cognitive load status includes low load, moderate load, and high load;
[0020] S16: Integrate the analysis results of steps S12 to S15 and use the pre-trained machine learning model to build a personalized portrait.
[0021] In a preferred embodiment, the step S2: when a drag operation signal is received, matching candidate models from a preset 3D model library and matching interactive elements from a preset interactive database based on an AI algorithm, integrating and generating a plurality of AR training resources, and storing the resources, includes:
[0022] S21: parsing the operation intention of the drag operation signal based on a predefined rule library and an intention recognition algorithm, where the operation intention includes model addition, model replacement, scenario construction, or fault setting;
[0023] S22: Applying a spatial analysis algorithm to accurately calculate a spatial bounding box of the target area defined by the drag operation signal;
[0024] S23: Preliminarily screening a set of models with matching functions from the preset three-dimensional model library according to the identified operation intention;
[0025] S24: Calculating the spatial overlap ratio between the bounding box of each model in the filtered model set and the bounding box of the target area;
[0026] S25: Select the top N models with the highest spatial overlap ratio as candidate models, where N is a preset positive integer;
[0027] S26: Filtering a set of interactive elements with matching functions from the preset interaction database according to the identified operation intention and preset keyword matching rules;
[0028] S27: Intelligently combine and associate the matched candidate models and interactive element sets to generate AR training resources and store them in the resource library.
[0029] In a preferred embodiment, the step of S3: generating customized training content based on matching the personalized portrait with corresponding AR training resources, and displaying the customized training content in AR form on the trainee's real field of view through the head-mounted AR device, includes:
[0030] S31: Extracting the student's current skill proficiency level, learning preference feature vector, and cognitive style identifier from the personalized portrait;
[0031] S32: Calculate the matching score between the trainee's skill proficiency level and the preset difficulty level of the AR training resource;
[0032] S33: Calculate the similarity score between the learner's learning preference feature vector and the AR training resource content type feature vector;
[0033] S34: Calculate the compatibility score between the trainee’s cognitive style identification and the preset presentation mode of the AR training resources;
[0034] S35: Based on the matching score, the similarity score, and the adaptability score, a multi-criteria decision algorithm with preset weights is used to generate a comprehensive priority ranking of all available AR training resources;
[0035] S36: Selecting the best matching AR training resource combination based on the comprehensive priority ranking to generate customized training content for the current stage;
[0036] S37: Through the display module of the head-mounted AR device, the customized training content is accurately superimposed and displayed in the trainee's real environment field of view in a preset AR format.
[0037] In a preferred embodiment, the step S4: performing centimeter-level 3D modeling and spatial positioning of the trainee's rail transit environment based on an improved YOLOv8 model combined with SLAM technology, includes:
[0038] S41: Using the built-in sensor of the head-mounted AR device to collect environmental image streams and motion data in real time;
[0039] S42: Apply the improved YOLOv8 model to perform real-time target detection on the collected image stream to accurately identify the equipment, facilities and trainee locations in the rail transit environment;
[0040] S43: Combined with the SLAM algorithm, the recognition result of step S42 is integrated with the motion data to construct and update the three-dimensional model of the environment with centimeter-level accuracy in real time, and simultaneously and accurately track the six-degree-of-freedom posture of the head-mounted AR device in the environment.
[0041] In a preferred embodiment, the step of S5: upon receiving a student's voice or gesture interaction request, combining the 3D modeling and spatial positioning results to provide virtual information display, intelligent question and answer, and remote expert assistance, includes:
[0042] S51: Capture and recognize students’ voice commands or preset gestures in real time;
[0043] S52: Combining the centimeter-level three-dimensional environment model generated in step S4 and the student's current spatial positioning result, analyzing the student's interaction request intention and the environmental entity or spatial location it points to;
[0044] S53: Based on the parsed intention and position, overlay and display relevant virtual information at the correct spatial position in the field of view of the head-mounted AR device;
[0045] S54: Calling a preset knowledge base to provide intelligent questions and answers related to the request;
[0046] S55: When the request exceeds the system's preset processing capacity, an audio and video communication channel with the remote expert is automatically established, and the student's real-time AR vision and environmental three-dimensional model are shared with the expert.
[0047] In a preferred embodiment, the step S6: continuously collecting trainee training data during the training process, wherein the training data includes operation data, interaction data, and multimodal data, and evaluating the learning status in real time based on the training data, dynamically updating the personalized profile, and triggering the adjustment of the training content, includes:
[0048] S61: The operation data includes operation steps, time, and error points;
[0049] S62: The interaction data includes query content and question and answer records;
[0050] S63: Based on the training data, applying a preset learning status evaluation model to calculate evaluation indicators, wherein the evaluation indicators include the trainee's mastery of the current training content, operational proficiency, and cognitive load level;
[0051] S64: Dynamically update the student personalized profile constructed in step S1 based on the evaluation indicator results;
[0052] S65: Based on the updated personalized profile and the current learning status assessment results, determine whether the need to adjust the training content is triggered;
[0053] S66: When an adjustment requirement is identified, step S3 is triggered to generate new customized training content.
[0054] The second object of the present invention is achieved through the following technical solutions:
[0055] Module 1: When multimodal data of students is collected through head-mounted AR devices, a personalized profile of the students is constructed based on machine learning algorithms;
[0056] The second module: When a drag operation signal is received, it uses AI algorithms to match candidate models from a preset 3D model library and interactive elements from a preset interactive database, integrating and generating several AR training resources and storing them;
[0057] The third module: generating customized training content based on matching the personalized portrait with corresponding AR training resources, and displaying the customized training content in AR form overlaid on the trainee's real field of view through the head-mounted AR device;
[0058] Module 4: Based on the improved YOLOv8 model combined with SLAM technology, centimeter-level 3D modeling and spatial positioning of the trainees' rail transit environment;
[0059] Module 5: When receiving a student's voice or gesture interaction request, the system combines the 3D modeling and spatial positioning results to provide virtual information display, intelligent question and answer, and remote expert assistance;
[0060] Module 6: Continuously collect students' training data during the training process. The training data includes operation data, interaction data and multimodal data. Based on the training data, the learning status is evaluated in real time, the personalized portrait is dynamically updated and the adjustment of the training content is triggered.
[0061] The third object of the present invention is achieved through the following technical solutions:
[0062] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned AI-driven rail transit AR training and operation optimization method are implemented.
[0063] The fourth object of the present invention is achieved through the following technical solutions:
[0064] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned rail transit AR training and operation optimization method based on AI intelligent driving.
[0065] In summary, this application includes at least one of the following beneficial technical effects:
[0066] 1. When the system collects multimodal data (eye movement trajectory, operation behavior, voice interaction, physiological signals) of trainees through head-mounted AR devices, a personalized portrait is constructed based on machine learning algorithms. Specifically: by analyzing eye movement trajectories to identify visual attention patterns and interface interaction depth, parsing operation behavior data to quantify step intervals, error rates and fluency, exploring problem types and knowledge blind spots in voice interaction, and evaluating cognitive load status (low / medium / high) in real time based on physiological signals. Integrating the above multi-dimensional analysis results, a pre-trained model (such as random forest or neural network) is used to generate a dynamically updated student portrait. Effect: Break through the homogeneous limitations of traditional training, accurately portray individual differences of trainees, and lay a data foundation for customized training.
[0067] 2. In response to drag operation signals (gesture / controller circling), the system quickly generates AR training resources based on AI algorithms. First, the operation intention (such as model addition, fault setting) is parsed through the rule library and intent recognition algorithm, and the bounding box of the target area is calculated in combination with the spatial analysis algorithm; then, the function matching model is screened from the three-dimensional model library, the spatial overlap ratio between its bounding box and the target area is calculated, and the Top-N candidate models are selected; simultaneously, the interactive element library is matched based on the intent keywords, and finally the model and interactive element are intelligently associated to generate resources and store them. Effect: Replaces the tedious process of manual modeling, realizes "what you see is what you get" resource construction, and significantly improves the efficiency and flexibility of creating specific training content for rail transit scenarios (such as equipment failure simulation). BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 This is a flowchart of an implementation of an embodiment of an AI-driven rail transit AR training and operation optimization method of the present application;
[0069] Figure 2This is a flowchart for implementing step S1 in an embodiment of an AI-driven rail transit AR training and operation optimization method of the present application;
[0070] Figure 3 This is a flowchart for implementing step S2 in an embodiment of an AI-driven rail transit AR training and operation optimization method of the present application;
[0071] Figure 4 This is a principle block diagram of a computer device of the present application. DETAILED DESCRIPTION
[0072] The following is combined with Figure 1-4 This application is described in further detail.
[0073] In one embodiment, if Figure 1 As shown, this application discloses an AI-driven rail transit AR training and operation optimization method, which specifically includes the following steps:
[0074] S1: When multimodal data of students is collected through head-mounted AR devices, a personalized portrait of the students is constructed based on machine learning algorithms;
[0075] S2: When a drag operation signal is received, the AI algorithm is used to match candidate models from the preset 3D model library and interactive elements from the preset interaction database, and several AR training resources are integrated and stored.
[0076] S3: Generate customized training content based on the personalized portrait and the corresponding AR training resources, and display the customized training content in AR form overlaid on the trainee's real field of view through the head-mounted AR device;
[0077] S4: Based on the improved YOLOv8 model combined with SLAM technology, centimeter-level 3D modeling and spatial positioning of the trainee's rail transit environment are performed;
[0078] S5: When receiving a student's voice or gesture interaction request, combining the 3D modeling and spatial positioning results, providing virtual information display, intelligent question and answer, and remote expert assistance;
[0079] S6: Continuously collect trainees' training data during the training process, wherein the training data includes operation data, interaction data and multimodal data, evaluate the learning status in real time based on the training data, dynamically update the personalized portrait and trigger the adjustment of the training content.
[0080] In this embodiment, the system uses multimodal data to drive the construction of personalized portraits (S1), combines intelligent resource generation (S2) with centimeter-level environmental perception (S4), and achieves precise delivery of customized AR training content (S3); when responding to interactive requests, it integrates high-precision spatial positioning to provide virtual-reality fusion guidance (S5), and dynamically optimizes teaching in a closed loop based on continuously collected training data (S6).
[0081] like Figure 2 As shown, step S1 includes:
[0082] S11: The multimodal data includes eye movement trajectory data, operation behavior data, voice interaction data and physiological signal data;
[0083] S12: Analyze the eye movement trajectory data to identify the trainee's visual focus pattern and the duration and frequency of their focus shifts on AR interface elements;
[0084] S13: Analyze the time series of the operation behavior data, and calculate the time intervals, error rates, and operation fluency indicators between key operation steps;
[0085] S14: Analyze the content semantics of the voice interaction data to identify the types of questions raised by the students, areas of concern, and potential knowledge blind spots;
[0086] S15: Analyze the physiological signal data to assess the student's real-time cognitive load status, where the cognitive load status includes low load, moderate load, and high load;
[0087] S16: Integrate the analysis results of steps S12 to S15 and use the pre-trained machine learning model to build a personalized portrait.
[0088] In this embodiment, when the system collects multimodal data (eye movement trajectory, operation behavior, voice interaction, physiological signals) of the trainees through the head-mounted AR device, a personalized portrait is constructed based on the machine learning algorithm. Specifically: by analyzing the eye movement trajectory to identify the visual attention pattern and the depth of interface interaction, parsing the operation behavior data to quantify the step interval, error rate and fluency, exploring the problem types and knowledge blind spots in the voice interaction, and evaluating the cognitive load status (low / medium / high) in real time based on physiological signals. Integrating the above multi-dimensional analysis results, a pre-trained model (such as random forest or neural network) is used to generate a dynamically updated student portrait. Effect: Break through the homogeneity limitations of traditional training, accurately portray the individual differences of trainees, and lay a data foundation for customized training.
[0089] like Figure 3 As shown, step S2 includes:
[0090] S21: parsing the operation intention of the drag operation signal based on a predefined rule library and an intention recognition algorithm, where the operation intention includes model addition, model replacement, scenario construction, or fault setting;
[0091] S22: Applying a spatial analysis algorithm to accurately calculate a spatial bounding box of the target area defined by the drag operation signal;
[0092] S23: Preliminarily screening a set of models with matching functions from the preset three-dimensional model library according to the identified operation intention;
[0093] S24: Calculating the spatial overlap ratio between the bounding box of each model in the filtered model set and the bounding box of the target area;
[0094] S25: Select the top N models with the highest spatial overlap ratio as candidate models, where N is a preset positive integer;
[0095] S26: Filtering a set of interactive elements with matching functions from the preset interaction database according to the identified operation intention and preset keyword matching rules;
[0096] S27: Intelligently combine and associate the matched candidate models and interactive element sets to generate AR training resources and store them in the resource library.
[0097] In this embodiment, for drag operation signals (gesture / controller circling), the system quickly generates AR training resources based on AI algorithms. First, the operation intention (such as model addition, fault setting) is parsed through the rule library and intent recognition algorithm, and the target area bounding box is calculated in combination with the spatial analysis algorithm; then, the function matching model is screened from the three-dimensional model library, the spatial overlap ratio between its bounding box and the target area is calculated, and the Top-N candidate model is selected; synchronously, the interactive element library is matched based on the intent keyword, and finally the model and interactive element are intelligently associated to generate resources and store them. Effect: Replaces the tedious process of manual modeling, realizes "what you see is what you get" resource construction, and significantly improves the efficiency and flexibility of creating specific training content for rail transit scenarios (such as equipment failure simulation).
[0098] The S3 step includes:
[0099] S31: Extracting the student's current skill proficiency level, learning preference feature vector, and cognitive style identifier from the personalized portrait;
[0100] S32: Calculate the matching score between the trainee's skill proficiency level and the preset difficulty level of the AR training resource;
[0101] S33: Calculate the similarity score between the learner's learning preference feature vector and the AR training resource content type feature vector;
[0102] S34: Calculate the compatibility score between the trainee’s cognitive style identification and the preset presentation mode of the AR training resources;
[0103] S35: Based on the matching score, the similarity score, and the adaptability score, a multi-criteria decision algorithm with preset weights is used to generate a comprehensive priority ranking of all available AR training resources;
[0104] S36: Selecting the best matching AR training resource combination based on the comprehensive priority ranking to generate customized training content for the current stage;
[0105] S37: Through the display module of the head-mounted AR device, the customized training content is accurately superimposed and displayed in the trainee's real environment field of view in a preset AR format.
[0106] In this embodiment, the system dynamically customizes and precisely delivers training content based on personalized portraits. The system extracts skill proficiency, learning preferences, and cognitive style characteristics from the portraits, and calculates their matching, similarity, and compatibility scores with the difficulty, content type, and presentation method (visual / auditory / interactive) of the AR resources. A weighted multi-criteria decision-making algorithm (such as TOPSIS) is used to prioritize resources and assemble the optimal training content. The AR device accurately overlays virtual information onto the trainee's real-world field of view. The result: This addresses the pain point of "one-size-fits-all" training, ensuring that content is deeply aligned with the trainee's abilities, preferences, and cognitive characteristics, improving learning efficiency and immersion.
[0107] Step S4 includes:
[0108] S41: Using the built-in sensor of the head-mounted AR device to collect environmental image streams and motion data in real time;
[0109] S42: Apply the improved YOLOv8 model to perform real-time target detection on the collected image stream to accurately identify the equipment, facilities and trainee locations in the rail transit environment;
[0110] S43: Combined with the SLAM algorithm, the recognition result of step S42 is integrated with the motion data to construct and update the three-dimensional model of the environment with centimeter-level accuracy in real time, and simultaneously and accurately track the six-degree-of-freedom posture of the head-mounted AR device in the environment.
[0111] In this example, an improved YOLOv8 model (introducing an attention mechanism to optimize rail transit equipment recognition) is integrated with SLAM technology to achieve centimeter-level reconstruction and positioning of the environment. AR device sensors (RGB-D camera, IMU) are used to collect environmental image streams and motion data. YOLOv8 is improved to detect key equipment such as signal lights and tracks in real time. The SLAM algorithm fuses detection results with motion data to construct a centimeter-level accurate 3D environmental model and simultaneously track the equipment's six-degree-of-freedom pose. This approach overcomes the challenge of virtual information dislocation in complex rail transit scenarios, providing underlying support for precise operational guidance (such as switch adjustment) and spatial interaction.
[0112] Step S5 includes:
[0113] S51: Capture and recognize students’ voice commands or preset gestures in real time;
[0114] S52: Combining the centimeter-level three-dimensional environment model generated in step S4 and the student's current spatial positioning result, analyzing the student's interaction request intention and the environmental entity or spatial location it points to;
[0115] S53: Based on the parsed intention and position, overlay and display relevant virtual information at the correct spatial position in the field of view of the head-mounted AR device;
[0116] S54: Calling a preset knowledge base to provide intelligent questions and answers related to the request;
[0117] S55: When the request exceeds the system's preset processing capacity, an audio and video communication channel with the remote expert is automatically established, and the student's real-time AR vision and environmental three-dimensional model are shared with the expert.
[0118] In this embodiment, when responding to student voice or gesture requests, the system combines a high-precision environmental model with positioning data to analyze intent. After recognizing the command, it locates the entity or location it points to and overlays virtual information (such as device parameters) at the corresponding spatial coordinates. It then draws upon the knowledge base to provide intelligent Q&A. When questions exceed a preset level of complexity, it automatically establishes a remote expert channel and shares the student's AR field of view and environmental model. The result: precise, "point-and-answer" interaction is achieved, breaking down knowledge barriers between on-site and remote locations through "first-person perspective sharing," improving the efficiency of complex troubleshooting.
[0119] Step S6 includes:
[0120] S61: The operation data includes operation steps, time, and error points;
[0121] S62: The interaction data includes query content and question and answer records;
[0122] S63: Based on the training data, applying a preset learning status evaluation model to calculate evaluation indicators, wherein the evaluation indicators include the trainee's mastery of the current training content, operational proficiency, and cognitive load level;
[0123] S64: Dynamically update the student personalized profile constructed in step S1 based on the evaluation indicator results;
[0124] S65: Based on the updated personalized profile and the current learning status assessment results, determine whether the need to adjust the training content is triggered;
[0125] S66: When an adjustment requirement is identified, step S3 is triggered to generate new customized training content.
[0126] In this embodiment, operational data (steps / time / error points), interaction data (queries / questions and answers), and multimodal data are continuously collected during training. A learning state assessment model is used to calculate mastery, proficiency, and cognitive load metrics in real time. Based on this data, the student profile is dynamically updated and the need for content adjustment (e.g., strengthening weak links) is determined. If adjustments are necessary, S3 is re-executed to generate new customized content. The result: a closed loop of "assessment-feedback-optimization" is formed, ensuring that training content evolves dynamically as students' abilities grow, maximizing training effectiveness.
[0127] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0128] In one embodiment, a rail transit AR training and operation optimization system based on AI intelligent driving is provided. The rail transit AR training and operation optimization system based on AI intelligent driving corresponds to the rail transit AR training and operation optimization method based on AI intelligent driving in the above embodiment. The rail transit AR training and operation optimization system based on AI intelligent driving includes:
[0129] Module 1: When multimodal data of students is collected through head-mounted AR devices, a personalized profile of the students is constructed based on machine learning algorithms;
[0130] The second module: When a drag operation signal is received, it uses AI algorithms to match candidate models from a preset 3D model library and interactive elements from a preset interactive database, integrating and generating several AR training resources and storing them;
[0131] The third module: generating customized training content based on matching the personalized portrait with corresponding AR training resources, and displaying the customized training content in AR form overlaid on the trainee's real field of view through the head-mounted AR device;
[0132] Module 4: Based on the improved YOLOv8 model combined with SLAM technology, centimeter-level 3D modeling and spatial positioning of the trainees' rail transit environment;
[0133] Module 5: When receiving a student's voice or gesture interaction request, the system combines the 3D modeling and spatial positioning results to provide virtual information display, intelligent question and answer, and remote expert assistance;
[0134] Module 6: Continuously collect students' training data during the training process. The training data includes operation data, interaction data and multimodal data. Based on the training data, the learning status is evaluated in real time, the personalized portrait is dynamically updated and the adjustment of the training content is triggered.
[0135] Optionally, also include:
[0136] Module 7: The multimodal data includes eye movement data, operation behavior data, voice interaction data and physiological signal data;
[0137] Module 8: Analyze the eye movement data to identify the trainees’ visual focus patterns and their dwell time and shift frequency on AR interface elements;
[0138] Module 9: Analyze the time series of the operation behavior data and calculate the time intervals, error rates and operation fluency indicators between key operation steps;
[0139] Module 10: Analyze the content and semantics of the voice interaction data to identify the types of questions raised by students, areas of concern, and potential knowledge blind spots;
[0140] Module 11: Analyze the physiological signal data to assess the trainee's real-time cognitive load status, where the cognitive load status includes low load, moderate load, and high load;
[0141] Module 12: Integrate the analysis results of steps S12 to S15 and use the pre-trained machine learning model to build a personalized portrait.
[0142] Optionally, also include:
[0143] Module 13: Analyzes the operation intention of the drag operation signal based on a predefined rule library and intention recognition algorithm. The operation intention includes model addition, model replacement, scenario construction or fault setting.
[0144] Module 14: Applying a spatial analysis algorithm to accurately calculate the spatial bounding box of the target area defined by the drag operation signal;
[0145] Module 15: Preliminary screening of a set of functionally matching models from the preset three-dimensional model library based on the identified operation intention;
[0146] Module 16: Calculate the spatial overlap ratio between the bounding box of each model in the selected model set and the bounding box of the target area;
[0147] Module 17: Select the top N models with the highest spatial overlap ratio as candidate models, where N is a preset positive integer;
[0148] Module 18: Filtering a set of interactive elements with matching functions from the preset interactive database according to the identified operation intention and preset keyword matching rules;
[0149] Module 19: Intelligently combine and associate the matched candidate models and interactive element sets, generate AR training resources and store them in the resource library.
[0150] Optionally, also include
[0151] Module 20: Extracting the student's current skill proficiency level, learning preference feature vector, and cognitive style identifier from the personalized portrait;
[0152] Module 21: Calculate the matching score between the trainee's skill proficiency level and the preset difficulty level of the AR training resources;
[0153] Module 22: Calculate the similarity score between the learner’s learning preference feature vector and the AR training resource content type feature vector;
[0154] Module 2 and 3: Calculate the compatibility score between the students’ cognitive style identification and the preset presentation mode of AR training resources;
[0155] Module 24: Based on the matching score, the similarity score, and the adaptability score, a multi-criteria decision-making algorithm with preset weights is used to generate a comprehensive priority ranking of all available AR training resources;
[0156] Module 25: Based on the comprehensive priority ranking, select the best matching AR training resource combination and generate customized training content for the current stage;
[0157] Module 26: Through the display module of the head-mounted AR device, the customized training content is accurately superimposed and displayed in the trainee's real environment field of view in a preset AR format.
[0158] Optionally, also include:
[0159] Module 27: uses the built-in sensors of the head-mounted AR device to collect environmental image streams and motion data in real time;
[0160] Module 28: Apply the improved YOLOv8 model to perform real-time target detection on the collected image stream and accurately identify the equipment, facilities and trainee locations in the rail transit environment;
[0161] Module 29: Combined with the SLAM algorithm, the recognition result of step S42 is integrated with the motion data to construct and update the three-dimensional model of the environment with centimeter-level accuracy in real time, and simultaneously and accurately track the six-degree-of-freedom posture of the head-mounted AR device in the environment.
[0162] Optionally, also include:
[0163] Module 30: Real-time capture and recognition of students’ voice commands or preset gestures;
[0164] Trinity module: Combines the centimeter-level 3D environment model generated in step S4 with the student's current spatial positioning result to analyze the student's interaction request intention and the environmental entity or spatial location it points to;
[0165] Module 32: Based on the parsed intent and location, relevant virtual information is overlaid and displayed at the correct spatial location in the field of view of the head-mounted AR device;
[0166] Module 33: Calls the pre-built knowledge base to provide intelligent questions and answers related to the request;
[0167] Modules 3 and 4: When the request exceeds the system's preset processing capacity, an audio and video communication channel with the remote expert is automatically established, and the student's real-time AR vision and environmental three-dimensional model are shared with the expert.
[0168] Optionally, also include:
[0169] Module 35: The operation data includes operation steps, time, and error points;
[0170] 36 module: The interaction data includes query content and question and answer records;
[0171] 37 module: Based on the training data, a preset learning status assessment model is applied to calculate assessment indicators, which include the trainee's mastery of the current training content, operational proficiency, and cognitive load level;
[0172] 38 module: Dynamically update the personalized student portrait constructed in step S1 based on the evaluation index results;
[0173] Module 39: Based on the updated personalized profile and current learning status assessment results, determine whether the need for training content adjustment is triggered;
[0174] Module 40: When an adjustment need is identified, step S3 is triggered to generate new customized training content.
[0175] Regarding the specific definition of a rail transit AR training and operation optimization system based on AI intelligence drive, please refer to the definition of a rail transit AR training and operation optimization method based on AI intelligence drive above, which will not be repeated here. Each module in the above-mentioned rail transit AR training and operation optimization system based on AI intelligence drive can be implemented in whole or in part through software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0176] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store training content. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a rail transit AR training and operation optimization method based on AI intelligent drive is realized.
[0177] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a rail transit AR training and operation optimization method based on AI intelligent driving is implemented.
[0178] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, a rail transit AR training and operation optimization method based on AI intelligent driving is provided.
[0179] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0180] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
Claims
1. A rail transit AR training and operation optimization method based on AI intelligent driving, characterized by: Including steps: S1: When multimodal data of students is collected through head-mounted AR devices, a personalized portrait of the students is constructed based on machine learning algorithms; S2: When a drag operation signal is received, the AI algorithm is used to match candidate models from the preset 3D model library and interactive elements from the preset interaction database, and several AR training resources are integrated and stored. S3: Generate customized training content based on the personalized portrait and the corresponding AR training resources, and display the customized training content in AR form overlaid on the trainee's real field of view through the head-mounted AR device; S4: Based on the improved YOLOv8 model combined with SLAM technology, centimeter-level 3D modeling and spatial positioning of the trainee's rail transit environment are performed; S5: When receiving a student's voice or gesture interaction request, combining the 3D modeling and spatial positioning results, providing virtual information display, intelligent question and answer, and remote expert assistance; S6: Continuously collect trainees' training data during the training process, wherein the training data includes operation data, interaction data and multimodal data, evaluate the learning status in real time based on the training data, dynamically update the personalized portrait and trigger the adjustment of the training content.
2. The rail transit AR training and operation optimization method based on AI intelligent driving according to claim 1 is characterized in that: S1: When the multimodal data of the student is collected through the head-mounted AR device, the step of constructing a personalized portrait of the student based on the machine learning algorithm includes: S11: The multimodal data includes eye movement trajectory data, operation behavior data, voice interaction data and physiological signal data; S12: Analyze the eye movement trajectory data to identify the trainee's visual focus pattern and the duration and frequency of their focus shifts on AR interface elements; S13: Analyze the time series of the operation behavior data, and calculate the time intervals, error rates, and operation fluency indicators between key operation steps; S14: Analyze the content semantics of the voice interaction data to identify the types of questions raised by the students, areas of concern, and potential knowledge blind spots; S15: Analyze the physiological signal data to assess the student's real-time cognitive load status, where the cognitive load status includes low load, moderate load, and high load; S16: Integrate the analysis results of steps S12 to S15 and use the pre-trained machine learning model to build a personalized portrait.
3. The rail transit AR training and operation optimization method based on AI intelligent driving according to claim 1 is characterized in that: The step S2: when a drag operation signal is received, matching candidate models from a preset 3D model library and matching interactive elements from a preset interactive database based on an AI algorithm, integrating and generating a plurality of AR training resources and storing them, includes: S21: parsing the operation intention of the drag operation signal based on a predefined rule library and an intention recognition algorithm, where the operation intention includes model addition, model replacement, scenario construction, or fault setting; S22: Applying a spatial analysis algorithm to accurately calculate a spatial bounding box of the target area defined by the drag operation signal; S23: Preliminarily screening a set of models with matching functions from the preset three-dimensional model library according to the identified operation intention; S24: Calculating the spatial overlap ratio between the bounding box of each model in the filtered model set and the bounding box of the target area; S25: Select the top N models with the highest spatial overlap ratio as candidate models, where N is a preset positive integer; S26: Filtering a set of interactive elements with matching functions from the preset interaction database according to the identified operation intention and preset keyword matching rules; S27: Intelligently combine and associate the matched candidate models and interactive element sets to generate AR training resources and store them in the resource library.
4. The AI-driven rail transit AR training and operation optimization method according to claim 1 is characterized in that: The step S3: generating customized training content based on matching the personalized portrait with corresponding AR training resources, and displaying the customized training content in AR form on the trainee's real field of view through the head-mounted AR device, includes: S31: Extracting the student's current skill proficiency level, learning preference feature vector, and cognitive style identifier from the personalized portrait; S32: Calculate the matching score between the trainee's skill proficiency level and the preset difficulty level of the AR training resource; S33: Calculate the similarity score between the learner's learning preference feature vector and the AR training resource content type feature vector; S34: Calculate the compatibility score between the trainee’s cognitive style identification and the preset presentation mode of the AR training resources; S35: Based on the matching score, the similarity score, and the adaptability score, a multi-criteria decision algorithm with preset weights is used to generate a comprehensive priority ranking of all available AR training resources; S36: Selecting the best matching AR training resource combination based on the comprehensive priority ranking to generate customized training content for the current stage; S37: Through the display module of the head-mounted AR device, the customized training content is accurately superimposed and displayed in the trainee's real environment field of view in a preset AR format.
5. The AI intelligent driven rail transit AR training and operation optimization method according to claim 1 is characterized in that: S4: Based on the improved YOLOv8 model combined with SLAM technology, the steps of performing centimeter-level 3D modeling and spatial positioning of the trainee's rail transit environment include: S41: Using the built-in sensor of the head-mounted AR device to collect environmental image streams and motion data in real time; S42: Apply the improved YOLOv8 model to perform real-time target detection on the collected image stream to accurately identify the equipment, facilities and trainee locations in the rail transit environment; S43: Combined with the SLAM algorithm, the recognition result of step S42 is integrated with the motion data to construct and update the three-dimensional model of the environment with centimeter-level accuracy in real time, and simultaneously and accurately track the six-degree-of-freedom posture of the head-mounted AR device in the environment.
6. The AI intelligent driven rail transit AR training and operation optimization method according to claim 1 is characterized in that: The step S5: upon receiving a student's voice or gesture interaction request, providing virtual information display, intelligent question and answer, and remote expert assistance in combination with the 3D modeling and spatial positioning results, includes: S51: Capture and recognize students’ voice commands or preset gestures in real time; S52: Combining the centimeter-level three-dimensional environment model generated in step S4 and the student's current spatial positioning result, analyzing the student's interaction request intention and the environmental entity or spatial location it points to; S53: Based on the parsed intention and position, overlay and display relevant virtual information at the correct spatial position in the field of view of the head-mounted AR device; S54: Calling a preset knowledge base to provide intelligent questions and answers related to the request; S55: When the request exceeds the system's preset processing capacity, an audio and video communication channel with the remote expert is automatically established, and the student's real-time AR vision and environmental three-dimensional model are shared with the expert.
7. The AI intelligent driven rail transit AR training and operation optimization method according to claim 1 is characterized in that: S6: continuously collecting trainee training data during the training process, wherein the training data includes operation data, interaction data, and multimodal data, and evaluating the learning status in real time based on the training data, dynamically updating the personalized profile, and triggering the adjustment of the training content, including: S61: The operation data includes operation steps, time, and error points; S62: The interaction data includes query content and question and answer records; S63: Based on the training data, applying a preset learning status evaluation model to calculate evaluation indicators, wherein the evaluation indicators include the trainee's mastery of the current training content, operational proficiency, and cognitive load level; S64: Dynamically update the student personalized profile constructed in step S1 based on the evaluation indicator results; S65: Based on the updated personalized profile and the current learning status assessment results, determine whether the need to adjust the training content is triggered; S66: When an adjustment requirement is identified, step S3 is triggered to generate new customized training content.
8. An AI-driven rail transit AR training and operation optimization system, characterized by: include: Module 1: When multimodal data of students is collected through head-mounted AR devices, a personalized profile of the students is constructed based on machine learning algorithms; The second module: When a drag operation signal is received, it uses AI algorithms to match candidate models from a preset 3D model library and interactive elements from a preset interactive database, integrating and generating several AR training resources and storing them; The third module: generating customized training content based on matching the personalized portrait with corresponding AR training resources, and displaying the customized training content in AR form overlaid on the trainee's real field of view through the head-mounted AR device; Module 4: Based on the improved YOLOv8 model combined with SLAM technology, centimeter-level 3D modeling and spatial positioning of the trainees' rail transit environment; Module 5: When receiving a student's voice or gesture interaction request, the system combines the 3D modeling and spatial positioning results to provide virtual information display, intelligent question and answer, and remote expert assistance; Module 6: Continuously collect students' training data during the training process. The training data includes operation data, interaction data and multimodal data. Based on the training data, the learning status is evaluated in real time, the personalized portrait is dynamically updated and the adjustment of the training content is triggered.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the rail transit AR training and operation optimization method based on AI intelligent driving are implemented as described in claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, the steps of the rail transit AR training and operation optimization method based on AI intelligent driving are implemented as described in claims 1-7.
Citation Information
Cited By
Self-adaptive training method and system based on service data closed-loop feedback
CN121684740A