An AI and VR-based interactive training assistance system for automotive production lines
The interactive training system combining AI and VR solves the problems of high cost and high risk in traditional automotive production line training, achieving high-fidelity physical interaction and real-time intelligent teaching, thus improving training quality and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MACHINERY DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-26
AI Technical Summary
Existing automotive production line training relies on physical resources, resulting in high costs and risks. Virtual simulation systems lack high-fidelity physical interaction experiences and real-time intelligent teaching feedback.
An AI and VR-based interactive training system is adopted, which collects motion poses and voice streams through hardware interactive terminals, builds a high-fidelity production line environment by combining VR virtual simulation modules, and provides real-time teaching feedback by using AI intelligent tutor modules, including speech recognition, semantic understanding and generation, and digital human-driven systems, to realize physical interaction logic and teaching guidance.
Improve operational standardization and troubleshooting capabilities with zero safety risks, reduce training resource investment, ensure the professionalism and consistency of teaching content, shorten the skills mastery cycle, and improve training efficiency.
Smart Images

Figure CN122090677A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of virtual simulation technology, specifically to an interactive training assistance system for automotive production lines based on AI and VR. Background Technology
[0002] Currently, as the automotive manufacturing industry accelerates its evolution towards intelligent and lean manufacturing, the complexity of production line processes is increasing, placing higher demands on the skill proficiency and standardization of frontline operators. Automotive production lines encompass multiple core processes such as welding, painting, and final assembly, involving the assembly of numerous precision parts and the collaborative operation of automated equipment. The operational procedures are complex and subject to strict standard operating procedures. Therefore, establishing an efficient, safe, and standardized pre-job training system has become a crucial link in ensuring production quality and operational safety.
[0003] To address the personnel training needs of the aforementioned automotive production lines, existing technologies primarily rely on on-site practical training or multimedia teaching methods. In the practical training mode, trainees directly practice using physical parts and equipment on the production line, reinforcing their memory through repeated operation. In multimedia or basic virtual simulation applications, computer graphics technology is typically used to construct static 3D workshop models. Trainees browse equipment structures through screens or head-mounted displays, and consult electronic operation manuals and process videos, thereby establishing a basic understanding of production line layout and component construction.
[0004] However, existing training technologies still have many shortcomings. First, they consume significant physical resources and carry high risks. Physical production line training requires expensive vehicle prototypes and production schedules, resulting in high raw material costs. Furthermore, in high-voltage electricity or heavy machinery operation training, misoperation can easily lead to equipment damage or even personal injury accidents. Second, virtual interaction lacks in-depth physical feedback. Existing simulation systems often only provide visual displays and lack mechanical logic based on rigid body dynamics. Trainees cannot perceive the collision boundaries and assembly tolerances of parts in the virtual environment, leading to a disconnect between the operational feel and real assembly. In addition, teaching guidance lacks real-time and intelligent features. Traditional systems cannot accurately recognize trainees' voice requests in simulated high-noise industrial environments and lack intelligent error correction mechanisms that can dynamically link to standard operating procedures. This makes it difficult for trainees to receive targeted technical guidance after leaving human instructors, resulting in long skill mastery periods and difficulty in ensuring standardized procedures.
[0005] Therefore, the present invention provides an interactive training assistance system for automotive production lines based on AI and VR to address the shortcomings of existing technologies. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides an interactive training assistance system for automotive production lines based on AI and VR. This system solves the technical problems of high costs and risks caused by the reliance on physical resources in traditional automotive production line training, as well as the lack of high-fidelity physical interaction experience and real-time intelligent teaching feedback in existing virtual simulation systems.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solution: an interactive training assistance system for automobile production lines based on AI and VR, comprising a hardware interactive terminal, a VR virtual simulation module, an AI intelligent tutor module, and a common technical support layer.
[0008] The hardware interactive terminal is used to collect the user's motion and posture data and voice stream in physical space in real time, and present the rendered virtual training screen to the user.
[0009] The VR virtual simulation module is used to construct a virtual workshop scene containing automotive production line equipment, receive the motion pose data, process the physical interaction logic between the user and virtual parts, and transmit the virtual workshop scene containing the interaction results to the hardware interaction terminal.
[0010] The AI intelligent tutor module is used to receive the voice stream, generate teaching feedback based on the production line standard operating procedures, and generate drive commands to control the virtual digital human embedded in the virtual workshop scene, providing real-time voice explanations and action demonstrations to the user in conjunction with the physical interaction logic.
[0011] The public technology support layer is used to provide development engine support for the VR virtual simulation module and the AI intelligent tutor module, and to adapt to the hardware interaction terminal through standard interfaces to establish a data communication channel.
[0012] By adopting the above technical solution, a high-fidelity production line environment is constructed using a VR virtual simulation module, combined with real-time guidance from an AI intelligent tutor module, deeply integrating motion capture in the physical space with logical feedback in the virtual space. The system not only solves the problem of the disconnect between theory and practice in traditional training but also achieves interactive, accompanying learning through virtual digital humans. Therefore, it achieves the beneficial effect of improving trainees' operational standardization and troubleshooting capabilities with zero safety risks.
[0013] Preferably, the hardware interaction terminal includes a VR head-mounted display and an interactive controller. The VR virtual simulation module includes a virtual workshop scene unit, a physics engine unit, and an interaction control unit. The virtual workshop scene unit is used to render the 3D model data containing the automotive production line equipment using a general rendering pipeline to generate a visualized virtual workshop scene. The physics engine unit is used to calculate the collision, gravity, and rigid body dynamics effects between the virtual components, and manage the interaction relationships between objects at different levels in the virtual workshop scene through a collision detection matrix. The interaction control unit is used to convert the motion pose data input from the hardware interaction terminal into interactive events in the virtual environment, and execute raycasting interaction mode and direct grabbing mode according to the interactive events.
[0014] Preferably, when the virtual workshop scene unit performs lighting rendering and restores the realistic material texture of the automotive production line equipment, it follows the rendering equation to process the lighting rendering calculation. This process first defines the outgoing light intensity radiated from any surface point in the virtual workshop scene towards the viewing direction, as the basis for lighting calculation. The rendering equation comprehensively calculates the self-illumination intensity of the arbitrary surface point, as well as the integral of the bidirectional reflection distribution function and the incident light intensity in the hemispherical integral domain, thereby simulating realistic light and shadow interaction. The bidirectional reflection distribution function is used to define the reflection properties of the material surface, and the influence of the projected area is calculated by combining the cosine value of the angle between the incident light direction and the surface normal, ultimately determining the pixel color value of the arbitrary surface point under the current viewing angle.
[0015] Preferably, the interactive control unit utilizes the hardware interactive terminal to execute the ray interaction mode and complete non-contact interactive confirmation of a valid target. This process includes obtaining the real-time position of the interactive handle in the hardware interactive terminal as the emission origin, emitting a virtual ray along the normalized direction vector pointed to by the interactive handle, and mapping the virtual ray onto the virtual workshop scene as the user's pointing input, thereby establishing the user's long-distance pointing path in the virtual workshop scene. The geometric trajectory point set of the virtual ray is calculated based on preset ray length parameters, and it is detected in real time whether the geometric trajectory point set overlaps with the collision surface of an object in the virtual workshop scene, thereby determining the interaction intent. When a valid target is detected, a haptic feedback signal is triggered to the interactive handle, and the outline of the valid target is visually highlighted.
[0016] Preferably, the interactive control unit, in conjunction with the physics engine unit, executes the direct grasping mode to achieve precise positioning and assembly. This process includes pre-setting target mounting positions for the detachable virtual parts on the vehicle body in the virtual workshop scene, and setting adsorption area detectors at these target mounting positions to provide a spatial reference for supporting adsorption-based auxiliary assembly logic. When the user grasps the virtual part near the target mounting position using the direct grasping mode, the adsorption area detectors calculate in real time the Euclidean distance between the current position of the virtual part and the target mounting position, as well as the angle difference between the current posture quaternion and the target posture quaternion. If both the Euclidean distance and the angle difference are less than a preset tolerance threshold, a semi-transparent pre-installation phantom is visually displayed, indicating the alignment status and guiding the user to perform precise operations. When a signal is received that the user has released the grasp button, the position and posture of the virtual part are forcibly aligned instantly to the target mounting position, and the physical movement of the virtual part is locked.
[0017] By adopting the above technical solutions, and utilizing a physics engine and rendering equations that follow the laws of physics, the visual texture and physical feedback of the industrial site are restored. Through adsorption area determination and ray detection logic, the problems of precision assembly difficulties and inconvenience of long-distance operation caused by the lack of realistic force feedback in the virtual environment are solved, thereby achieving a high-precision virtual operation experience.
[0018] Preferably, the AI intelligent tutor module includes a speech recognition unit, a professional knowledge base, a semantic understanding and generation unit, and a digital human driving unit. The speech recognition unit receives the speech stream collected by the hardware interaction terminal and converts it into text data. The professional knowledge base pre-stores the production line standard operating procedures and safety specifications data used to generate the teaching feedback. The semantic understanding and generation unit uses a large language model combined with the production line standard operating procedures and safety specifications data in the professional knowledge base to generate the response text corresponding to the teaching feedback. The digital human driving unit receives the response text, generates audio, and drives the facial expressions and body movements of the virtual digital human.
[0019] Preferably, the speech recognition unit performs domain-enhanced data preparation and noise-resistant training. This process includes constructing a pre-defined domain dataset containing automotive manufacturing terminology, and mixing clean speech commands with simulated industrial background noise from the automotive production line equipment at different signal-to-noise ratios to generate an enhanced dataset. The speech recognition model is then fine-tuned using this enhanced dataset, with the optimization objective being to minimize the cross-entropy loss between the predicted and real text sequences, thereby endowing the speech recognition model with the ability to resist industrial noise interference. During the decoding stage, a cluster search strategy is employed, and additional probability weights are given to automotive production line-specific terms appearing in the predicted probability distribution to correct phoneme confusion caused by noise. Finally, high-accuracy text data is generated and passed to the semantic understanding and generation unit.
[0020] Preferably, the semantic understanding and generation unit, in conjunction with the professional knowledge base, performs retrieval enhancement generation and outputs professional technical guidance. This process includes calling a text embedding model to convert the text data output by the speech recognition unit into a query vector, retrieving the knowledge vector with the closest spatial distance to the query vector from the vector database, thereby performing a pre-retrieval for retrieval enhancement generation. The cosine of the angle between the query vector and the knowledge vector is calculated as a similarity score, and highly relevant production line standard operating procedure fragments are extracted from the professional knowledge base based on the similarity score as knowledge enhancement content. A prompt word template containing an instruction area, a context area, and a question area is constructed. The retrieved production line standard operating procedure fragments are filled into the context area of the prompt word template, and the user's question is filled into the question area of the prompt word template. The information is then input into the large language model for reasoning, ultimately generating an answer to the user's request for professional technical guidance.
[0021] Preferably, the digital human driving unit drives the virtual digital human to perform facial expression synchronization and restore the speaking state. This process includes extracting time-varying acoustic feature vectors from the audio generated by the digital human driving unit using an audio feature extractor, which serve as the input source for driving facial movements. The acoustic feature vectors are processed by a facial motion inference network to predict the weight coefficients of each hybrid deformation base shape required to achieve facial expression synchronization at the current time step. The final positions of the facial mesh vertices are calculated based on the principle of linear superposition. The sum of the products of the base vertex position coordinates of the virtual digital human model and the vertex displacement increment vectors of each hybrid deformation base shape and their corresponding weight coefficients is superimposed, thereby updating the facial mesh in real time and driving expression changes.
[0022] Preferably, the digital human driving unit drives the virtual digital human to perform pointing interaction based on inverse kinematics and achieve visual guidance. This process includes receiving and parsing structured instructions generated by the semantic understanding and generation unit, extracting the action type and focus target, and triggering state transitions in the animation state machine to prepare for the execution of pointing interaction based on inverse kinematics. When the structured instruction requires pointing to a target component in the virtual workshop scene, the position coordinates of the target component in the world coordinate system are obtained and set as the target effector position of the inverse kinematics solver. The rotation angles of each joint in the virtual digital human's arm skeletal chain are calculated inversely using inverse kinematics technology, and these rotation angles are applied to drive the skeletal model, enabling the virtual digital human's fingertips to accurately point to the target component in three-dimensional space.
[0023] By adopting the above technical solutions, through noise-enhanced training and retrieval-enhanced generation techniques, the difficulties in recognition under high-noise industrial environments and the illusion problem of large models are effectively overcome, ensuring the professionalism and accuracy of teaching content. At the same time, facial driving and inverse kinematics techniques based on acoustic features endow virtual tutors with expressiveness and spatial indication capabilities, enhancing the immersiveness and guidance efficiency of human-computer interaction.
[0024] This invention provides an interactive training assistance system for automotive production lines based on AI and VR. It has the following beneficial effects: 1. This invention constructs a high-fidelity virtual workshop scene through a VR virtual simulation module and uses a physics engine unit to calculate the dynamic effects of virtual parts, enabling users to simulate high-risk operation processes of an automobile production line in a completely virtual environment without relying on expensive physical prototype vehicles and physical sites. This avoids equipment damage and personnel injury caused by misoperation, achieving zero-risk training for high-risk operations while reducing the material investment in training resources.
[0025] 2. This invention integrates domain-enhanced speech recognition and retrieval enhancement generation technology through an AI intelligent tutor module; the system can effectively filter out background noise interference in industrial environments, accurately identify speech streams containing professional terminology, and generate standardized feedback based strictly on production line standard operating procedures in the professional knowledge base; ensuring the professionalism and consistency of teaching content, solving the teaching deviations caused by differences in teacher level in traditional training, and effectively improving the overall training quality through intelligent real-time error correction.
[0026] 3. This invention utilizes the inverse kinematics technology of the digital human driving unit to achieve precise visual guidance, and combines it with the ray interaction mode of the interactive control unit and the assembly logic assisted by the adsorption area; the system transforms abstract theoretical specifications into intuitive action demonstrations of virtual digital humans and precise operational experiences with physical feedback, supporting trainees to conduct personalized and repeated practice for specific processes; it enhances trainees' memory of spatial structures and process flows, thereby shortening the skill mastery cycle and improving training efficiency. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the system architecture according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the interaction process and training method of the system in an embodiment of the present invention; Figure 3 This is a schematic diagram of the comprehensive response simulation of the core algorithm of the system in this embodiment of the invention.
[0028] Among them, 10. Hardware interaction terminal; 11. VR head-mounted display; 12. Interactive handle; 20. VR virtual simulation module; 21. Virtual workshop scene unit; 22. Interactive control unit; 23. Physics engine unit; 30. AI intelligent tutor module; 31. Speech recognition unit; 32. Semantic understanding and generation unit; 33. Digital human driving unit; 34. Professional knowledge base; 40. Public technology support layer. Detailed Implementation
[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] See attached document Figure 1 This invention provides an interactive training assistance system for automotive production lines based on AI and VR. The system mainly consists of a hardware interactive terminal 10, a VR virtual simulation module 20, an AI intelligent tutor module 30, and a common technical support layer 40. The hardware interactive terminal 10 serves as the system's physical input and output interface, configured to capture user actions and voice information and present virtual images. The VR virtual simulation module 20 is configured to construct a three-dimensional virtual space and process physical interaction logic. The AI intelligent tutor module 30 is configured to process natural language, retrieve knowledge from the knowledge base, and generate digital human feedback. The common technical support layer 40 serves as the underlying infrastructure, providing development engine support and version management functions for the aforementioned modules.
[0031] The hardware interaction terminal 10 uses a virtual reality all-in-one device that supports the OpenXR standard. In this embodiment, the hardware interaction terminal 10 specifically uses the Pico 4 Enterprise VR all-in-one device. The hardware interaction terminal 10 includes a VR head-mounted display 11 and an interactive controller 12. The VR head-mounted display 11 has six degrees of freedom tracking capabilities, configured to locate the user's position and head posture in physical space in real time, and transmit the rendered virtual image to the user's eyes. The interactive controller 12 also supports six degrees of freedom tracking, configured to capture the user's hand movements, including position movement, rotation, and button trigger signals. The hardware interaction terminal 10 communicates with the VR virtual simulation module 20 wirelessly or via wired means, transmitting the user's pose data and operation commands to the system, and receiving visual and auditory signals from the system.
[0032] The VR virtual simulation module 20 is built on the Unity 2022 LTS engine provided by the public technology support layer 40. It is developed in C# and follows the OpenXR standard to ensure cross-hardware platform compatibility. The VR virtual simulation module 20 internally includes a virtual workshop scene unit 21, an interaction control unit 22, and a physics engine unit 23. The virtual workshop scene unit 21 uses the Universal Render Pipeline (URP) for rendering, ensuring high frame rates while presenting industrial-grade lighting and shadow effects. It contains 3D model data for welding, painting, and final assembly production lines. The physics engine unit 23 is built on the Unity PhysX physics engine and is configured to calculate collision, gravity, and rigid body dynamics effects between virtual components. The interaction control unit 22 integrates the XR Interaction Toolkit and is configured to implement interactive logic such as raycasting, direct grabbing, and UI touch control. The logical flow of the VR virtual simulation module 20 includes two parallel functional paths: a workshop production line introduction path and a vehicle structure recognition path. Users input commands through the hardware interaction terminal 10 to select the corresponding path for learning.
[0033] The AI-powered intelligent tutor module 30 operates independently of or integrated into the VR virtual simulation module 20, configured to provide intelligent teaching guidance. The AI-powered intelligent tutor module 30 integrates a speech recognition unit 31, a semantic understanding and generation unit 32, a digital human driving unit 33, and a professional knowledge base 34. The speech recognition unit 31 is configured to receive user speech streams collected by the hardware interaction terminal 10, use the Whisper Tiny model to denoise the speech in industrial noise environments, and convert it into text data. The semantic understanding and generation unit 32 employs a retrieval-enhanced RAG generation mechanism, combining the ChatGLM-6B large language model and the professional knowledge base 34 to generate accurate teaching response text. The professional knowledge base 34 pre-stores SOP standard operating procedures and safety specification data. The digital human driving unit 33 is configured to receive the response text, generate audio using text-to-speech technology, and drive the facial expressions and lip movements of the virtual digital human using Audio2Face technology, ultimately presenting a virtual tutor image with voice explanation and action demonstration functions in the VR virtual simulation module 20. The output signals of the AI intelligent tutor module 30 are directed to the workshop production line introduction path and the vehicle structure recognition path, respectively, to achieve real-time guidance in different training stages.
[0034] The common technical support layer 40 vertically integrates the VR virtual simulation module 20 and the AI intelligent tutor module 30, providing a unified technical foundation for the entire system. The common technical support layer 40 specifies that the project is developed based on the Unity 2022 LTS Long Term Support engine, ensuring the system's stability and maintainability. Simultaneously, the common technical support layer 40 uses the Git version control system to manage project code, 3D assets, and configuration files, supporting multi-user collaborative development and historical version rollback. This layer also handles OpenXR interface adaptation, decoupling upper-layer business logic from lower-layer hardware devices. When the hardware interaction terminal 10 is replaced with other devices supporting the OpenXR standard, there is no need to refactor the core code logic.
[0035] See attached document Figure 1 Within the internal logic of the VR virtual simulation module 20, the virtual workshop scene unit 21 serves as the fundamental visual carrier for constructing an immersive experience. The virtual workshop scene unit 21 is developed based on the Unity 2022 LTS Long Term Support engine, employing a left-handed coordinate system. It sets the ratio of virtual units to real-world physical units at a 1:1 ratio, meaning one unit in Unity corresponds to one meter in the real world. This setting ensures that the spatial distances, component sizes, and equipment heights in the virtual environment remain strictly consistent with those of a real automotive production line, providing an accurate spatial reference for subsequent physical interactions.
[0036] Virtual workshop scene unit 21 uses the Universal Render Pipeline (URP) as its core rendering architecture. The URP customizes the rendering process through scripted rendering pipeline technology, offering higher execution efficiency on mobile VR devices compared to traditional built-in rendering pipelines. Virtual workshop scene unit 21 is equipped with a physically based rendering material system. This system utilizes a bidirectional reflection distribution function to simulate the physical behavior of light on object surfaces, thereby reproducing the realistic optical textures of metal, glass, rubber, and car paint in the virtual environment.
[0037] To achieve a highly realistic visual effect, the virtual workshop scene unit 21 follows a rendering equation when processing lighting calculations. This equation describes the luminance radiated from any point in the scene towards the viewer. The specific calculation formula is as follows: ; in, Indicates position Along the direction The brightness of the emitted light; Indicates position The self-illuminating brightness at that location; Indicated by normal line The central unit hemisphere region; This represents the bidirectional reflection distribution function, used to define the reflection properties of a material surface; Indicates along direction The intensity of incident light; The cosine of the angle between the incident light direction and the surface normal is used to calculate the impact on the projected area. Through real-time or pre-calculation using this formula, the virtual workshop scene unit 21 can accurately present the lighting effects of complex light sources on the equipment surface within the workshop.
[0038] The construction process of the virtual workshop scene unit 21 begins with the collection of real production line data. Developers import industrial CAD data from the automotive production line, including geometric data of welding robots, painting spray guns, assembly conveyor chains, and vehicle body structures. Because the original industrial CAD data contains a high density of vertices and faces, it cannot run smoothly directly on the hardware interactive terminal 10. Therefore, the virtual workshop scene unit 21 introduces a lightweight model processing workflow. This workflow uses topology reconstruction technology to convert high-poly meshes into low-poly meshes while preserving the model's external contour features, and generates normal maps to simulate the detailed bumps and depressions of the high-poly model on the low-poly surface.
[0039] The virtual workshop scene unit 21 is internally divided into three sub-scenes: welding production line, painting production line, and final assembly production line. The welding production line sub-scene focuses on simulating the motion trajectory of industrial robots and the effect of welding sparks. It uses a particle system to generate sparks during welding and binds kinematic parent-child hierarchical relationships to the robot joints. The painting production line sub-scene focuses on simulating a cleanroom environment and the car paint spraying process. It uses reflective probe technology to capture the surrounding environment in real time and generate dynamic reflections on the car body surface, simulating the visual difference between wet and dry film. The final assembly production line sub-scene focuses on simulating assembly line operations and parts assembly. It constructs a complete assembly line including chassis assembly stations and interior installation stations, and performs independent mesh separation and collision body settings for each interactive component.
[0040] To maintain a stable frame rate under the computing power limitations of the hardware interactive terminal 10, the virtual workshop scene unit 21 employs multi-level detail (LOD) technology and occlusion culling technology. LOD technology dynamically switches the precision level of object models based on the distance between the camera and the object. When the user is far from a device, the system renders a low-precision model; when the user moves closer, the system seamlessly switches to a high-precision model. Occlusion culling technology detects the occlusion relationships of objects within the view frustum in real time, stopping rendering calculations for background objects completely occluded by foreground objects, thereby reducing the load on the graphics processor. Furthermore, the virtual workshop scene unit 21 uses a hybrid lighting strategy. For static walls, floors, and large equipment, lightmap baking technology is used to pre-calculate shadows and indirect lighting and store them as textures; for dynamic robotic arms and parts, real-time lighting calculations are used to balance image quality and performance.
[0041] See attached document Figure 1 In the architecture of the VR virtual simulation module 20, the interaction control unit 22 and the physics engine unit 23 work together to provide users with an operating experience that conforms to the laws of physics. The interaction control unit 22 is built based on the XR InteractionToolkit development toolkit and is responsible for converting the sensor data input from the hardware interaction terminal 10 into interactive events in the virtual environment. The physics engine unit 23 is built based on the Unity PhysX physics simulation engine and is responsible for handling the mechanical interactions between virtual objects.
[0042] The physics engine unit 23 parameterizes the physical properties of all 3D models within the virtual workshop. For virtual components, the physics engine unit 23 adds rigid body components and collider components. The rigid body component defines the component's dynamic properties, such as mass, drag, angular drag, and whether it is affected by gravity. The collider component defines the physical boundary shape of the component. For simple components (such as bolts and washers), box-shaped or spherical colliders are used to reduce computational overhead; for complex components (such as engine blocks and door panels), mesh colliders are used to accurately fit their geometric shapes. The physics engine unit 23 manages the interaction relationships between objects at different levels through a collision detection matrix, dividing virtual hands, interactive components, static equipment, and environmental walls into different physical layers, and pre-setting collision trigger logic between each layer, thereby avoiding unnecessary physical calculations and preventing clipping phenomena.
[0043] The interaction control unit 22 implements two core interaction modes: ray interaction mode and direct grabbing mode. Ray interaction mode is mainly used for selecting distant objects and operating the UI interface. In this mode, the system emits a virtual ray along the direction the handle is pointing, with the real-time position of the interaction handle 12 as the origin. The interaction control unit 22 calculates the intersection of this ray with colliders in the scene in real time. The geometric trajectory of the ray is defined to satisfy the parametric equations: ; in, The length parameter on the ray is Spatial coordinates of time; This indicates the origin of the ray's emission, corresponding to the center coordinates of the interactive handle 12 in the virtual space; This represents the normalized direction vector of the ray, corresponding to the current direction of the interaction handle 12; The length parameter of the ray is zero, and its value ranges from zero to the maximum detection distance set. The interactive control unit 22 determines whether a target is selected by checking whether the point set calculated by the detection parameter equation overlaps with the surface of the collider of an object in the scene. When the ray detects a valid target, the system triggers a haptic feedback signal to the interactive handle 12 and visually highlights the outline of the target object.
[0044] The direct gripping mode is primarily used for precision operation and assembly training of close-range parts. When the virtual model of the interactive handle 12 overlaps with the collision body of the part, and the user presses the grip button on the handle, the interactive control unit 22 establishes a parent-child hierarchical relationship or physical joint connection between the virtual hand and the part. To simulate realistic weight and inertia, the physics engine unit 23 applies a velocity-based physics-driven mechanism during the gripping process. The system calculates the position and attitude differences of the interactive handle 12 between the current frame and the previous frame, derives the instantaneous velocity and angular velocity, and applies these velocity vectors to the gripped rigid body part, so that the part follows the hand movement while retaining the physical collision effect, preventing the part from penetrating other virtual objects.
[0045] For the two core training functions of vehicle structure recognition and component cognition, the VR virtual simulation module 20 has designed specific interactive logic. In the component cognition function, the system implements interactive observation logic. When the user selects a component using a ray and triggers the cognition command, the system clones a copy of the component to the observation point in front of the user's field of vision and temporarily suspends the gravity attribute of the copy. The user can control the free rotation of the component copy in the X, Y, and Z axes by operating the joystick of the interactive handle 12 or by rotating the wrist. At the same time, the system supports dual-axis zoom gestures; the user can zoom in on the component details by stretching their hands outward and zoom out on the panoramic view by clasping their hands together. During this process, the interactive control unit 22 reads the displacement increment of the handle in real time and maps it to the scaling factor of the component model.
[0046] In the vehicle structure recognition function, the system implements a disassembly and assembly logic based on the adsorption area. The system pre-defines a target mounting position for each detachable component on the vehicle body and sets up an adsorption area detector at that position. During disassembly, the user directly grasps the component and moves it out of the adsorption area; the system automatically decouples the component from the vehicle body. During assembly, when the user grasps the component and approaches the correct target mounting position, the adsorption area detector calculates the Euclidean distance between the component's current position and the target position, as well as the angle difference between the current quaternion and the target quaternion. If both the distance and angle difference are less than a preset tolerance threshold, the system will visually present a semi-transparent pre-installed phantom. If the user releases the grip button at this time, the physics engine unit 23 will forcibly align the component's position and orientation to the target mounting position and lock its physical movement, completing the virtual assembly operation. This mechanism ensures that trainees can accurately understand the relative positional relationships and assembly sequence of each component within the vehicle when performing complex structural recognition exercises.
[0047] See attached document Figure 1In the architecture of the AI intelligent tutor module 30, the speech recognition unit 31 serves as the entry point for human-computer interaction, configured to process the real-time audio stream collected by the hardware interaction terminal 10. The core processing engine of the speech recognition unit 31 adopts the Whisper tiny model. This model is a weakly supervised learning speech recognition model based on the Transformer architecture, characterized by small parameter count and fast inference speed, and can adapt to the local computing power of VR all-in-one devices or the low-latency deployment requirements of the cloud. The speech recognition unit 31 first establishes an audio data channel with the hardware interaction terminal 10, receiving monophonic pulse code modulation data with a sampling rate of 16000 Hz. To ensure the continuity and integrity of the data, the speech recognition unit 31 maintains a circular buffer internally to temporarily store the input audio frames and slices the audio data according to a preset time window (e.g., 30 seconds).
[0048] The speech recognition unit 31 performs feature extraction on the received raw audio data, converting it into a log-Mel spectrogram. Specifically, the system first performs a short-time Fourier transform on the audio signal, using a Hanning window function to reduce spectral leakage. Then, the system maps the spectrum onto a Mel scale, which simulates the nonlinear perception characteristics of the human ear at different frequencies. Finally, the system takes the logarithm of the energy values of the Mel spectrum to generate a Log-Mel spectrogram. This spectrogram is input as tensor data to the encoder part of the Whisper tiny model. The encoder consists of two one-dimensional convolutional layers and a subsequent sinusoidal positional coding layer, configured to extract the time-frequency features of the audio and preserve the positional information of the sequence.
[0049] To address the high-intensity industrial noise present in automotive production line environments (such as spot welding machine noise, conveyor belt friction noise, and pneumatic tool exhaust noise), the speech recognition unit 31 did not simply rely on traditional signal processing noise reduction algorithms. Instead, it adopted a domain-adaptive fine-tuning strategy to improve the model's robustness in noisy environments. The developers constructed a domain-specific dataset containing automotive manufacturing terminology. This dataset was generated using data augmentation techniques, specifically by mixing clean speech command audio with industrial background noise audio collected from real workshops at different signal-to-noise ratios. The mixing ratios covered a range from -5 dB to +20 dB to simulate noise interference levels at different workstations.
[0050] The speech recognition unit 31 uses the aforementioned augmented dataset to perform full-parameter fine-tuning or parameter-efficient fine-tuning (such as LoRA) on the Whisper Tiny model. During fine-tuning, the optimization objective of the model is to minimize the cross-entropy loss between the predicted text sequence and the real text sequence. The formula for this loss function is as follows: ; in, This represents the loss function value of the model; This represents the set of weight parameters for the Whisper tiny model; Indicates the total length of the target text sequence; Indicates at time step The actual target word at the location; Indicates at time step All previously generated historical word sequences; This represents the Log-Mel spectrogram features of the input, which include industrial noise. This represents the conditional probability that the current word is the true target word, given the input spectrogram features and the historical word sequence. The weight parameter set is continuously updated through the backpropagation algorithm. This enables the model to learn the ability to separate and identify valid speech signals from noisy speech features, thereby suppressing the interference of background industrial noise on the recognition results.
[0051] After fine-tuning, the speech recognition unit 31 employs a cluster search strategy during the decoding stage. The decoder receives the output features from the encoder and generates text tags autoregressively. During decoding, the system is configured with a bias mechanism for automotive production line-specific vocabulary (such as torque wrench, body-in-white, and front longitudinal beam). When candidate tags for these specific terms appear in the probability distribution predicted by the model, they are given additional probability weights, thereby improving the accuracy of terminology recognition. Finally, the speech recognition unit 31 standardizes the decoded text string (e.g., removes punctuation and unifies capitalization) and transmits it as an output signal to the downstream semantic understanding and generation unit 32. This process enables high-precision capture and transcription of trainee speech commands in noisy virtual or mixed reality industrial environments.
[0052] See attached document Figure 1In the architecture of the AI intelligent tutor module 30, the semantic understanding and generation unit 32 and the professional knowledge base 34 are connected via a high-speed data bus, jointly forming the cognitive and reasoning center of the system. The semantic understanding and generation unit 32 is designed based on a retrieval-enhanced generation technology architecture, aiming to solve the problems of knowledge illusion and timeliness lag in traditional generative large models in vertical industrial fields. The professional knowledge base 34, as a structured data storage container, pre-stores a large amount of unstructured and semi-structured text data. This data covers standard operating procedure documents, production safety manuals, equipment maintenance technical guidelines, historical fault case libraries, and workshop emergency response plans in the automotive manufacturing field. During system initialization, the semantic understanding and generation unit 32 is equipped with a data preprocessing pipeline responsible for format cleaning, noise reduction, and block processing of the aforementioned raw documents. Block processing uses a sliding window algorithm to divide long texts into fixed-length semantic segments, each containing 512 or 1024 characters, while preserving a certain amount of overlap between adjacent segments to maintain the continuity of contextual semantics.
[0053] The semantic understanding and generation unit 32 integrates a text embedding model, which is responsible for mapping the segmented text fragments to a high-dimensional vector space. The embedding model encodes each text fragment, generating a dense floating-point vector with 768 or 1024 dimensions. These vectors are stored as knowledge indexes in a vector database that supports approximate nearest neighbor search. When the speech recognition unit 31 outputs a user's natural language text query, the semantic understanding and generation unit 32 first calls the embedding model to convert the query text into a query vector of the same dimension. Subsequently, the system performs a vector retrieval operation, searching the vector database for several knowledge vectors that are spatially closest to the query vector.
[0054] To quantify the semantic relevance between the query vector and the knowledge vectors stored in the knowledge base, the semantic understanding and generation unit 32 uses cosine similarity as an evaluation metric. The system calculates the cosine value of the angle between the query vector and each candidate knowledge vector, using the following formula: ; in, This represents the similarity score between the query vector and a specific knowledge vector. The value usually ranges from negative one to positive one, and the closer the value is to positive one, the more semantically related the meaning is. This represents the vector corresponding to the query text of the current user; This represents the knowledge vector corresponding to a text fragment stored in the professional knowledge base 34; Represents the total number of dimensions of the vector; and Representing vectors respectively and In the The system ranks all candidate knowledge fragments based on the calculated similarity scores and selects the top three to five text fragments with the highest similarity as the external knowledge context.
[0055] The core generation engine of the semantic understanding and generation unit 32 employs the ChatGLM-6B large language model. This model, based on a general language model architecture, has 6 billion parameters and supports both Chinese and English bilingual processing. The system constructs an input sequence containing prompt word templates, which explicitly delineate the instruction area, context area, and question area. The semantic understanding and generation unit 32 fills the context area with highly relevant text fragments and the question area with the user's original question, then inputs the combined complete prompt word sequence into the ChatGLM-6B model. The model utilizes its pre-trained attention mechanism to perform joint reasoning by combining the input's professional contextual information with the user's question.
[0056] When generating answers, the ChatGLM-6B model employs an autoregressive decoding strategy, predicting each word in the output sequence one by one. Since the input includes precise SOP terms or safety specifications retrieved from the professional knowledge base 34, the model is forced to pay attention to this external evidence during the generation process. This ensures that the generated answers strictly adhere to the technical standards and safety regulations of the automotive production line, preventing the model from generating seemingly plausible but erroneous information based solely on internal memory. The final generated text answer, after being verified by a compliance filter, is passed to the digital human driving unit 33 to drive the virtual tutor's speech synthesis and facial expression animation, achieving a closed-loop intelligent question-answering system based on an authoritative knowledge base.
[0057] See attached document Figure 1 The digital human driving unit 33, serving as the visual output terminal of the AI intelligent tutor module 30, is responsible for generating and driving a highly realistic virtual tutor avatar in real time within the three-dimensional space constructed by the VR virtual simulation module 20. The digital human driving unit 33 pre-stores digital human model data reconstructed from high-precision 3D scans. This model includes skeletal binding information, skin texture mapping, and hybrid deformation data for facial expression control. During system operation, the digital human driving unit 33 receives output signals from the semantic understanding and generation unit 32. These signals contain two data streams: a synthesized audio stream for speech playback and a semantic command stream for motion control.
[0058] For driving facial expressions, the digital human driving unit 33 employs audio-driven facial animation technology based on deep neural networks, namely Audio2Face. The digital human driving unit 33 integrates an audio feature extractor and a facial motion inference network. When receiving a pulse-code modulation stream of synthesized audio data, the audio feature extractor extracts audio frames using a sliding window approach and extracts Mel-frequency cepstral coefficients as acoustic feature vectors. The facial motion inference network receives this acoustic feature vector as input, processes it through a multi-layer fully connected network or recurrent neural network, and predicts the facial muscle motion parameters corresponding to the current time step. To achieve subtle lip-sync and facial expression changes, the digital human driving unit 33 employs a blend shape technique, namely BlendShape. This technique predefines a set of basic facial shape variants, each corresponding to a specific facial muscle movement, such as raising the corners of the mouth, opening the jaw, and furrowing the eyebrows. The output of the facial motion inference network is directly mapped to the linear weight coefficients of these basic shape variants.
[0059] In each frame rendering loop, the digital human driving unit 33 calculates the final positions of the facial mesh vertices based on the predicted weight coefficients. The vertex position calculation follows the principle of linear superposition, and its formula is as follows: ; in, This represents the final position coordinates of the facial mesh vertices in the local coordinate system after calculation; This represents the coordinates of the base vertex positions of the digital human model in a silent, neutral state. This indicates the total number of predefined hybrid deformation base shapes, typically set to 52 to meet ARKit standards or higher precision standards; Indicates the first The weight coefficients corresponding to the basic shape of the hybrid deformation are output in real time by the facial motion inference network, and the value range is limited to between zero and one. Indicates the first The formula represents the vertex displacement increment vector of the hybrid deformable base shape relative to the base state. Through this formula, the digital human driving unit 33 can ensure that the virtual tutor's lip movements and speech pronunciation are synchronized at the millisecond level, while simultaneously exhibiting natural blinking and micro-expressions.
[0060] For driving body movements, the digital human driving unit 33 employs an intent-driven animation state machine mechanism. While generating the response text, the semantic understanding and generation unit 32 utilizes natural language processing technology to classify the text content by intent and recognize entities, generating structured instructions that include the action type and the target of interest. For example, when the response involves engine disassembly steps, the system generates an action instruction pointing to the engine. The digital human driving unit 33 receives this instruction and triggers a state transition within its internal animation state machine. The state machine smoothly transitions from an idle state to an explanation state, demonstration state, or warning state based on the instruction type.
[0061] To achieve precise spatial interaction, the digital human driving unit 33 incorporates inverse kinematics technology. When an instruction requires the virtual instructor to point to a specific component in the virtual workshop scene unit 21, the digital human driving unit 33 acquires the component's position coordinates in the world coordinate system and sets these coordinates as the target effector position for the inverse kinematics solver. The solver uses an iterative algorithm to calculate the rotation angles of the shoulder, elbow, and wrist joints in the virtual instructor's arm skeletal chain, ensuring that the fingertips accurately point to the target component while maintaining the biomechanical rationality of the body posture. This multimodal interactive driving method allows the virtual instructor to not only provide explanations through voice but also to use body language and facial expressions to offer immersive and present-like teaching guidance in the virtual space.
[0062] See attached document Figure 2 This invention provides an interactive process and training method for an AI and VR-based interactive training assistance system for automotive production lines. This method, through the collaborative work of a hardware interactive terminal 10, a VR virtual simulation module 20, and an AI intelligent tutor module 30, provides trainees with a closed-loop training process from immersive access to intelligent feedback.
[0063] In step S1, the system performs immersive access and scene initialization. The trainee wears the VR headset 11 and holds the interactive controller 12, activating the hardware interaction terminal 10. The system first performs a hardware self-test and connection handshake to ensure smooth data transmission between the VR headset 11 and the computing unit. The VR virtual simulation module 20 then starts, loading the virtual workshop scene unit 21 built based on the Unity engine. During this process, the system reads preset automotive production line scene data, including the 3D model assets and lighting maps of the welding workshop, painting workshop, and final assembly workshop. Simultaneously, the hardware interaction terminal 10 activates the six-degree-of-freedom tracking function, establishing a mapping relationship between the virtual space coordinate system and the physical space coordinate system, mapping the trainee's head movements and hand gestures to the virtual avatar in real time. The AI intelligent tutor module 30 initializes synchronously, loading a virtual digital human model at a preset position in the virtual scene and establishing a voice listening channel, putting the system in standby mode.
[0064] In step S2, the system performs a dual-path learning mode selection operation. After initialization, the VR virtual simulation module 20 renders a holographic interactive menu in front of the learner's field of vision. The learner uses the interactive handle 12 to emit a virtual ray to point to the menu option and presses the confirmation button to select. The system provides two parallel learning paths: workshop production line introduction and vehicle structure recognition. If the learner selects the workshop production line introduction mode, the VR virtual simulation module 20 calls the roaming logic and plans a virtual tour route that runs through each production workshop according to the production process sequence. The virtual digital human will follow this route to provide fixed-point explanations. If the learner selects the vehicle structure recognition mode, the system switches to an independent display workstation scene, loads a high-precision 3D model of the whole vehicle or assembly parts in the center of the scene, and activates the grabbing and disassembly function permissions of the interactive control unit 22 to prepare for subsequent refined operations.
[0065] In step S3, the system performs a virtual-real fusion interactive operation. This step is the core of the training, covering component recognition operations and high-risk scenario simulation. During component recognition, trainees use the interactive handle 12 to approach virtual components, and the interactive control unit 22, in conjunction with the physics engine unit 23, performs real-time collision detection. When the virtual hand model contacts the component's collision object and triggers a grasping signal, the system binds the component's model coordinates to the handle coordinates, allowing trainees to drag, rotate, or disassemble the component using six degrees of freedom. The interactive control unit 22 records the trainee's operation trajectory and the component's pose changes in real time. For high-risk scenario simulation, such as in the welding robot's working area, the system sets up an invisible electronic fence trigger. When a trainee enters the danger zone without disconnecting the power or wearing virtual protective equipment, the physics engine unit 23 detects a positional conflict and immediately triggers the virtual accident logic. The VR virtual simulation module 20 simulates sparks or equipment impact effects through visual effects and controls the interactive handle 12 to generate strong tactile vibration feedback, thereby enhancing the trainee's safety awareness with zero risk of personal injury.
[0066] In step S4, the system executes a closed-loop operation of AI real-time guidance and feedback. Throughout the interaction, the AI intelligent tutor module 30 runs continuously. When a student encounters difficulties during operation and asks a question via voice, the voice recognition unit 31 collects audio data and converts it into a text stream. The semantic understanding and generation unit 32 receives the text stream and, combined with the current context of the student's situation, retrieves relevant standard operating procedures or troubleshooting documents from the professional knowledge base 34. The system generates accurate guidance text using a large language model and transmits it to the digital human driving unit 33. The digital human driving unit 33 drives the virtual tutor to face the student, providing answers through synthesized voice and corresponding gesture guidance, such as pointing to the button to be operated or showing the correct disassembly / assembly angle. Furthermore, the system has an active error correction function. When the interaction control unit 22 detects that the student's operation steps do not conform to the standard procedures in the professional knowledge base 34, such as an incorrect bolt tightening sequence, the AI intelligent tutor module 30 immediately interrupts the current process, issues a voice warning through the virtual tutor, and highlights the correct operation object in the student's field of vision until the student completes the correct operation, thus forming a real-time teaching feedback closed loop.
[0067] See attached document Figure 1 and attached Figure 2 Although the specific embodiments of the present invention have been described in detail above, it should be understood that these descriptions are only for explaining the principles and applications of the present invention and should not be construed as limiting the scope of protection of the present invention. Those skilled in the art can make various modifications, substitutions, or variations to the hardware devices, software algorithms, development platforms, and application scenarios in the system without departing from the spirit and scope of the present invention.
[0068] Regarding the selection of the hardware interaction terminal 10, although the Pico 4 Enterprise virtual reality all-in-one machine is specifically used in this embodiment, the technical solution of the present invention is not limited to a specific brand or model of device. The hardware interaction terminal 10 can be any head-mounted display device that supports the OpenXR general standard and has six degrees of freedom spatial positioning and tracking capabilities. For example, the hardware interaction terminal 10 can be replaced with the Meta Quest series, HTC VIVE Focus series, or other virtual reality head-mounted displays with equivalent rendering computing capabilities and sensor accuracy. As long as the device can run the VR virtual simulation module 20 and communicate with the AI intelligent tutor module 30, and can provide real-time feedback on the user's head posture and position information through its built-in or external sensors. In addition, the computing unit of the hardware interaction terminal 10 can be an all-in-one processor integrated in the head-mounted display, or a high-performance external computer workstation connected via streaming cable or wireless network.
[0069] Regarding the development environment of the VR virtual simulation module 20 and the common technology support layer 40, this embodiment of the invention describes an implementation method based on the Unity 2022 LTS engine and URP rendering pipeline, but this is not the only implementation path. The VR virtual simulation module 20 can also be built based on Unreal Engine, CryEngine, or other real-time 3D engines that support extended reality development. For example, if Unreal Engine is used as the development platform, the virtual workshop scene unit 21 can use its Nanite virtual geometry technology to handle industrial models with high polygon counts, and the physics engine unit 23 can be replaced with the Chaos physics engine to achieve the same physical collision detection and dynamic simulation effects. The key is that the selected engine must be able to support the XRInteraction Toolkit or a similar interactive development framework to ensure that the functional logic of the interactive control unit 22 is realized.
[0070] For the algorithm models within the AI intelligent tutor module 30, feasible alternative technical solutions exist for each sub-unit. While the speech recognition unit 31 preferentially uses the Whisper Tiny model to adapt to edge computing resources, it can also be replaced with other lightweight deep neural network speech recognition models, or, where network conditions permit, access high-precision speech recognition API services in the cloud. The ChatGLM-6B model integrated in the semantic understanding and generation unit 32 can also be replaced with the BERT model for intent classification tasks, or with other open-source or self-developed generative pre-trained transformation models such as LLaMA and Baichuan for text generation tasks, depending on actual needs. The core criterion for replacement is that the model must have the ability to fine-tune for specific industrial verticals and compatibility with RAG retrieval enhancement generation mechanisms.
[0071] Regarding the interaction methods of the interactive control unit 22, in addition to the physical button and ray interaction based on the interactive handle 12 described in the embodiment, the system can be expanded or replaced with more natural interaction forms. The system can utilize the external camera sensor of the hardware interactive terminal 10 to integrate computer vision gesture recognition algorithms, thereby realizing hand skeleton tracking interaction without a handle, allowing users to directly grasp parts in virtual space with both hands. In addition, the interactive control unit 22 can also combine eye-tracking technology to assist in object selection or attention analysis by capturing the user's gaze point, thereby triggering the specific feedback mechanism of the AI intelligent tutor module 30.
[0072] While this embodiment primarily focuses on automotive production line training, the system architecture of this invention boasts broad applicability and scalability. By replacing the 3D assets in the virtual workshop scene unit 21 and updating the SOP documents in the professional knowledge base 34 to corresponding domain operating standards, this system can be seamlessly migrated to training scenarios in other industrial manufacturing fields such as aerospace assembly, shipbuilding welding processes, and precision electronic instrument assembly. This cross-domain application migration does not alter the core technological logic of the deep integration of VR and AI, and therefore also falls within the protection scope of this invention.
[0073] See attached document Figure 3 This embodiment constructs a virtual training scenario for the precision assembly process of engine pistons. In this scenario, the system needs to simultaneously handle physically based lighting rendering, semantic-based knowledge retrieval, and weighted digital human-driven operations. To verify the logical consistency of the technical solution and the accuracy of the mathematical model in this invention, this embodiment performs normalized numerical simulations on the algorithms for the above three key dimensions. (Appendix) Figure 3 The system's response characteristics under different input parameters are displayed in a unified two-dimensional coordinate system, where the horizontal axis represents the normalized input control variables, ranging from zero to one; and the vertical axis represents the normalized system response output value, also ranging from zero to one.
[0074] Appendix Figure 3 The solid line in the graph corresponds to the verification results (digital human driving response) of the hybrid deformation algorithm of the digital human driving unit 33. In this simulation test, the input control variable on the horizontal axis represents the hybrid deformation weight coefficient output by the facial motion inference network, and the vertical axis represents the displacement response of the facial mesh vertices. The formula for calculating the digital human vertex position involved in this invention is as follows: ; Appendix Figure 3 The solid line in the diagram represents a straight line with a constant slope, indicating that as the weighting coefficient increases... As the displacement increment increases from zero to one, the displacement increment of the vertex changes uniformly. This result verifies that the digital human driving unit 33 can accurately perform linear superposition operations, ensuring that when the virtual tutor explains the precautions for piston installation, the size of its mouth opening and closing and the intensity of its voice are strictly linearly synchronized, avoiding stiff facial expressions or clipping phenomena caused by nonlinear distortion.
[0075] Appendix Figure 3 The long dashed line in the graph corresponds to the lighting rendering algorithm verification result (lighting rendering response) of virtual workshop scene unit 21. In this test, the horizontal axis represents the cosine of the angle between the incident ray direction and the normal direction of the object surface, and the vertical axis represents the physically based reflected light intensity. The core terms of the rendering equation involved in this invention are as follows: ; In simulating the material of a real metal piston, the system introduces a specular reflection component, resulting in a non-linear specular response to light. (See attached image) Figure 3 The long dashed line in the image shows an exponentially steep upward trend in the region where the input value is close to one. This verifies that the system can correctly handle the integral calculation and material property mapping in the rendering equation. This allows students to see the metallic highlights that flash violently with small changes in angle when rotating and observing the piston, thereby improving the texture and realism of the virtual parts.
[0076] Appendix Figure 3 The short dashed lines in the graph correspond to the verification results (AI semantic retrieval response) of the retrieval enhancement algorithm of the semantic understanding and generation unit 32. In this verification scenario, the horizontal axis represents the semantic proximity between the query vector and the knowledge base vector, and the vertical axis represents the similarity score calculated by the system. The cosine similarity calculation formula involved in this invention is as follows: ; Appendix Figure 3 The short dashed line in the graph illustrates a highly sensitive response curve, particularly in the high similarity region where the horizontal axis value is greater than 0.8, where the response value rapidly approaches one. This verifies the effectiveness of the cosine similarity formula in handling high-dimensional vectors, demonstrating that the system can keenly capture key technical terms in user voice questions and accurately match them with standard operating procedure clauses in the knowledge base 34. This non-linear, highly discriminative response characteristic ensures that when learners ask specific questions such as the orientation of piston ring notches, the system can quickly filter out low-relevance documents, providing only the correct installation specifications with high confidence to the generative model, thus guaranteeing the authority and accuracy of the teaching content.
[0077] In summary, Appendix Figure 3 The simulation results further reveal the collaborative optimization mechanism of this system when processing multimodal complex data, and its specific applications are reflected in the following three aspects: First, the exponential specular highlights in lighting rendering response are not only for enhancing visual aesthetics, but also for providing crucial industrial-grade visual cues. In precision processes such as engine piston assembly, the fluidity of the metallic luster on component surfaces as it changes with angle is a vital indicator for operators to judge workpiece posture and surface flatness. By reproducing this non-linear physical optical property, the system allows trainees to fine-tune the piston's cylinder angle using reflected highlights, just as in a real environment. This enables them to train realistic hand-eye coordination in virtual space, effectively reducing the skill loss rate when transferring from virtual training to physical operation.
[0078] Secondly, the strictly linear nature of the digital human-driven response is designed to ensure psychological comfort and focus during human-computer interaction. In extended VR training sessions, non-linear delays or exaggerated distortions in the virtual instructor's lip movements and speech can trigger the uncanny valley effect, leading to distraction. The linear superposition algorithm validated in this system ensures that the virtual instructor's facial expressions remain natural, stable, and predictable during lengthy explanations of standard operating procedures (SOPs), thereby reducing the cognitive load on trainees and allowing them to fully concentrate on learning the key technical points.
[0079] Finally, the high sensitivity threshold of the AI semantic retrieval response effectively constitutes a safety knowledge firewall for the system. In automotive production line training, incorrect guidance can lead to serious safety accidents. This response curve shows that the system is highly selective in retrieving knowledge, only using retrieved SOP fragments with a high semantic similarity (greater than 0.8) to generate data. This mechanism effectively suppresses the machine illusion phenomenon that large language models are prone to in industrial vertical fields, ensuring that every suggestion provided by the AI intelligent tutor regarding torque values, tightening sequence, or safety clearances is verifiable and strictly compliant, thus providing an underlying logical guarantee for zero-risk training of high-risk operations.
Claims
1. An interactive training assistance system for automotive production lines based on AI and VR, characterized in that, include: The hardware interactive terminal (10) is used to collect the user's motion and posture data and voice stream in the physical space in real time, and to present the rendered virtual training screen to the user. The VR virtual simulation module (20) is used to construct a virtual workshop scene containing automobile production line equipment, receive the motion pose data and process the physical interaction logic between the user and the virtual parts, and transmit the virtual workshop scene containing the interaction results to the hardware interaction terminal (10). The AI intelligent tutor module (30) is used to receive the voice stream, generate teaching feedback based on the production line standard operating procedure, and generate drive instructions to control the virtual digital human embedded in the virtual workshop scene, and provide the user with real-time voice explanation and action demonstration in conjunction with the physical interaction logic; The public technology support layer (40) is used to provide development engine support for the VR virtual simulation module (20) and the AI intelligent tutor module (30), and to adapt the hardware interaction terminal (10) through the standard interface to establish a data communication channel.
2. The AI and VR-based interactive training assistance system for automotive production lines according to claim 1, characterized in that, The hardware interaction terminal (10) includes a VR head-mounted display (11) and an interaction controller (12). The VR virtual simulation module (20) includes: The virtual workshop scene unit (21) is used to render the three-dimensional model data containing the automobile production line equipment using a general rendering pipeline to generate a visualized virtual workshop scene. The physics engine unit (23) is used to calculate the collision, gravity and rigid body dynamics effects between the virtual parts and to manage the interaction relationships between objects at different levels in the virtual workshop scene through the collision detection matrix. The interactive control unit (22) is used to convert the motion pose data input by the hardware interactive terminal (10) into interactive events in the virtual environment, and to execute the ray interaction mode and the direct grab mode according to the interactive events.
3. The AI and VR-based interactive training assistance system for automotive production lines according to claim 2, characterized in that, The process by which the interactive control unit (22) executes the ray interaction mode using the hardware interactive terminal (10) includes: The real-time position of the interactive handle (12) in the hardware interactive terminal (10) is obtained as the emission origin. A virtual ray is emitted along the normalized direction vector pointed by the interactive handle (12), and the virtual ray is mapped to the virtual workshop scene as the user's pointing input, thereby establishing the user's long-distance pointing path in the virtual workshop scene. The set of geometric trajectory points of the virtual ray is calculated based on the preset ray length parameters. The system detects in real time whether the set of geometric trajectory points overlaps with the surface of the collision body of the object in the virtual workshop scene, thereby determining the interaction intent. When a valid target is detected, a haptic feedback signal is triggered to the interactive handle (12), and the outline of the valid target is highlighted visually, thereby completing the non-contact interactive confirmation of the valid target.
4. The AI and VR-based interactive training assistance system for automotive production lines according to claim 2, characterized in that, The process by which the interactive control unit (22) works with the physics engine unit (23) to execute the direct grabbing mode includes: Preset target mounting positions for the detachable virtual parts on the vehicle body in the virtual workshop scenario, and set an adsorption area detector at the target mounting position to provide a spatial reference for supporting auxiliary assembly logic based on the adsorption area; When the user grabs the virtual component close to the target mounting position through the direct grab mode, the adsorption area detector is used to calculate in real time the Euclidean distance between the current position of the virtual component and the target mounting position, as well as the angle difference between the current attitude quaternion and the target attitude quaternion. If both the Euclidean distance and the angle difference are less than a preset tolerance threshold, a semi-transparent pre-installed ghost image is visually displayed, indicating the alignment status, thereby guiding the user to perform the operation; When the signal from the user releasing the grip button is received, the position and posture of the virtual component are forcibly aligned to the target installation position instantly, and the physical movement of the virtual component is locked, ultimately achieving positioning and assembly.
5. The AI and VR-based interactive training assistance system for automotive production lines according to claim 1, characterized in that, The AI intelligent tutor module (30) includes: The speech recognition unit (31) is used to receive the speech stream collected by the hardware interactive terminal (10) and convert the speech stream into text data; A professional knowledge base (34) is used to pre-store the production line standard operating procedures and safety specifications data used to generate the teaching feedback; The semantic understanding and generation unit (32) is used to generate the response text corresponding to the teaching feedback by combining the production line standard operating procedures and the safety specification data in the professional knowledge base (34) with the large language model; The digital human driving unit (33) is used to receive the reply text, generate audio, and drive the facial expressions and body movements of the virtual digital human.
6. The AI and VR-based interactive training assistance system for automotive production lines according to claim 5, characterized in that, The process by which the speech recognition unit (31) converts the speech stream into text data includes: Construct a pre-defined domain dataset containing automotive manufacturing terminology, mix clean speech commands with industrial background noise simulating the automotive production line equipment at different signal-to-noise ratios to generate an enhanced dataset, and complete the data preparation based on domain enhancement to support noise-resistant training. The speech recognition model is fine-tuned using the augmented dataset. The optimization objective is to minimize the cross-entropy loss between the predicted text sequence and the real text sequence, thereby giving the speech recognition model the ability to resist industrial noise interference. In the decoding stage, a cluster search strategy is adopted, and additional probability weights are given to automotive production line-specific words appearing in the predicted probability distribution to correct phoneme confusion caused by noise. Finally, the text data with high accuracy is generated and passed to the semantic understanding and generation unit (32).
7. The AI and VR-based interactive training assistance system for automotive production lines according to claim 5, characterized in that, The process by which the semantic understanding and generation unit (32) generates the response text corresponding to the teaching feedback includes: The text embedding model is called to convert the text data output by the speech recognition unit (31) into a query vector as the query text. The knowledge vector that is closest to the query vector in the vector database is retrieved, thereby performing the pre-retrieval generated by the retrieval enhancement. The cosine of the angle between the query vector and the knowledge vector is calculated as a similarity score. Based on the similarity score, a highly relevant segment of the production line standard operating procedure is extracted from the professional knowledge base (34) as knowledge enhancement content. A prompt word template containing an instruction area, a context area, and a question area is constructed. The retrieved standard operating procedure fragments of the production line are filled into the context area of the prompt word template, and the questions raised by the user are filled into the question area of the prompt word template. The information is then input into the large language model for reasoning, and finally, a professional technical guidance answer for replying to the user is generated.
8. The AI and VR-based interactive training assistance system for automotive production lines according to claim 5, characterized in that, The process by which the digital human driving unit (33) drives the facial expressions of the virtual digital human includes: An audio feature extractor is used to extract time-varying acoustic feature vectors from the audio generated by the digital human driving unit, which are then used as the input source to drive facial movements. The acoustic feature vector is processed by a facial motion inference network to predict the weight coefficients of each hybrid deformation base shape required to achieve facial expression synchronization at the current time step. The final position of the facial mesh vertices is calculated based on the principle of linear superposition. The sum of the product of the basic vertex position coordinates of the virtual digital human model and the product of the vertex displacement increment vector of each of the hybrid deformation basic shapes and the corresponding weight coefficients is superimposed to update the facial mesh in real time and drive expression changes.
9. The AI and VR-based interactive training assistance system for automotive production lines according to claim 5, characterized in that, The process by which the digital human driving unit (33) drives the body movements of the virtual digital human includes: Receive and parse the structured instructions generated by the semantic understanding and generation unit (32), extract the action type and the target of interest, trigger the state transition in the animation state machine, and thus prepare to execute the pointing interaction based on inverse kinematics; When the structured instruction requires pointing to a target component in the virtual workshop scene, the position coordinates of the target component in the world coordinate system are obtained and set as the target effector position of the inverse kinematics solver. The rotation angles of each joint in the skeletal chain of the virtual digital human arm are calculated in reverse using inverse kinematics technology, and the rotation angles are applied to drive the skeletal model so that the fingertips of the virtual digital human point to the target component in three-dimensional space.
10. The AI and VR-based interactive training assistance system for automotive production lines according to claim 2, characterized in that, The process by which the virtual workshop scene unit (21) renders the 3D model data containing the automobile production line equipment includes: Lighting rendering calculations are performed in accordance with the rendering equations. The outgoing light intensity radiated from any surface point in the virtual workshop scene toward the viewing direction is defined as the basis for lighting calculations. The rendering equation comprehensively calculates the self-illumination brightness of any surface point, as well as the integral of the bidirectional reflection distribution function and the incident light brightness in the hemispherical integral domain, thereby simulating real light and shadow interaction. The reflection properties of the material surface are defined using the bidirectional reflection distribution function. The influence of the projected area is calculated by combining the cosine value of the angle between the incident light direction and the surface normal. Finally, the pixel color value of any surface point under the current viewpoint is determined.