Semantic-based intelligent object interaction and operation method in VR environment

By capturing multimodal interactive input in real time in the VR environment, recognizing user intent, and dynamically adapting to virtual tools and operating modes, the problem of rigid interaction in existing VR training platforms is solved. This enables intelligent semantic operation and automatic task advancement, improving user experience and training efficiency.

CN122172962APending Publication Date: 2026-06-09BEIJING JUNHE CHUANGXIANG TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING JUNHE CHUANGXIANG TECH DEV CO LTD
Filing Date
2026-02-26
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing VR mechanical simulation and maintenance training platforms have rigid interactive logic and lack flexibility. Users must strictly follow a pre-edited sequence of steps and cannot understand the user's abstract intentions, resulting in an unintelligent and unnatural training experience. In particular, the operation is cumbersome and has a low fault tolerance rate in complex equipment or multi-person collaborative scenarios.

Method used

By capturing multimodal interactive input in real time, identifying user intent, and querying the interaction adaptation rule base, the system dynamically adapts to virtual tools and operating modes, providing intelligent guidance and real-time operation assistance, and enabling semantic operation execution and automatic evolution of task status.

Benefits of technology

It enhances the naturalness and immersion of VR operation, reduces misoperation, allows for flexible adjustment of training paths, and significantly improves the efficiency and intelligence level of VR simulation training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122172962A_ABST
    Figure CN122172962A_ABST
Patent Text Reader

Abstract

This invention discloses a semantic-based intelligent object interaction and operation method in a VR environment, including S1, user intent and target recognition; S2, dynamic adaptation of interaction strategies; S3, intelligent guidance and prompt generation; S4, semantic operation execution and assistance; and S5, automatic evolution of task state. By introducing a semantic understanding layer, this invention upgrades traditional trigger-response VR interaction to an intent-understanding-adaptation-guidance intelligent interaction paradigm. The system can understand the abstract intent of "what to do," rather than simply recognizing the simple command "where to tap," thereby automatically matching tools, adapting operation methods, and providing context-related intelligent guidance. This greatly reduces the learning and memory burden of VR operation and improves the naturalness, fluency, and immersion of the interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality simulation and interaction technology, specifically to a semantic-based intelligent object interaction and operation method in a VR environment. Background Technology

[0002] Existing VR mechanical simulation and maintenance training platforms primarily rely on preset, discrete sequence of steps and fixed object properties for their interactive logic. User operations must strictly follow pre-edited "blueprint-style disassembly and assembly sequences," using designated "maintenance tools" to trigger specific "maintenance parts" to complete the steps.

[0003] This interaction method heavily relies on the pre-production editing by project developers, resulting in a rigid and inflexible process. For example, to unscrew a nut, the user must first accurately select "wrench" from the tool list, and then use the wrench to click (trigger) the nut model that is pre-marked as removable. The system cannot understand the user's abstract intent to "loosen" or "remove," nor can it automatically adapt the correct tool and operation method based on the scene context. This makes the VR training experience less intelligent and natural, with a steep learning curve, especially for complex equipment or multi-person collaborative scenarios, resulting in cumbersome operations and a low tolerance for error. Summary of the Invention

[0004] To address this, the present invention provides a semantic-based intelligent object interaction and operation method in a VR environment, in order to solve the problems of insufficient intelligence and naturalness in the VR training experience in the prior art, steep learning curves, and cumbersome operation and low fault tolerance, especially for complex equipment or multi-person collaborative scenarios.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A semantic-based intelligent object interaction and manipulation method in a VR environment includes the following steps:

[0007] S1. User Intent and Target Recognition: Real-time capture of multimodal interactive input from the user in the VR scene, the interactive input including at least one of user gaze focus, voice commands, predefined gestures, and controller pointing; based on the interactive input and attribute information of virtual objects in the current VR scene, identify the user's current interactive intent and at least one target virtual object; the interactive intent includes disassembly, assembly, inspection, maintenance, and operation;

[0008] S2. Dynamic Adaptation of Interaction Strategy: Based on the identified interaction intent and the type, specifications, and status information of the target virtual object, the system queries a pre-set interaction adaptation rule base to automatically determine and adapt the virtual tools, operation modes, and action parameters suitable for the current interaction; the operation modes include grabbing, rotating, translating, spraying, and wiping.

[0009] S3. Intelligent guidance and prompt generation: Based on the adapted virtual tools and operation mode, dynamic interactive guidance information is generated in the VR scene; the guidance information includes: dynamically presenting or highlighting the adapted virtual tool model on the user's virtual hand, visually highlighting the key operation area of ​​the target virtual object, and overlaying augmented reality operation prompt graphics or text in the scene.

[0010] S4. Semantic Operation Execution and Assistance: Responding to physical actions performed by the user through the interactive handle that conform to the operation mode, driving the virtual tool to interact with the target virtual object, and providing real-time operation assistance based on the physics engine and semantic rules; the real-time operation assistance includes: automatic snap-in alignment of the tool and the target object, real-time visual feedback of the operation force and angle, and real-time detection of operation compliance.

[0011] S5. Automatic task status evolution: After the semantic operation corresponding to the interaction intent is detected, the status of the target virtual object is automatically updated, and the task objective of the next stage is automatically advanced or prompted according to the predefined work process logic.

[0012] Preferably, in step S1, recognizing the interaction intent and the target virtual object based on the interactive input specifically includes:

[0013] S11. Eye Focus Determination: The first candidate virtual object on which the user's current eye focus falls is determined by the eye-tracking module of the head-mounted display device or scene ray collision detection.

[0014] S12, Multimodal instruction parsing: Simultaneously analyzes the text content of the received voice instruction or the meaning of the predefined gesture, and parses out the action verbs and object nouns;

[0015] S13. Intent and Target Fusion Determination: Match and fuse the candidate object determined by the gaze focus with the object noun parsed by the multimodal instruction. If they are consistent or have an inclusion relationship, the object is determined as the target virtual object, and the parsed action verb is mapped to the corresponding interaction intent. If the instruction does not specify the object, the first candidate virtual object is taken as the target virtual object.

[0016] Preferably, the interaction adaptation rule base contains a knowledge base of multiple adaptation rules. Each rule has the following structure: a condition part and an action part. The condition part consists of a combination of intent type, object type, object specification parameters, and the object's current state. The action part defines the corresponding virtual tool type, specific tool parameters, recommended operation mode, default action parameters, and expected operation feedback effect. The query and adaptation process in step S2 is as follows: the intent, target object type, specification, and state output in step S1 are used as query conditions, and a match is made in the rule base. If a unique rule is matched, its action part is executed. If multiple rules are matched, the optimal rule is selected based on rule priority or context, or an option list is generated for user confirmation.

[0017] Preferably, the step S3 of dynamically displaying or highlighting the adapted virtual tool model on the user's virtual hand specifically involves:

[0018] S31. Tool Model Instantiation: Based on the adaptation results, load the corresponding 3D tool model from the virtual tool resource library;

[0019] S32. Hand posture adaptation: Based on the type of the tool and the predefined grip point information, adjust the position and rotation of the tool model so that it is bound to the user's virtual hand model in a correct ergonomic posture.

[0020] S33, Dynamic Attachment and Highlighting: Dynamically attach the bound tool model to the virtual hand and apply a continuous highlighting or contour lighting effect to the tool model until the current interaction intent is completed or the user actively switches tools.

[0021] Preferably, providing real-time operation assistance in step S4 specifically includes:

[0022] S41, Intelligent Adsorption Alignment: When it is detected that the distance between the virtual tool and the target virtual object is less than a first threshold and the angle between the tool orientation and the ideal operating direction is less than a second threshold, a gentle force field or position interpolation is automatically applied so that the tool tip is precisely aligned with and adsorbed onto the operating interface of the target object.

[0023] S42. Visualization of the operation process: When the user performs rotation or translation operations, the current operation parameter values ​​are calculated and displayed in real time. The operation parameter values ​​include rotation angle, displacement distance, and applied virtual force value. The current value is compared with the preset standard value range, and real-time feedback is provided in the form of progress bar, color change or numerical superposition.

[0024] S43. Real-time compliance monitoring: During operation, continuously monitor whether the operating parameters exceed the safety range, whether the tool experiences unexpected collisions, and whether the operation sequence conforms to the process constraints. When potential or actual violations are detected, generate visual, audible, or tactile warnings.

[0025] Preferably, the intelligent adsorption alignment in step S41 further includes a direction pre-alignment step: before the tool enters the adsorption distance, i.e., according to the spatial direction of the target object's operation interface, a semi-transparent tool shadow in the correct orientation is previewed near the user's virtual hand or on the tool model.

[0026] Preferably, the determination criterion for detecting the completion of the semantic operation corresponding to the interaction intent in step S5 is at least one of the following:

[0027] Physical state achieved: The physics engine detects that the target virtual object has been separated from its assembly position and moved to a specified distance, or has been installed in the correct position from its disassembled state and meets the connection constraints.

[0028] Operation parameters meet the standards: The system detects that the operation parameters performed by the user have met or exceeded the preset completion standards, such as the rotation angle reaching the specified number of revolutions or the smearing action covering the specified area.

[0029] User explicit confirmation: Receives a completion confirmation command from the user via voice or gamepad button;

[0030] The process logic satisfies the following: when the current interaction intent is a process step, all its sub-actions are determined by the system to be completed.

[0031] Preferably, the method further includes step S0: constructing a scene semantic knowledge base, which is executed before the interaction begins, and is used to inject semantically understandable attribute information into virtual objects in the VR scene; the attribute information includes at least: the object's unique identifier, general type name, specific specification parameters, functional role in the device, assembly relationship with other objects, acceptable set of interactive intents, and corresponding standard operation parameter range.

[0032] The present invention has the following advantages: By introducing a semantic understanding layer, the present invention upgrades the traditional trigger-response VR interaction to an intelligent interaction paradigm of intent-understanding-adaptation-guidance. The system can understand the abstract intent of the user "what to do", rather than just recognizing the simple instruction of "where to click", thereby automatically matching tools, adapting operation methods and providing context-related intelligent guidance. This greatly reduces the learning and memory burden of VR operation and improves the naturalness and immersion of the interaction.

[0033] In complex processes such as maintenance training and equipment operation, this invention can effectively reduce misoperation, guide users to follow the correct procedures, and flexibly adjust the training path according to the user's intention, significantly improving the efficiency and intelligence level of VR simulation training. Attached Figure Description

[0034] To more intuitively illustrate the prior art and this application, exemplary drawings are provided below. It should be understood that the specific shapes and structures shown in the drawings should not generally be regarded as limiting conditions for implementing this application; for example, based on the technical concept disclosed in this application and the exemplary drawings, those skilled in the art are able to easily make conventional adjustments or further optimizations to the addition / reduction / classification, specific shapes, positional relationships, connection methods, size ratios, etc. of certain units (components).

[0035] Figure 1 A flowchart illustrating a semantic-based intelligent object interaction and operation method in a VR environment provided in this application embodiment. Detailed Implementation

[0036] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these embodiments are merely for further explanation of the present invention and should not be construed as limiting the scope of protection of the present invention. Technical engineers in the field can make some non-essential improvements and adjustments to the present invention based on the above-described content. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Please see Figure 1 A semantic-based intelligent object interaction and manipulation method in a VR environment includes the following steps:

[0038] S1. User Intent and Target Recognition: Real-time capture of multimodal interactive input from the user in the VR scene, the interactive input including at least one of user gaze focus, voice commands, predefined gestures, and controller pointing; based on the interactive input and attribute information of virtual objects in the current VR scene, identify the user's current interactive intent and at least one target virtual object; the interactive intent includes disassembly, assembly, inspection, maintenance, and operation;

[0039] S2. Dynamic Adaptation of Interaction Strategy: Based on the identified interaction intent and the type, specifications, and status information of the target virtual object, the system queries a pre-set interaction adaptation rule base to automatically determine and adapt the virtual tools, operation modes, and action parameters suitable for the current interaction; the operation modes include grabbing, rotating, translating, spraying, and wiping.

[0040] S3. Intelligent guidance and prompt generation: Based on the adapted virtual tools and operation mode, dynamic interactive guidance information is generated in the VR scene; the guidance information includes: dynamically presenting or highlighting the adapted virtual tool model on the user's virtual hand, visually highlighting the key operation area of ​​the target virtual object, and overlaying augmented reality operation prompt graphics or text in the scene.

[0041] S4. Semantic Operation Execution and Assistance: Responding to physical actions performed by the user through the interactive handle that conform to the operation mode, driving the virtual tool to interact with the target virtual object, and providing real-time operation assistance based on the physics engine and semantic rules; the real-time operation assistance includes: automatic snap-in alignment of the tool and the target object, real-time visual feedback of the operation force and angle, and real-time detection of operation compliance.

[0042] S5. Automatic task status evolution: After the semantic operation corresponding to the interaction intent is detected, the status of the target virtual object is automatically updated, and the task objective of the next stage is automatically advanced or prompted according to the predefined work process logic.

[0043] In implementation, this method integrates multimodal inputs such as gaze, voice, and gestures in step S1, and combines them with predefined semantic attributes of objects in the scene (such as "nut" and "to be disassembled") to proactively "understand" what the user wants to "do" (such as "remove this"), rather than passively waiting for clicks. Step S2 introduces an "interaction adaptation rule base," enabling the system to automatically decide "what tool to use" (such as a 10mm wrench) and "how to operate" (such as counterclockwise rotation) based on the recognized "intent" and "target object attributes" (such as an M10 nut). This replaces the rigid mode in traditional VR where users must manually select tools from a list and strictly follow pre-programmed steps. Step S3 transforms the adaptation results into intuitive guidance (such as highlighting the correct wrench in the user's hand and displaying a rotation arrow on the nut), greatly reducing the user's memory and search burden. Step S4 provides intelligent assistance during execution (such as automatic alignment and force feedback), improving operational accuracy and realism. Step S5 enables automatic progression of the task state. This series of steps forms a complete closed loop from "intent understanding" to "intelligent adaptation" and then to "guided execution," upgrading VR interaction from mechanical operation based on "preset instructions" to flexible and natural dialogue based on "semantic intent," significantly improving interaction efficiency, training effectiveness, and the immersive experience of the user.

[0044] For example, when the intent is "disassembly" and the target object attributes are "Type: Nut, Specification: M10", the rule base will adapt to "Tool Type: Open-end Wrench, Tool Parameters: Opening 10mm, Operation Mode: Rotation around axis (loosening)". In the multimodal interactive input, the voice command can be, for example, "Remove this nut", and the gesture can be, for example, a pinching and then rotating wrist action. In S3, the AR operation prompt graphics can include arrow animations indicating the direction of rotation, dashed boxes indicating the area to be painted, etc. The real-time operation assistance in S4, for example, when the user brings the wrench close to the nut, automatically fine-tunes the wrench posture to make it perfectly fit into the nut's edge; when the user rotates the handle, a comparison bar between the rotation angle and the standard torque is displayed in real time. In S5, when the system determines that the intent to "disassemble the nut" has been achieved through physical detection or semantic judgment (such as the nut displacement exceeding a threshold, voice confirmation completion), it automatically marks the nut status as "disassembled", checks this step in the task list, and highlights the next part to be disassembled.

[0045] In step S1, recognizing the interaction intent and target virtual object based on interactive input specifically includes:

[0046] S11. Eye Focus Determination: The first candidate virtual object on which the user's current eye focus falls is determined by the eye-tracking module of the head-mounted display device or scene ray collision detection.

[0047] S12, Multimodal instruction parsing: Simultaneously analyzes the text content of the received voice instruction or the meaning of the predefined gesture, and parses out the action verbs and object nouns;

[0048] S13. Intent and Target Fusion Determination: Match and fuse the candidate object determined by the gaze focus with the object noun parsed by the multimodal instruction. If they are consistent or have an inclusion relationship, the object is determined as the target virtual object, and the parsed action verb is mapped to the corresponding interaction intent. If the instruction does not specify the object, the first candidate virtual object is taken as the target virtual object.

[0049] By initially identifying candidate objects through "eye focus determination," and then extracting action and object keywords from speech or gestures through "multimodal command parsing," the system finally performs "fusion judgment," improving the accuracy and robustness of intent and target recognition. For example, when a user looks at a nut and says "remove it," the system can reliably bind "remove" to the nut at the eye focus; even if the object is omitted in the command (such as only saying "remove"), the system can still default to executing the disassembly intent on the object at the eye focus. This approach is closer to the natural collaborative communication habits of humans, reducing operational ambiguity caused by unclear target specification, and making human-computer interaction smoother and more intuitive.

[0050] The interaction adaptation rule base contains a knowledge base of multiple adaptation rules. Each rule has the following structure: a condition part and an action part. The condition part consists of a combination of intent type, object type, object specification parameters, and the object's current state. The action part defines the corresponding virtual tool type, specific tool parameters, recommended operation mode, default action parameters, and expected operation feedback effect. The query and adaptation process in step S2 is as follows: the intent, target object type, specification, and state output in step S1 are used as query conditions, and the results are matched in the rule base. If a unique rule is matched, its action part is executed. If multiple rules are matched, the optimal rule is selected based on rule priority or context, or an option list is generated for user confirmation.

[0051] Encoding expert knowledge or operating procedures into structured condition-action rules makes the system's intelligent decision-making process predictable, manageable, and easy to maintain and extend. When equipment or procedures are updated, maintenance personnel only need to add, delete, or modify the corresponding rule entries, without rewriting the core logic. Simultaneously, by setting rule priorities or selecting rules based on context, the system can handle complex or boundary situations, ensuring that appropriate and optimal interaction strategies are assigned to the user's semantic intent, guaranteeing the consistency and correctness of interaction behavior.

[0052] In step S3, dynamically displaying or highlighting the adapted virtual tool model on the user's virtual hand specifically involves:

[0053] S31. Tool Model Instantiation: Based on the adaptation results, load the corresponding 3D tool model from the virtual tool resource library;

[0054] S32. Hand posture adaptation: Based on the type of the tool and the predefined grip point information, adjust the position and rotation of the tool model so that it is bound to the user's virtual hand model in a correct ergonomic posture.

[0055] S33, Dynamic Attachment and Highlighting: Dynamically attach the bound tool model to the virtual hand and apply a continuous highlighting or contour lighting effect to the tool model until the current interaction intent is completed or the user actively switches tools.

[0056] Through a process of "instantiation-pose adaptation-dynamic attachment and highlighting," immediate, clear, and error-free visual feedback is provided. Users do not need to search through menus; the correct tool will automatically appear in their hand in an ergonomic posture and be highlighted, clearly instructing them to use it. This not only significantly reduces tool preparation time but also avoids operation interruptions or errors caused by selecting the wrong tool, making the interaction process more coherent and efficient.

[0057] Step S4 provides real-time operation assistance, specifically including:

[0058] S41, Intelligent Adsorption Alignment: When it is detected that the distance between the virtual tool and the target virtual object is less than a first threshold and the angle between the tool orientation and the ideal operating direction is less than a second threshold, a gentle force field or position interpolation is automatically applied so that the tool tip is precisely aligned with and adsorbed onto the operating interface of the target object.

[0059] S42. Visualization of the operation process: When the user performs rotation or translation operations, the current operation parameter values ​​are calculated and displayed in real time. The operation parameter values ​​include rotation angle, displacement distance, and applied virtual force value. The current value is compared with the preset standard value range, and real-time feedback is provided in the form of progress bar, color change or numerical superposition.

[0060] S43. Real-time compliance monitoring: During operation, continuously monitor whether the operating parameters exceed the safety range, whether the tool experiences unexpected collisions, and whether the operation sequence conforms to the process constraints. When potential or actual violations are detected, generate visual, audible, or tactile warnings.

[0061] Intelligent adsorption alignment reduces the operational difficulty of precision assembly and improves the success rate; the visualization of the operation process transforms abstract forces and angles into intuitive graphical feedback, helping users to precisely control operations and comply with standard procedures; real-time compliance detection acts like an expert providing real-time guidance, promptly preventing and correcting operational errors. The combination of these three elements lowers the professional threshold for VR operation while ensuring the standardization and safety of training or operation processes, enabling even beginners to quickly master complex operational techniques.

[0062] The intelligent adsorption alignment in step S41 also includes a pre-alignment step: before the tool enters the adsorption range, based on the spatial orientation of the target object's operating interface, a semi-transparent, correctly oriented tool silhouette is previewed near the user's virtual hand or on the tool model to guide the user to adjust the actual tool to the approximate correct orientation. The preview silhouette serves as a demonstration of the target posture, providing directional guidance to the user before actual contact with the object, allowing for advance adjustment of gestures and tool orientation. This further simplifies fine-tuning operations (such as inserting a screwdriver into a narrow slot), reduces repeated trial and error, and makes the entire interaction process smoother and more efficient, significantly improving operational efficiency and user experience satisfaction.

[0063] The determination criteria for detecting the completion of the semantic operation corresponding to the interaction intent in step S5 are at least one of the following:

[0064] Physical state achieved: The physics engine detects that the target virtual object has been separated from its assembly position and moved to a specified distance, or has been installed in the correct position from its disassembled state and meets the connection constraints.

[0065] Operation parameters meet the standards: The system detects that the operation parameters performed by the user have met or exceeded the preset completion standards, such as the rotation angle reaching the specified number of revolutions or the smearing action covering the specified area.

[0066] User explicit confirmation: Receives a completion confirmation command from the user via voice or gamepad button;

[0067] The process logic satisfies the following: when the current interaction intent is a process step, all its sub-actions are determined by the system to be completed.

[0068] The above solution enables the system to flexibly and accurately perceive task progress and adapt to the needs of different interaction scenarios. By combining multiple judgment methods such as physical state, parameter attainment, user confirmation, and process logic, the system can handle both goals with clear physical state changes, such as "tightening the nut," and steps that rely more on processes or confirmation, such as "completing the check."

[0069] It also includes step S0: Scene semantic knowledge base construction, which is executed before the interaction begins, and is used to inject semantically understandable attribute information into virtual objects in the VR scene; the attribute information includes at least: the object's unique identifier, general type name, specific specification parameters, functional role in the device, assembly relationship with other objects, acceptable set of interactive intentions and corresponding standard operation parameter range, thereby improving the interpretability and maintainability of the system, and making the entire interaction model built on clear knowledge.

[0070] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A semantic-based intelligent object interaction and manipulation method in a VR environment, characterized in that, Includes the following steps: S1. User Intent and Target Recognition: Real-time capture of multimodal interactive input from the user in the VR scene, the interactive input including at least one of user gaze focus, voice commands, predefined gestures, and controller pointing; based on the interactive input and attribute information of virtual objects in the current VR scene, identify the user's current interactive intent and at least one target virtual object; the interactive intent includes disassembly, assembly, inspection, maintenance, and operation; S2. Dynamic Adaptation of Interaction Strategy: Based on the identified interaction intent and the type, specifications, and status information of the target virtual object, the system queries a pre-set interaction adaptation rule base to automatically determine and adapt the virtual tools, operation modes, and action parameters suitable for the current interaction; the operation modes include grabbing, rotating, translating, spraying, and wiping. S3. Intelligent guidance and prompt generation: Based on the adapted virtual tools and operation mode, dynamic interactive guidance information is generated in the VR scene; The guidance information includes: dynamically presenting or highlighting the adapted virtual tool model on the user's virtual hand, visually highlighting the key operation areas of the target virtual object, and overlaying augmented reality operation prompt graphics or text in the scene; S4. Semantic Operation Execution and Assistance: Responding to physical actions performed by the user through the interactive handle that conform to the operation mode, driving the virtual tool to interact with the target virtual object, and providing real-time operation assistance based on the physics engine and semantic rules; the real-time operation assistance includes: automatic snap-in alignment of the tool and the target object, real-time visual feedback of the operation force and angle, and real-time detection of operation compliance. S5. Automatic task status evolution: After the semantic operation corresponding to the interaction intent is detected, the status of the target virtual object is automatically updated, and the task objective of the next stage is automatically advanced or prompted according to the predefined work process logic.

2. The semantic-based intelligent object interaction and operation method in a VR environment according to claim 1, characterized in that, In step S1, recognizing the interaction intent and target virtual object based on interactive input specifically includes: S11. Eye Focus Determination: The first candidate virtual object on which the user's current eye focus falls is determined by the eye-tracking module of the head-mounted display device or scene ray collision detection. S12, Multimodal instruction parsing: Simultaneously analyzes the text content of the received voice instruction or the meaning of the predefined gesture, and parses out the action verbs and object nouns; S13. Intent and Target Fusion Determination: Match and fuse the candidate object determined by the gaze focus with the object noun parsed by the multimodal instruction. If they are consistent or have an inclusion relationship, the object is determined as the target virtual object, and the parsed action verb is mapped to the corresponding interaction intent. If the instruction does not specify the object, the first candidate virtual object is taken as the target virtual object.

3. The semantic-based intelligent object interaction and operation method in a VR environment according to claim 1, characterized in that, The interaction adaptation rule base contains a knowledge base of multiple adaptation rules. Each rule has the following structure: a condition part and an action part. The condition part consists of a combination of intent type, object type, object specification parameters, and the object's current state. The action part defines the corresponding virtual tool type, specific tool parameters, recommended operation mode, default action parameters, and expected operation feedback effect. The query and adaptation process in step S2 is as follows: the intent, target object type, specification, and state output in step S1 are used as query conditions, and the results are matched in the rule base. If a unique rule is matched, its action part is executed. If multiple rules are matched, the optimal rule is selected based on rule priority or context, or an option list is generated for user confirmation.

4. The semantic-based intelligent object interaction and operation method in a VR environment according to claim 1, characterized in that, In step S3, dynamically displaying or highlighting the adapted virtual tool model on the user's virtual hand specifically involves: S31. Tool Model Instantiation: Based on the adaptation results, load the corresponding 3D tool model from the virtual tool resource library; S32. Hand posture adaptation: Based on the type of the tool and the predefined grip point information, adjust the position and rotation of the tool model so that it is bound to the user's virtual hand model in a correct ergonomic posture. S33, Dynamic Attachment and Highlighting: Dynamically attach the bound tool model to the virtual hand and apply a continuous highlighting or contour lighting effect to the tool model until the current interaction intent is completed or the user actively switches tools.

5. The semantic-based intelligent object interaction and operation method in a VR environment according to claim 1, characterized in that, Step S4 provides real-time operation assistance, specifically including: S41, Intelligent Adsorption Alignment: When it is detected that the distance between the virtual tool and the target virtual object is less than a first threshold and the angle between the tool orientation and the ideal operating direction is less than a second threshold, a gentle force field or position interpolation is automatically applied so that the tool tip is precisely aligned with and adsorbed onto the operating interface of the target object. S42. Visualization of the operation process: When the user performs rotation or translation operations, the current operation parameter values ​​are calculated and displayed in real time. The operation parameter values ​​include rotation angle, displacement distance, and applied virtual force value. The current value is compared with the preset standard value range, and real-time feedback is provided in the form of progress bar, color change or numerical superposition. S43. Real-time compliance monitoring: During operation, continuously monitor whether the operating parameters exceed the safety range, whether the tool experiences unexpected collisions, and whether the operation sequence conforms to the process constraints. When potential or actual violations are detected, generate visual, audible, or tactile warnings.

6. The semantic-based intelligent object interaction and operation method in a VR environment according to claim 5, characterized in that, The intelligent adsorption alignment in step S41 also includes a direction pre-alignment step: before the tool enters the adsorption distance, i.e., according to the spatial orientation of the target object's operation interface, a semi-transparent tool phantom that is in the correct orientation is previewed near the user's virtual hand or on the tool model.

7. The semantic-based intelligent object interaction and operation method in a VR environment according to claim 1, characterized in that, The determination criteria for detecting the completion of the semantic operation corresponding to the interaction intent in step S5 are at least one of the following: Physical state achieved: The physics engine detects that the target virtual object has been separated from its assembly position and moved to a specified distance, or has been installed in the correct position from its disassembled state and meets the connection constraints. Operation parameters meet the standards: The system detects that the operation parameters performed by the user have met or exceeded the preset completion standards, such as the rotation angle reaching the specified number of revolutions or the smearing action covering the specified area. User explicit confirmation: Receives a completion confirmation command from the user via voice or gamepad button; The process logic satisfies the following: when the current interaction intent is a process step, all its sub-actions are determined by the system to be completed.

8. The semantic-based intelligent object interaction and operation method in a VR environment according to claim 1, characterized in that, It also includes step S0: Scene semantic knowledge base construction, which is executed before the interaction begins, and is used to inject semantically understandable attribute information into virtual objects in the VR scene; The attribute information includes at least: the object's unique identifier, general type name, specific specification parameters, functional role in the device, assembly relationship with other objects, acceptable set of interactive intents, and corresponding standard operating parameter range.