Virtual reality intelligent education method and system based on industrial multi-modal large model driving
By scanning target objects in a virtual reality education system to obtain multimodal information and then analyzing it using a large multimodal model, the problem of lack of personalization and real-time feedback in virtual reality education systems is solved, thus achieving personalized learning experiences and efficient learning.
Patent Information
- Application Number
- CN202411608240.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-11-12
AI Technical Summary
Existing virtual reality education systems lack personalization and real-time feedback, and cannot dynamically adjust according to learners' real-time performance, resulting in low learning efficiency.
A virtual reality intelligent education method based on an industrial multimodal large model is adopted. Multimodal information is obtained by scanning target objects, and the multimodal large model is used for analysis, processing and feedback, and real-time correction of user questions.
It provides a personalized learning experience, enhances learning efficiency and engagement, enables real-time error correction and feedback, reduces training costs, minimizes reliance on physical equipment, and ensures safe and efficient learning.
Smart Images

Figure CN119600397B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent education, in particular to a virtual reality intelligent education method and system based on industrial multi-modal large model driving. BACKGROUND
[0002] Most existing education systems rely on pre-set static teaching content, lacking intelligent and personalized learning experience. Virtual reality technology can provide virtual teaching scenarios such as simulated laboratories and historical scene reproductions, improving students' learning interest and understanding ability.
[0003] However, traditional virtual reality education systems mainly rely on pre-set teaching content and cannot dynamically adjust according to learners' real-time performance. Learners' operation level, knowledge mastery, and other factors cannot affect the system's feedback and content generation, limiting the possibility of personalized learning. Moreover, traditional virtual reality education systems usually provide post-feedback, which cannot realize real-time guidance during learners' operation. When learners encounter problems, they often need to complete the entire procedure before receiving feedback, resulting in low learning efficiency.
[0004] Therefore, it is necessary to provide a new virtual reality intelligent education method and system based on industrial multi-modal large model driving. SUMMARY
[0005] In view of the above problems existing in the prior art, the purpose of the embodiments of the present application is to provide a virtual reality intelligent education method and system based on industrial multi-modal large model driving, which can provide personalized learning experience, provide safe, convenient, and efficient learning methods for users, and can correct errors and provide feedback in real time.
[0006] To achieve the above purpose, the technical solution adopted by the present application is: a virtual reality intelligent education method based on industrial multi-modal large model driving, comprising:
[0007] S1, in a virtual reality scene, scanning a target object using a scanning function;
[0008] S2, obtaining multi-modal information of the target object and encoding and formatting the multi-modal information;
[0009] S3, inputting the multi-modal information into a multi-modal large model for analysis and processing, generating a processing result and feeding back to the user;
[0010] S4, obtaining question information input by the user, inputting the question information into the multi-modal large model for extraction processing and generating answer information in real time, and the multi-modal large model intelligently correcting errors according to industrial knowledge involved in the question information;
[0011] S5, judge whether the user continues to ask questions, when the user continues to ask questions, then return to step S4, otherwise, close the system dialogue function.
[0012] Further, in S1, the target object is scanned in the virtual reality scene by using a scanning function, comprising:
[0013] Step S11, the user explicitly targets the object and calls the scanning function of the system;
[0014] Step S12, search and lock the target object in the virtual reality scene;
[0015] Step S13, the target object is scanned by the scanning tool built in the system according to the scanning technology adopted.
[0016] Further, the scanning technology includes laser scanning, structured light scanning and photogrammetry, the laser scanning obtains the three-dimensional coordinate information of the target object surface by emitting laser beam and measuring the reflection of laser on the object surface; the structured light scanning calculates the three-dimensional shape of the object by projecting a specific pattern of light onto the target object surface according to the deformation of the pattern; the photogrammetry reconstructs the three-dimensional shape of the object by taking a large number of photos of the object from different angles and using computer vision algorithm.
[0017] Further, in S2, the multi-modal information of the target object is obtained and encoded and formatted, comprising:
[0018] Step S21, obtain the multi-modal information of the target object, including the three-dimensional model information and the picture information of the target object;
[0019] Step S22, encode and format the multi-modal information, including classifying and integrating the multi-modal information first, and then encoding and formatting the integrated multi-modal information;
[0020] Wherein, obtaining the three-dimensional model information of the target object includes that the system accurately obtains the coordinate information of each point on the object surface by scanning technology, and further constructs the three-dimensional model of the object, and further presents the geometric shape of the object, including whether the overall contour is a regular geometric shape or a complex irregular shape; obtaining the three-dimensional model information of the target object also includes the shape of each component of the target object and the connection relationship between each component; obtaining the picture information of the target object includes that the system takes the picture information of the object from different angles by scanning function.
[0021] Further, in S3, the multi-modal information is input into the multi-modal large model for analysis and processing, and the processing result is fed back to the user, comprising:
[0022] Step S31, the multi-modal large model identifies and extracts features of the object according to the multi-modal information;
[0023] Step S32, the multi-modal large model analyzes the relevant industrial knowledge of the target object according to the results of identification and feature extraction;
[0024] Step S33, the relevant industrial knowledge of the target object analyzed is taken as a processing result, and the processing result is fed back to the user in the virtual reality scene.
[0025] Further, in S4, the user input question information is obtained, the question information is input to the multi-modal large model for extraction processing and real-time generation of answer information, and the multi-modal large model intelligently corrects the industrial knowledge involved in the question information, including:
[0026] Step S41, the user inputs the question information through the microphone device equipped by the system, and the system converts the audio mode question information into text form by using the speech recognition technology and inputs it into the multi-modal large model;
[0027] Step S42, the multi-modal large model analyzes the question information in text form in combination with the natural language processing technology, and extracts the key question information;
[0028] Step S43, the multi-modal large model comprehensively judges according to the key question information, in combination with the associated multi-modal information, the extracted language features, the fused knowledge graph and industrial knowledge, etc., and generates answer information in real time;
[0029] Step S44, the multi-modal large model identifies the industrial knowledge elements involved in the question information, establishes an intelligent correction mechanism based on industrial knowledge, and the multi-modal large model corrects the errors in the industrial knowledge involved in the user's question information.
[0030] Further, the establishment of the intelligent correction mechanism based on industrial knowledge includes product knowledge related correction, industrial process and process knowledge related correction, industrial equipment operation and maintenance knowledge related correction, and industrial standard and specification knowledge related correction.
[0031] The virtual reality intelligent education system driven by the industrial multi-modal large model is applied to the virtual reality intelligent education method driven by the industrial multi-modal large model, and the system comprises:
[0032] A target scanning module is configured to scan a target object in a virtual reality scene by using a scanning function;
[0033] An information acquisition module is configured to acquire multi-modal information of the target object and encode and format the multi-modal information;
[0034] An analysis processing module is configured to input the multi-modal information into the multi-modal large model for analysis processing.
[0035] A result feedback module is configured to feed back the generated processing result to the user.
[0036] A question and answer correction module is configured to acquire the question information input by the user, input the question information into the multi-modal large model for extraction processing and generate answer information in real time, and meanwhile, the multi-modal large model intelligently corrects the industrial knowledge involved in the question information.
[0037] A judgment returning module is configured to judge whether the user continues to ask questions, and if the user continues to ask questions, return, otherwise, close the system dialogue function.
[0038] The embodiment of the application further provides a network side server, comprising:
[0039] at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned virtual reality intelligent education method driven by an industrial multi-modal large model.
[0040] The embodiment of the application further provides a computer readable storage medium storing a computer program, and the computer program is executed by a processor to implement the above-mentioned virtual reality intelligent education method driven by an industrial multi-modal large model.
[0041] The beneficial effects of the application are as follows: in a virtual reality scene, a target object is scanned by using a scanning function; multi-modal information of the target object is acquired, and the multi-modal information is encoded and formatted; the multi-modal information is input into a multi-modal large model for analysis processing, a processing result is generated and fed back to the user; question information input by the user is acquired, the question information is input into the multi-modal large model for extraction processing and answer information is generated in real time, and meanwhile, the multi-modal large model intelligently corrects the industrial knowledge involved in the question information; it is judged whether the user continues to ask questions, and if the user continues to ask questions, the previous step is returned, otherwise, the system dialogue function is closed. The virtual reality intelligent education method driven by an industrial multi-modal large model can provide personalized learning experience, has strong interaction ability and flexibility, provides a safe, convenient and efficient learning mode for the user, can correct errors and provide feedback in real time, acquires multi-modal information by scanning an object, can effectively understand the needs of the user, and answers questions in real time through an intelligent dialogue function, not only can improve learning efficiency, but also can enhance the participation and immersion of the user; the multi-modal large model ensures the accuracy and richness of the information, so that the user can deeply understand the structure, function and operation method of the object. BRIEF DESCRIPTION OF DRAWINGS
[0042] The application will be further described below in conjunction with the accompanying drawings and embodiments.
[0043] In the drawings:
[0044] Figure 1 A flowchart of the virtual reality intelligent education method based on an industrial multi-modal large model driving provided for embodiment one of the application;
[0045] Figure 2 A module schematic diagram of the virtual reality intelligent education system based on an industrial multi-modal large model driving provided for embodiment two of the application;
[0046] Figure 3 A structural schematic diagram of a network side server according to the third embodiment of the application. DETAILED DESCRIPTION
[0047] In order to make the objectives, technical solutions and advantages of the embodiments of the application clearer, the technical solutions of the application will be described below in conjunction with the accompanying drawings, obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0048] First embodiment:
[0049] The first embodiment of the application provides a virtual reality intelligent education method based on an industrial multi-modal large model driving, comprising: in a virtual reality scene, scanning a target object by using a scanning function; obtaining multi-modal information of the target object, and encoding and formatting the multi-modal information; inputting the multi-modal information into a multi-modal large model for analysis and processing, generating a processing result and feeding back to a user; obtaining question information input by the user, inputting the question information into the multi-modal large model for extraction processing and generating answer information in real time, and the multi-modal large model intelligently correcting errors according to industrial knowledge involved in the question information; determining whether the user continues to ask questions, when the user continues to ask questions, returning to the previous step, otherwise, closing the system dialogue function. The virtual reality intelligent education method based on an industrial multi-modal large model driving of the application can provide personalized learning experience, has strong interaction ability and flexibility, provides a safe, convenient and efficient learning mode for the user, and can correct errors and provide feedback in real time.
[0050] The implementation details of the virtual reality intelligent education method based on an industrial multi-modal large model driving of the present embodiment will be described in detail below, the following content is only provided for the implementation details for easy understanding, and is not necessary for implementing the present solution, the specific process of the present embodiment is as followsFigure 1 The embodiment is applied to a virtual reality intelligent education system driven by an industrial multi-modal large model.
[0051] In step S1, a target object is scanned in a virtual reality scene by using a scanning function.
[0052] Specifically, the specific steps of scanning the target object in the virtual reality scene by using the scanning function are as follows.
[0053] In step S11, the user explicitly specifies the target object and calls the scanning function of the system.
[0054] Optionally, the system of the embodiment is equipped with a virtual reality system with a scanning function. Before scanning, the user needs to explicitly specify the target object to be analyzed. For example, in education and training, students need to understand the structure of a certain teaching model. The user can first roughly browse in the virtual reality scene, and then explicitly specify the location or general appearance of the target object; and then call the scanning function on the operation interface of the system. The system usually sets the scanning function in a more conspicuous position so that the user can easily find it. On the operation interface, the scanning function can appear in the form of an icon, a button, or a menu option.
[0055] As an example, the scanning function icon in the system is a camera lens-like pattern placed in the toolbar at the bottom of the screen; or an option in a virtual menu called out by operating the handle, and provided with a text label indicating scanning objects.
[0056] In step S12, the target object is searched for and locked in the virtual reality scene.
[0057] Optionally, the user finds the target object in the virtual reality scene by operating the interactive mode provided by the system; after finding the target object, the user needs to lock the target object by a specific operation so as to accurately scan it. The interactive mode provided by the system includes using a mouse to drag or moving a view angle by operating a motion capture handle.
[0058] As an example, in the virtual reality scene, the user needs to rotate the view angle and observe closely by moving the motion capture handle, so as to find the target object among a large number of teaching models and parts. When the target object is found, the cursor or the user's line of sight is stopped on the target object, and the mouse button or the trigger button of the motion capture handle is pressed to confirm, so as to make the virtual aiming frame accurately cover the target object. When the color of the aiming frame changes to the color indicating successful locking, it means that the target object has been locked.
[0059] In step S13, the target object is scanned by the scanning tool built in the system according to the scanning technology adopted by the scanning tool.
[0060] Optionally, the scanning techniques include laser scanning, structured light scanning, and photogrammetry.
[0061] As an example, laser scanning obtains the three-dimensional coordinate information of the target object surface by emitting a laser beam and measuring the reflection of the laser on the object surface.
[0062] As an example, structured light scanning calculates the three-dimensional shape of the object by projecting a specific pattern of light onto the target object surface and calculating the deformation of the pattern.
[0063] As an example, photogrammetry reconstructs the three-dimensional shape of the object by taking a large number of photos of the object from different angles and using computer vision algorithms.
[0064] Optionally, during the scanning process, the user needs to keep the scanning tool aligned with the target object and perform the corresponding operations according to the system prompts.
[0065] As an example, for some larger or complex-shaped objects, it may be necessary to slowly move the scanning tool around the object to ensure that all surfaces of the object are scanned completely; or press a certain button at a specific time interval to obtain scanning data at different angles as required by the system.
[0066] Step S2, obtain multi-modal information of the target object, and encode and format the multi-modal information.
[0067] Specifically, the specific steps of obtaining multi-modal information of the target object and encoding and formatting the multi-modal information are as follows:
[0068] Step S21, obtain multi-modal information of the target object, including three-dimensional model information and picture information of the target object.
[0069] Optionally, obtaining three-dimensional model information of the target object includes the system accurately obtaining coordinate information of each point on the object surface through scanning technology, and then constructing a three-dimensional model of the object, and then completely presenting the geometric shape of the object, including whether the overall contour is a regular geometric shape or a complex irregular shape.
[0070] Optionally, obtaining three-dimensional model information of the target object also includes the shapes of each component of the target object and the connection relationship between each component.
[0071] As an example, for a mechanical target object, the specific shapes of each gear, shaft, shell, and other components can be obtained, as well as how they cooperate and connect together.
[0072] Optionally, the three-dimensional model information of the target object can also include material properties of the target object. When scanning a metal object using a laser scanning method, the material properties such as the glossiness and roughness of the metal can be initially reflected on the model according to the reflection of the laser on the surface of the object; for objects of different materials, the model will present the material characteristics in different ways. For example, scanning a wooden object can show the texture of the wood, and scanning a plastic object can show the smoothness or transparency of the plastic.
[0073] Optionally, the picture information of the target object includes picture information of the object taken from different angles by the system through the scanning function. The picture can fully demonstrate the appearance of the object, showing the color, surface details, light and shadow effects, and other characteristics from the front, side, back, and various inclined angles. Obtaining the picture information of the target object can provide more detailed information than the three-dimensional model information.
[0074] Further, the multi-modal information of the target object can also include audio information and text information of the target object.
[0075] Optionally, the audio information of the target object includes the function of synchronously collecting audio when scanning the object, so that for some objects with sound-emitting components or that will emit sound during operation, the operating sound can be used as a supplementary understanding of the characteristics of the object, and in subsequent analysis or display, the user can more realistically perceive the running state of the object from the auditory level. The text information of the target object includes that the system can obtain the identification information such as the name, model, and brand of the object through association with existing databases or manual annotation, and the identification information is very important for identifying the identity and classification of the object, which facilitates the user to quickly understand the basic situation of the object in subsequent analysis and processing. In addition, through database association or manual annotation, the system can also obtain text information such as function description and operation guide of the object.
[0076] Step S22, encoding and formatting the multi-modal information, including classifying and integrating the multi-modal information first, and then encoding and formatting the integrated multi-modal information.
[0077] Optionally, classifying the multi-modal information includes dividing according to the type and characteristics of the information. For example, the three-dimensional model information of the object is classified as a category, and the picture information taken from different angles is classified according to the shooting angle, resolution, and other factors.
[0078] Optionally, the system integrates different categories of data to form a set of organically combined multi-modal information. In the integration process, not only should the model and the picture be ensured to correspond to each other in space, but also the semantic association between them should be established, such as associating a certain part on the model with the area in the picture showing the appearance of the part and the related text description.
[0079] Optionally, the encoding and formatting processing of the integrated multi-modal information includes encoding the three-dimensional model information in common formats such as FBX, OBJ, and adjusting the coordinate system, size, rotation angle, etc. For picture information, it is encoded in formats such as JPEG, PNG, and formatted in resolution, color balance, etc. If there is text information, it is encoded in UTF-8 format and its font, size, color, etc. appearance attributes and presentation mode are set.
[0080] Step S3, input the multi-modal information into the multi-modal large model for analysis and processing, generate the processing result and feedback to the user.
[0081] Specifically, the specific steps of inputting the multi-modal information into the multi-modal large model for analysis and processing, generating the processing result and feeding back to the user are as follows:
[0082] Step S31, the multi-modal large model identifies and extracts features of the object according to the multi-modal information.
[0083] Optionally, the multi-modal large model uses computer vision technology to identify the three-dimensional model information and picture information in the input multi-modal information. By analyzing the shape, color, texture, etc. visual features of the object, and comparing with the existing object model database, the type, name, etc. of the object are determined.
[0084] As an example, for a three-dimensional model of an industrial part and related pictures, the model can identify that it is a certain type of bolt or a certain type of gear.
[0085] Further, if the input multi-modal information includes text information, the object name, description, etc. mentioned in the input text information are analyzed and compared with the known industrial product knowledge base or name database.
[0086] Optionally, the modal large model performs feature extraction on the object according to the multi-modal information, including deep extraction of various visual features from the picture information and the three-dimensional model information of the object. In addition to basic shape, color, and texture features, it also includes key part features of the object, such as key connection parts of mechanical parts, operation panel parts of equipment, etc., posture features of the object, such as inclination angle and rotation state of the object in space, and motion features of the object.
[0087] In step S32, the multi-modal large model analyzes the relevant industrial knowledge of the target object based on the recognition and feature extraction results.
[0088] Optionally, the relevant industrial knowledge of the target object includes the overall function and usage mode of the target object. The multi-modal large model infers the role of each component in the overall function implementation of the target object by analyzing the mutual relationship between the component features of the target object and comparing with the known industrial object component function library, and then comprehensively infers the overall function of the target object based on the function of each component. The multi-modal large model extracts the size, interface type, material, etc. of the components in the three-dimensional model of the target object according to the component layout shown in the three-dimensional model information, compares them with the corresponding component features of other known industrial objects, judges the compatibility between them, and then infers the usage mode of the target object.
[0089] In step S33, the relevant industrial knowledge of the target object analyzed is taken as the processing result, and the processing result is fed back to the user in the virtual reality scene.
[0090] Optionally, in the virtual reality scene, a special information display area or mechanism is set to show the processing result to the user in real time and intuitively. The display area can be a fixed floating window located at a suitable position in the user's field of view, or a prompt box that appears dynamically according to the user's interaction with the target object.
[0091] As an example, when the user focuses on the target object in the virtual reality scene by gazing or pointing with the motion capture handle, the display window or prompt box of the processing result will automatically pop up to display the relevant text information.
[0092] Further, interactive functions are added to the displayed processing result in the virtual reality scene. As an example, the user can expand or collapse different parts of the processing result, such as the purpose of the object, the functional characteristics, etc., through the operation of the motion capture handle, so as to view a certain item in more detail or quickly browse the overall information; or when the user clicks on a specific component name mentioned in the text, the corresponding component on the target object will be highlighted in the virtual reality scene, allowing the user to more intuitively see the actual location and appearance of the described component on the object.
[0093] Step S4, obtain the user input question information, input the question information to the multi-modal large model for extraction processing and real-time generation of answer information, and the multi-modal large model intelligently corrects the industrial knowledge involved in the question information.
[0094] Specifically, the specific steps of obtaining the user's question information, inputting the question information to the multi-modal large model for extraction processing and real-time generation of answer information are as follows:
[0095] Step S41, the user inputs the question information through the microphone device equipped by the system, and the system converts the audio mode question information into text form by using speech recognition technology and inputs it into the multi-modal large model.
[0096] Step S42, the multi-modal large model analyzes the question information in text form in combination with natural language processing technology, and extracts the key question information.
[0097] As an example, it is clear that the user is asking about the function of a certain component of the object, the overall manufacturing process, or other aspects.
[0098] Step S43, the multi-modal large model makes a comprehensive judgment according to the key question information, in combination with the associated multi-modal information, the extracted language features, the fused knowledge graph and the industrial knowledge, and generates answer information in real time.
[0099] Since the multi-modal large model usually has a knowledge graph inside, it stores rich industrial domain knowledge and common information of various objects. When processing user questions, the extracted language features and semantic information will be fused with the relevant knowledge in the knowledge graph.
[0100] As an example, the user asks about the main function of the target object, and the multi-modal large model analyzes the component composition and possible function implementation of the device from the three-dimensional model information, references the details of the device appearance and operation interface displayed in the picture, and combines the function description of this type of device in the knowledge graph and the general cognition of this type of device function in the industrial field to generate detailed and accurate answers.
[0101] Specifically, the specific steps of the multi-modal large model intelligently correcting the industrial knowledge involved in the question information include:
[0102] Step S44, the multi-modal large model identifies the industrial knowledge elements involved in the question information, establishes an intelligent correction mechanism based on industrial knowledge, and the multi-modal large model corrects the errors in the industrial knowledge involved in the user's question information.
[0103] Specifically, an intelligent correction mechanism based on industrial knowledge is established, including product knowledge related correction, industrial process and technology knowledge related correction, industrial equipment operation and maintenance knowledge related correction, and industrial standard and specification knowledge related correction.
[0104] As an example, product knowledge related correction includes the model pointing out and correcting when the user's description of the characteristics of industrial products in the question does not match the knowledge mastered by the model, and when the user mistakenly classifies one industrial product into another product category, and the multimodal large model will correct the error.
[0105] As an example, industrial process and technology knowledge related correction includes the model giving correction suggestions when the user makes a mistake in describing the industrial production process and when the user mentions unreasonable industrial process parameters.
[0106] As an example, industrial equipment operation and maintenance knowledge related correction includes the multimodal large model giving correction suggestions when the user makes a mistake in describing the operation steps of industrial equipment and when the user mentions unreasonable industrial equipment maintenance periods.
[0107] As an example, and industrial standard and specification knowledge related correction includes the multimodal large model giving correction suggestions when the user makes a mistake in understanding industrial standards or specifications and when the user makes a mistake in applying industrial standards or specifications.
[0108] Further, by establishing an intelligent correction mechanism based on industrial knowledge, the multimodal large model can effectively correct errors related to industrial knowledge in the user's question information, accurately understand the user's intent, and give more accurate answers that conform to the actual industrial situation, helping users better understand industrial field related knowledge.
[0109] Step S5, determine whether the user continues to ask, if the user continues to ask, return to step S4, otherwise close the system dialogue function.
[0110] The first embodiment of the present application provides a virtual reality intelligent education method based on an industrial multi-modal large model, which comprises the following steps: in a virtual reality scene, a target object is scanned by using a scanning function; multi-modal information of the target object is obtained, and the multi-modal information is encoded and formatted; the multi-modal information is input into a multi-modal large model for analysis and processing, and a processing result is generated and fed back to a user; question information input by the user is obtained, the question information is input into the multi-modal large model for extraction processing and real-time generation of answer information, and the multi-modal large model intelligently corrects industrial knowledge involved in the question information; it is judged whether the user continues to ask questions, and if the user continues to ask questions, the previous step is returned, otherwise the system dialogue function is closed. The virtual reality intelligent education method based on the industrial multi-modal large model has strong interaction ability and flexibility, and can provide personalized learning experience for the user; the multi-modal information is obtained by scanning the object, the user's demand can be effectively understood, and the intelligent dialogue function is used to answer questions in real time, which not only improves the learning efficiency, but also enhances the participation and immersion of the user; the multi-modal large model ensures the accuracy and richness of the information, so that the user can deeply understand the structure, function and operation method of the object.
[0111] The virtual reality intelligent education method based on the industrial multi-modal large model can greatly reduce the cost in training. Compared with traditional physical training, the virtual reality system reduces the dependence on actual equipment and site, not only saving the expensive equipment maintenance cost, but also reducing the cost of site rental and tool purchase; the multi-modal large model is combined with the virtual reality technology to simulate the industrial operation scene, and the learner can train anytime and anywhere without being limited by time and space, and at the same time, through dynamic content generation, the system can automatically adjust the training difficulty to ensure that the learner learns at the most suitable pace, greatly improving the training efficiency; training in a virtual environment avoids the safety risks that may be caused by directly operating real equipment; through the virtual reality scene simulation of complex industrial equipment operation, the user can master the operation steps and safety specifications, and reduce the damage of equipment and accidents caused by misoperation.
[0112] The second embodiment of the present application provides a virtual reality intelligent education system based on an industrial multi-modal large model, which comprises the following steps: in a virtual reality scene, a target object is scanned by using a scanning function; multi-modal information of the target object is obtained, and the multi-modal information is encoded and formatted; the multi-modal information is input into a multi-modal large model for analysis and processing, and a processing result is generated and fed back to a user; question information input by the user is obtained, the question information is input into the multi-modal large model for extraction processing and real-time generation of answer information, and the multi-modal large model intelligently corrects industrial knowledge involved in the question information; it is judged whether the user continues to ask questions, and if the user continues to ask questions, the previous step is returned, otherwise the system dialogue function is closed. The virtual reality intelligent education method based on the industrial multi-modal large model has strong interaction ability and flexibility, and can provide personalized learning experience for the user; the multi-modal information is obtained by scanning the object, the user's demand can be effectively understood, and the intelligent dialogue function is used to answer questions in real time, which not only improves the learning efficiency, but also enhances the participation and immersion of the user; the multi-modal large model ensures the accuracy and richness of the information, so that the user can deeply understand the structure, function and operation method of the object.
[0113] As shown in Figure 2 The second embodiment of the present application provides a virtual reality intelligent education system based on an industrial multi-modal large model, which comprises the following steps: in a virtual reality scene, a target object is scanned by using a scanning function; multi-modal information of the target object is obtained, and the multi-modal information is encoded and formatted; the multi-modal information is input into a multi-modal large model for analysis and processing, and a processing result is generated and fed back to a user; question information input by the user is obtained, the question information is input into the multi-modal large model for extraction processing and real-time generation of answer information, and the multi-modal large model intelligently corrects industrial knowledge involved in the question information; it is judged whether the user continues to ask questions, and if the user continues to ask questions, the previous step is returned, otherwise the system dialogue function is closed. The virtual reality intelligent education method based on the industrial multi-modal large model has strong interaction ability and flexibility, and can provide personalized learning experience for the user; the multi-modal information is obtained by scanning the object, the user's demand can be effectively understood, and the intelligent dialogue function is used to answer questions in real time, which not only improves the learning efficiency, but also enhances the participation and immersion of the user; the multi-modal large model ensures the accuracy and richness of the information, so that the user can deeply understand the structure, function and operation method of the object.
[0114] Specifically, the target scanning module 201 is used to scan target objects in a virtual reality scene using a scanning function; the information acquisition module 202 is used to acquire multimodal information of the target object and encode and format the multimodal information; the analysis and processing module 203 is used to input the multimodal information into a multimodal large model for analysis and processing; the result feedback module 204 is used to feed back the generated processing results to the user; the question-and-answer correction module 205 is used to acquire the question information input by the user, input the question information into the multimodal large model for extraction and processing, and generate answer information in real time. At the same time, the multimodal large model performs intelligent error correction based on the industrial knowledge involved in the question information; and the judgment and return module 206 is used to determine whether the user should continue asking questions. If the user continues to ask questions, the system returns; otherwise, the system dialogue function is closed.
[0115] It is not difficult to see that this embodiment is a system implementation corresponding to the first embodiment, and this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.
[0116] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by this invention; however, this does not mean that other units are absent from this embodiment.
[0117] The third embodiment of the present invention relates to a network-side server, such as... Figure 3 As shown, it includes at least one processor 302; and a memory 301 communicatively connected to at least one processor 302; wherein the memory 301 stores instructions executable by at least one processor 302, the instructions being executed by at least one processor 302 to enable at least one processor 302 to perform the above-described data processing method.
[0118] The memory 301 and the processor 302 are connected in a bus manner, the bus can include any number of interconnected buses and bridges, the bus connects one or more processors 302 and various circuits of the memory 301 together. The bus can also connect various other circuits such as peripheral devices, voltage stabilizers and power management circuits together, which are well known in the art, and therefore, further description is not made herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements such as multiple receivers and transmitters, which provide a unit for communicating with various other devices on the transmission medium. The data processed by the processor 302 is transmitted on the wireless medium through the antenna, further, the antenna also receives data and transmits the data to the processor 302.
[0119] The processor 302 is responsible for managing the bus and general processing, and can also provide various functions including timing, peripheral interface, voltage regulation, power management and other control functions. The memory 301 can be used to store data used by the processor 302 in performing operations.
[0120] The fourth embodiment of the present application relates to a computer readable storage medium storing a computer program. The computer program is executed by the processor to realize the virtual reality intelligent education method based on industrial multi-modal large model driving in the first embodiment.
[0121] That is, those skilled in the art can understand that all or part of the steps of the methods in the above embodiments can be completed by a program instructing related hardware, the program is stored in a storage medium, and includes a plurality of instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk and various program code storage media.
[0122] The above-mentioned are only embodiments of the present application, and the common knowledge of the specific structure and characteristics in the scheme is not described too much, the ordinary skilled in the art knows all the ordinary technical knowledge in the field of the present application before the application date or the priority date, can know all the prior art in the field, and has the ability to apply the conventional experimental means before the date, the ordinary skilled in the art can improve and implement the present scheme under the inspiration given by the present application, some typical known structure or known method should not become the obstacle for the ordinary skilled in the art to implement the present application. It should be pointed out that for the skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which should be considered as the protection scope of the present application, which will not affect the effect and practicality of the present application. The protection scope of the present application should be subject to the content of its claims, and the specific implementation mode and the like in the specification can be used to explain the content of the claims.
[0123] The above-mentioned are only embodiments of the present application, and the common knowledge of the specific structure and characteristics in the scheme is not described too much, the ordinary skilled in the art knows all the ordinary technical knowledge in the field of the present application before the application date or the priority date, can know all the prior art in the field, and has the ability to apply the conventional experimental means before the date, the ordinary skilled in the art can improve and implement the present scheme under the inspiration given by the present application, some typical known structure or known method should not become the obstacle for the ordinary skilled in the art to implement the present application. It should be pointed out that for the skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which should be considered as the protection scope of the present application, which will not affect the effect and practicality of the present application. The protection scope of the present application should be subject to the content of its claims, and the specific implementation mode and the like in the specification can be used to explain the content of the claims.
Claims
1. A virtual reality intelligent education method based on an industrial multimodal large model, characterized in that, include: S1, in a virtual reality scene, uses the scanning function to scan the target object; S2, acquire the multimodal information of the target object, and encode and format the multimodal information; S3 inputs multimodal information into a large multimodal model for analysis and processing, generates processing results, and feeds them back to the user; S4: Obtain the question information input by the user, input the question information into the multimodal big model for extraction and processing, and generate answer information in real time. At the same time, the multimodal big model performs intelligent error correction based on the industrial knowledge involved in the question information. S5, determine whether to continue asking questions. If the user continues to ask questions, return to step S4; otherwise, close the system dialogue function. The scanning function used to scan target objects in a virtual reality scene includes: Step S11: The user identifies the target object and invokes the system's scanning function; Step S12: Search for and lock onto the target object in the virtual reality scene; Step S13: The target object is scanned using the system's built-in scanning tool according to the scanning technology it employs; The scanning technology includes laser scanning, structured light scanning, and photogrammetry. Laser scanning obtains the three-dimensional coordinate information of the target object's surface by emitting a laser beam and measuring the reflection of the laser on the object's surface. Structured light scanning projects light in a specific pattern onto the target object's surface and calculates the object's three-dimensional shape based on the deformation of the pattern. Photogrammetry reconstructs the object's three-dimensional shape by taking a large number of photographs of the object from different angles and then using computer vision algorithms. The acquisition of multimodal information of the target object, and the encoding and formatting of the multimodal information, includes: Step S21: Obtain multimodal information of the target object, including the three-dimensional model information and image information of the target object; Step S22, encoding and formatting the multimodal information, including first classifying and integrating the multimodal information, and then encoding and formatting the integrated multimodal information; The acquisition of the target object's 3D model information includes the system accurately acquiring the coordinate information of each point on the object's surface through scanning technology, thereby constructing a 3D model of the object and fully presenting the object's geometric shape, including whether its overall outline is a regular geometric shape or a complex irregular shape; the acquisition of the target object's 3D model information also includes the shape of each component of the target object and the connection relationship between each component; the acquisition of the target object's image information includes the image information of the object taken by the system from different angles through scanning function.
2. The virtual reality intelligent education method based on industrial multimodal large model driving according to claim 1, characterized in that, In S3, the step of inputting multimodal information into a large multimodal model for analysis and processing, generating processing results, and feeding them back to the user includes: Step S31: The multimodal large model identifies and extracts features from objects based on multimodal information; Step S32: Based on the results of recognition and feature extraction, the multimodal large model analyzes and derives relevant industrial knowledge about the target object; Step S33: The relevant industrial knowledge of the target object obtained from the analysis is used as the processing result, and the processing result is fed back to the user in the virtual reality scene.
3. The virtual reality intelligent education method based on industrial multimodal large model driving according to claim 1, characterized in that, In step S4, the process of acquiring user-inputted question information, inputting the question information into a multimodal large model for extraction and processing, and generating answer information in real time, while the multimodal large model performs intelligent error correction based on the industrial knowledge involved in the question information, including: Step S41: The user inputs a question through the microphone device provided with the system. The system uses speech recognition technology to convert the audio question into text and input it into the multimodal large model. Step S42: The multimodal large model, combined with natural language processing technology, analyzes the textual question information and extracts the key question information. Step S43: The multimodal big model makes a comprehensive judgment based on the key question information, combined with the associated multimodal information, extracted language features, fused knowledge graph and industrial knowledge, and generates answer information in real time. Step S44: The multimodal big data model identifies the industrial knowledge elements involved in the question information, establishes an intelligent error correction mechanism based on industrial knowledge, and corrects the errors in the industrial knowledge involved in the user's question information.
4. The virtual reality intelligent education method based on industrial multimodal large model driving according to claim 3, characterized in that, The establishment of an intelligent error correction mechanism based on industrial knowledge includes error correction related to product knowledge, error correction related to industrial process and technology knowledge, error correction related to industrial equipment operation and maintenance knowledge, and error correction related to industrial standards and specifications knowledge.
5. A virtual reality intelligent education system driven by an industrial multimodal large model, characterized in that: The system applied to the virtual reality intelligent education method based on industrial multimodal large model driven by claim 1, the system comprising: The target scanning module is used to scan target objects in a virtual reality scene using a scanning function; The information acquisition module is used to acquire multimodal information of the target object and encode and format the multimodal information. The analysis and processing module is used to input multimodal information into a large multimodal model for analysis and processing; The results feedback module is used to provide feedback on the generated processing results to the user. The question-and-answer error correction module is used to obtain the question information input by the user, input the question information into the multimodal big model for extraction and processing, and generate answer information in real time. At the same time, the multimodal big model performs intelligent error correction based on the industrial knowledge involved in the question information. The return decision module is used to determine whether to continue asking questions. If the user continues to ask questions, the module returns; otherwise, the system's chat function is closed.
6. A network-side server, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform the virtual reality intelligent education method based on industrial multimodal large model driven as described in any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the virtual reality intelligent education method based on industrial multimodal large model driven by any one of claims 1 to 4.
Citation Information
Patent Citations
Man-machine interaction method and device based on large model and medium
CN117992587A
Question and answer method and device based on multi-modal industrial large model
CN118761458A