Manufacturing workshop virtual-real fusion safety training method combining machine vision and large language model

Through the integrated virtual and real safety training method of machine vision and large language models, the problems of low efficiency, poor results and safety hazards in traditional training methods are solved, and efficient, accurate and safe equipment training is achieved.

CN120108240APending Publication Date: 2025-06-06ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510167385.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-15
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional manufacturing workshop equipment training methods have problems of inefficiency, poor results and safety risks, especially in training for complex equipment and high-risk operations.

Method used

The virtual and real fusion safety training method combined with machine vision and large language models is adopted, and equipment image analysis and part recognition are performed through semantic segmentation models, and natural dialogue and dynamic three-dimensional feedback are achieved in combination with large language models, improving the interactiveness and intelligence level of training.

Benefits of technology

It improves the efficiency, accuracy and user experience of workshop equipment training, reduces the risk of training, and realizes an immersive training environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108240A_ABST
    Figure CN120108240A_ABST
Patent Text Reader

Abstract

The invention provides a manufacturing workshop virtual-real fusion safety training method combining machine vision and a large language model. According to the method, a virtual-real fusion training module of production equipment in an augmented reality (AR) environment, a production equipment semantic understanding module based on machine vision, a man-machine intelligent dialogue module combining a knowledge graph and a large language model, and a training database are included. A user shoots a production scene in a manufacturing workshop through a camera of AR equipment, and equipment contained in the scene is segmented through a semantic understanding module; matching the corresponding three-dimensional model and the corresponding information from the database, loading the three-dimensional model and the corresponding information in the AR environment, and taking the matched equipment information as the input of a large language model; the trainee completes training of cognition, operation and other contents through the virtual-real fusion training module according to the prompt of the training information; in the training process, the man-machine intelligent dialogue module provides natural language questions and answers and dynamic three-dimensional feedback for trainees according to equipment information and user questions, and the training effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent workshop equipment training, and in particular to a virtual-reality fusion safety training method for a manufacturing workshop that combines machine vision with a large language model. Background Art

[0002] In order to ensure production efficiency and production safety, employees in the manufacturing workshop need to undergo rigorous operation training before actually operating the processing equipment. Traditional training methods usually rely on paper teaching materials, video demonstrations and physical operations, but these methods still have many problems in terms of efficiency, effectiveness and safety, especially in the training of complex equipment and high-risk operations. There are safety hazards. In order to make up for the shortcomings of traditional training methods, smart workshops and unmanned factories have gradually introduced advanced technical means such as augmented reality (AR) to provide manufacturing workshop employees with a more immersive and interactive training experience while ensuring safety. Through AR equipment, employees can overlay virtual information in the actual working environment and understand the equipment structure, operation process and maintenance methods more intuitively. However, traditional AR training systems rely on manually preset equipment models, are limited to displaying static content, lack natural human-machine interaction and intelligent feedback with trainers, resulting in workshop personnel after training still finding it difficult to deeply master equipment operation skills and safety points.

[0003] With the development of artificial intelligence (AI), technologies such as machine vision for scene understanding and natural human-computer dialogue through large models have gradually been combined with intelligent manufacturing to form a new generation of intelligent manufacturing technology system. The combination of AI and machine vision gives computers the ability to see and understand the world, and the combination of AI and natural language gives computers the ability to conduct generative dialogues with people. To this end, the present invention proposes a virtual-real integration safety training method for manufacturing workshops that combines machine vision with a large language model. Through machine vision, the virtual training system of the manufacturing workshop understands the photographed production equipment, as well as the properties, uses, operating difficulties, operation and maintenance difficulties of the production equipment, and the relationship with other equipment; through large language models and knowledge graphs, the virtual training system can conduct natural dialogues with trainees and answer questions raised by trainees; finally, the above technology is combined with the AR system to complete the virtual-real integration of production equipment operation and operation and maintenance safety training without turning on the production equipment. Summary of the invention

[0004] In view of the dangers of workshop equipment training and the lack of interaction and feedback with users in traditional AR training systems. The present invention provides a virtual-reality fusion safety training method for manufacturing workshops that combines machine vision with a large language model. It uses a semantic segmentation model to perform equipment image analysis and part recognition, and can automatically segment equipment and parts and match them with three-dimensional models. At the same time, through precise alignment with physical entities, it provides a more natural virtual-reality combination effect, and completes production equipment operation and maintenance safety training without turning on the production equipment. In addition, the system is also combined with a large language model to allow users to conduct interactive questions and answers on safety training knowledge through AR devices, further improving the interactivity and intelligence of training. Through the present invention, the efficiency, accuracy and user experience of workshop equipment training can be greatly improved, while reducing the danger of industrial workshop training.

[0005] The technical solution of the present invention is implemented as follows:

[0006] A virtual-reality fusion safety training method for manufacturing workshops combining machine vision and a large language model, the specific steps are as follows:

[0007] S1. Create a training database containing basic information of workshop equipment, equipment 3D models and training knowledge for each module to call.

[0008] S2. Establish a machine vision-based production equipment semantic understanding module to segment the semantic information contained in the equipment and part images used in training;

[0009] S3. Build a human-machine intelligent dialogue module based on the knowledge graph of RAG architecture and the large language model to answer the questions users encounter in training in real time and provide dynamic three-dimensional feedback in the augmented reality environment;

[0010] S4. Create an augmented reality virtual training system that includes the cognitive, operational, and maintenance scenarios required for equipment training, and implement image transmission and voice interactive operations.

[0011] Furthermore, step S2 is specifically as follows:

[0012] S21. Take images of equipment and parts and create annotated datasets

[0013] First, use photography equipment to take photos of the equipment and its parts, buttons, and nameplates that need to be recognized under different lighting environments and from different angles; to ensure the training effect, all equipment and parts that need to be trained must be included, and the samples of different equipment and parts images must be kept uniform. Secondly, use labelme software to mark the polygonal outlines of the equipment and parts that need to be recognized on the picture, give the corresponding label name, and generate a json file with the same name as the picture; convert the json file into a txt format that YOLOV11 can recognize, so as to create a data set for training.

[0014] S22. Train the semantic segmentation model and integrate it into the AR training system

[0015] First, divide the training set, validation set, and test set, and define the key parameters of model training in the .yaml file of the dataset: number of categories, category name, training set, validation set, and test set path. Train the YOLOV11 semantic segmentation network to obtain a semantic segmentation model that can segment images taken by AR devices. Secondly, use the trained semantic segmentation model in the workshop image taken by the AR device to segment the production equipment, process equipment, transportation equipment and other equipment objects involved in the image, and superimpose different colors on the original image, while displaying the corresponding confidence box and label name, and then output the label name in the form of a string and transmit it to the training system in the AR device, so that the virtual safety training system can match the corresponding label of the equipment 3D model and basic equipment information.

[0016] Furthermore, step S3 is specifically as follows:

[0017] S31. Build a safety training knowledge base in the form of text and knowledge graph

[0018] Connect the big language model with the safety training knowledge base, and limit the format of the knowledge base to text files and knowledge graphs for easy retrieval by the big language model; convert text content such as precautions before training, equipment information (such as name, model, power consumption, origin, etc.), equipment operation and use methods, maintenance methods, etc. into unstructured text files; convert the relationship between equipment parts, the relationship between equipment, operation steps, etc. into structured neo4j knowledge graphs. For example Figure 3 , express the composition of a robotic arm workbench in the form of a structured neo4j knowledge graph, define the robotic arm workbench and the equipment placed on it as nodes with different attributes, and establish a relationship from the robotic arm workbench to each device; and so on, divide the hierarchical composition of the equipment and convert it into the relationship between nodes, so that the large language model can retrieve related content through nodes and relationships.

[0019] S32. Build a security training question-answering system and integrate text knowledge base and knowledge graph

[0020] 1) The training-related knowledge base and knowledge graph are used as the external knowledge source of the large language model. The architecture diagram is as follows: Figure 4 ; The text files are vectorized and stored in the knowledge base, and the relationships between devices and between devices and parts are stored in the graph database using the neo4j knowledge graph; when answering questions, the user's questions are vectorized, and the answers with the highest similarity are retrieved from the vectorized knowledge base and graph database and output; the text is vectorized using an embedding model, and the cosine similarity is used to calculate the similarity between document blocks. The similarity formula is:

[0021] Let the document vector be v 1 and v 2 , and its cosine similarity is:

[0022]

[0023] Where: v 1 ·v 2 is the dot product of the document vector, ||v 1 || and ||v 2 || are the modulus of the vector (i.e. the length of the vector). This formula can be used to retrieve the most similar documents from the vectorized knowledge base.

[0024] 2) For text data, first process the text documents in the database, load the text documents from the specified folder, and use LangChain's TextLoader to load; use the text segmenter CharacterTextSplitter to divide the text into blocks, and set appropriate chunk_sizeh and chunk_overlap to maintain the text continuity between blocks. After the segmentation is completed, use the embedding model to generate the corresponding vector embedding for each block and store it in the vector database; after the text document is processed, define a retriever component to provide additional context information based on the user's query and the semantic similarity between the embedded blocks; create a prompt template to specify how the large language model uses the retrieved context to answer questions, for example: "Use the retrieved text fragment to answer the question, if you don't know the answer, answer that you don't know", so as to avoid the large language model not answering questions based on the document; finally, build a RAG process chain to connect the retriever, prompt template and the specified LLM. After defining the RAG chain, you can call it for generation.

[0025] 3) For the knowledge graph, first initialize the Neo4j graph database and specify the large language model. Through the GraphCypherQAChain function of LangChain, combine the Neo4j graph database and the large language model (LLM) to allow the graph database to be queried through natural language. When the user asks a question, the natural language question is converted into a Cypher statement that Neo4j can recognize for query. After the query is completed, the results retrieved from the graph database are returned.

[0026] S33. Create an intelligent agent and integrate it with the AR training system to achieve keyword extraction and dynamic 3D feedback

[0027] Use Langchain to create an intelligent agent, and create a tool through LLM Agent. The large language model and the code that meets the corresponding functions are embedded in the tool. While outputting the answer to the user, the answer given by the large language model is extracted by keywords, and the keywords are transmitted to the training system in AR through the TCP protocol, matching the corresponding dynamic three-dimensional feedback, such as the highlighted flashing of the button, the arrow guidance of the required operation position, and the model flashing of the equipment composition cognition. For example: set the keyword extraction rule to "device name + step + object". When the user asks: "How to turn on the teach pendant", the answer given by the large language model is: "You need to press the red button to turn on the teach pendant". At the same time, LLM Agent will extract the keyword "teach pendant + power on + button" according to the answer and rules, and transmit this information to the AR training system, matching the corresponding dynamic three-dimensional arrow feedback pointing to the red button of the teach pendant, guiding the user to complete the training further.

[0028] Further, step S4 is specifically as follows:

[0029] S41. Use Unity3D to create cognitive, operational, and maintenance scenarios required for training;

[0030] Use 3D modeling software (such as Solidworks, Creo) to create 3D models of the equipment and parts included in the training. To ensure the fusion effect of the virtual 3D model and the physical entity, the modeling accuracy should retain its characteristics as much as possible. Import the built 3D model into Blender for format conversion, convert it into .FBX format, and import it into Unity3D, add buttons, nameplates and other content corresponding to the equipment to the model; Create a Canve panel and add image components and TextMeshPro components to it. The image component is used to preview the photos taken by the AR device and the result image after the semantic segmentation model segmented the image. The TextMeshPro component is used to receive the label name returned in the semantic segmentation image result. The system displays the corresponding 3D model according to the returned label name; At the same time, add the gesture interaction function in the MRTK mixed reality development kit, and write C# code to implement the corresponding functions, such as the on and off of the device button indicator light, the disassembly and assembly of the device, and the start and stop of the device; Create an Amination animation to reflect the composition and operation of the equipment, so that users can interact with the model and the buttons on the model; Through the above process, the effect of controlling and recognizing the device in an augmented reality environment can be achieved.

[0031] S42. Use Vuforia virtual-reality fusion method to achieve the fusion of virtual 3D models and real physical entities in augmented reality environment;

[0032] Use Vuforia's Model Target Generator (MTG) software to import the CAD 3D model of the equipment or parts designed with the 3D modeling software, create a virtual object of the Model Target in the software, specify the advanced model target database, select the model length, width, and height fields to ensure that they are consistent with the size of the real physical model, specify the material and color of the model, set the guide view to the Digital Eyewear corresponding to the AR glasses, and select the recognition range as needed, such as 360°, 180°, etc.; train the model on the Vuforia cloud, save the trained model on the cloud and export it in .unitypackage format.

[0033] Import the Vuforia SDK and the trained .unitypackage file into the augmented reality development environment in Unity3D, overlap the trained model outline with the 3D CAD model, select the preview outline to display in the augmented reality environment, set the trigger conditions, and deploy it to the AR device when the model outline coincides with the physical entity outline. By comparing the model preview outline with the real physical entity outline, further guide the user to complete the virtual-real fusion. After the comparison is successful, the 3D CAD model of the device appears and coincides with its physical entity. The user can operate this model that coincides with the real physical entity in the augmented reality environment to complete its recognition, operation, disassembly and assembly tasks.

[0034] The virtual-real fusion method is a feature detection method, and its detection formula is as follows:

[0035] Let the error of the i-th pair of points be e i , then e i =p i -(Rp' i +t), where p a is a point on the 3D model in the augmented reality environment, p i is a point on a real physical entity, R is the rotation matrix, and t is the transfer vector (x 1 ,y 1 ,z 1 ), generate a model with contours through MTG, match the extracted points with the real points, so as to realize the superposition of 3D CAD model and real physical entity, that is, virtual-real fusion; the superposition formula is:

[0036]

[0037] When virtual and real are integrated, the system will call the spatial information of the AR device to continuously track the model, allowing the AR device to maintain virtual and real integration from multiple angles.

[0038] S43. Create interactive panels and dynamic three-dimensional feedback to achieve interaction between various scenes in the augmented reality virtual training system and the safety training knowledge question and answer prediction model.

[0039] First, create an interactive panel: open Unity3D, check the Microphone option in Capabilities in the Player configuration, add a Canvas canvas to the scene, add the Mixed Reality InputModule component in MRTK to the Main Camera to call the MixedRealityKeyboard component that provides methods for starting and closing the system keyboard and interacting with text entered through the keyboard. This component can call the ShowKeyboard(string text="",bool multiLine=false)HideKeyboard() function to show and hide the keyboard, and handle the OnCommitText, OnShowKeyboard, OnHideKeyboard keyboard display, hiding, and Enter key events to be processed; create two TextMeshPro components in the scene, respectively for receiving user questions and answers from the safety training knowledge question and answer large language model; add the TMP_KeyboardInputField component in the MixedRealityKeyboard component to the TextMeshPro component used to receive user questions.

[0040] Secondly, construct dynamic three-dimensional feedback:

[0041] Dynamic 3D feedback includes model flashing, highlighted buttons on the model, and arrows pointing to the model. In Unity3D, scripts with different functions are added to the model to achieve dynamic 3D feedback. For model flashing, the transparency of the material is modified, the flashing time interval is controlled, and the alpha value of the material is changed to help users locate key components and their locations. For model button highlighting, C# code is used to control the material of GameObject to achieve it, and the highlight change of the button material is used as dynamic 3D feedback during training. The arrow pointing to the model uses the DirectionalIndicator Class direction solver, starting from the reference point of SolverHandler Tracked Target and pointing to the DirectionalTarget target point. Assume that the position of the target point is T(x T ,y T ,z T ), the arrow points to the position P(x P ,y P ,z P ), the direction vector can be calculated as:

[0042]

[0043] Then, the direction indicator SolverHandler is used to calculate the direction to the target. When this object is not in the view, the indicator is pointed to the location of the model to help users find the location of the equipment or parts. At the same time, C# code is added to each model. Dynamic three-dimensional feedback will only be displayed when the keywords transmitted by LLM Agent to the AR virtual training system are received.

[0044] After completing the above content, deploy the scene to the AR device. Users can use the voice input function on the system keyboard to convert voice into text, thereby improving the efficiency of users asking questions to the large language model. After receiving the question, the large language model will search the training knowledge base and knowledge graph, and then transmit the answer to the TextMeshPro component used to receive the answer, and display it on the interactive panel in the AR environment. At the same time, the keywords corresponding to the answer given by the large language model will also match the corresponding dynamic three-dimensional feedback and be displayed in the virtual training system on the AR device, guiding users to further complete the training and master the training content.

[0045] Beneficial effects of the present invention:

[0046] The virtual-reality fusion safety training method for manufacturing workshops that combines machine vision with a large language model can automatically segment equipment and parts and match them with three-dimensional models through semantic segmentation models to allow users to further understand the equipment; at the same time, through precise alignment with physical entities, it provides a more natural virtual-reality combination effect, allowing users to control equipment through gestures, providing users with an immersive training environment. In addition, by combining with a large language model, users are allowed to conduct interactive Q&A on safety training knowledge through AR devices, and dynamic three-dimensional feedback is given in the AR environment, further improving the interactivity and intelligence of the training. The present invention can reduce the manpower input in workshop equipment training and the danger of industrial workshop training, while improving training efficiency and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 : Overall flow chart of the present invention.

[0048] Figure 2 : System structure diagram of the present invention.

[0049] Figure 3 : Schematic diagram of the knowledge graph of safety training knowledge in an embodiment of the present invention.

[0050] Figure 4 : The architecture diagram of the safety training knowledge question and answer prediction model of the present invention. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] like Figure 1 As shown, the example of the present invention provides a virtual-reality fusion safety training method for a manufacturing workshop combining machine vision with a large language model, and the specific steps are as follows:

[0053] S1. Create a training database containing basic information of workshop equipment, equipment 3D models and training knowledge for each module to call.

[0054] S2. Establish a machine vision-based production equipment semantic understanding module to segment the semantic information contained in the equipment and part images used in training;

[0055] S21. Take images of equipment and parts and create annotated datasets

[0056] First, use photography equipment to take photos of the equipment and its parts, buttons, and nameplates that need to be recognized under different lighting environments and from different angles; to ensure the training effect, all the equipment and parts that need to be trained must be included, and the samples of different equipment and parts images must be kept uniform. Secondly, use labelme software to mark the polygonal outlines of the equipment and parts that need to be recognized on the picture, give the corresponding label name, and generate a json file with the same name as the picture; convert the json file into a txt format that YOLOV11 can recognize, so as to create a data set for training.

[0057] S22. Train the semantic segmentation model and integrate it into the AR training system

[0058] First, divide the training set, validation set, and test set, and define the key parameters of model training in the .yaml file of the dataset: number of categories, category name, training set, validation set, and test set path. Train the YOLOV11 semantic segmentation network to obtain a semantic segmentation model that can segment images taken by AR devices. Secondly, use the trained semantic segmentation model in the workshop image taken by the AR device to segment the production equipment, process equipment, transportation equipment and other equipment objects involved in the image, and superimpose different colors on the original image, while displaying the corresponding confidence box and label name, and then output the label name in the form of a string and transmit it to the training system in the AR device, so that the virtual safety training system can match the corresponding label of the equipment 3D model and basic equipment information.

[0059] S3. Build a human-machine intelligent dialogue module based on the knowledge graph of RAG architecture and the large language model to answer the questions users encounter in training in real time and provide dynamic three-dimensional feedback in the augmented reality environment;

[0060] S31. Build a safety training knowledge base in the form of text and knowledge graph

[0061] Connect the big language model with the safety training knowledge base, and limit the format of the knowledge base to text files and knowledge graphs for easy retrieval by the big language model; convert text content such as precautions before training, equipment information (such as name, model, power consumption, origin, etc.), equipment operation and use methods, maintenance methods, etc. into unstructured text files; convert the relationship between equipment parts, the relationship between equipment, operation steps, etc. into structured neo4j knowledge graphs. For example Figure 3 , express the composition of a robotic arm workbench in the form of a structured neo4j knowledge graph, define the robotic arm workbench and the equipment placed on it as nodes with different attributes, and establish a relationship from the robotic arm workbench to each device; and so on, divide the hierarchical composition of the equipment and convert it into the relationship between nodes, so that the large language model can retrieve related content through nodes and relationships.

[0062] S32. Build a security training question-answering system and integrate text knowledge base and knowledge graph

[0063] 1) The training-related knowledge base and knowledge graph are used as the external knowledge source of the large language model. The principle diagram is as follows Figure 4 ; The text files are vectorized and stored in the knowledge base, and the relationships between devices and between devices and parts are stored in the graph database using the neo4j knowledge graph; when answering questions, the user's questions are vectorized, and the answers with the highest similarity are retrieved from the vectorized knowledge base and graph database and output; the text is vectorized using an embedding model, and the cosine similarity is used to calculate the similarity between document blocks. The similarity formula is:

[0064] Let the document vector be v 1 and v 2 , and its cosine similarity is:

[0065]

[0066] Where: v 1 ·v 2 is the dot product of the document vector, ||v 1 || and ||v 2 || are the modulus of the vector (i.e. the length of the vector). This formula can be used to retrieve the most similar documents from the vectorized knowledge base.

[0067] 2) For text data, first process the text documents in the database, load the text documents from the specified folder, and use LangChain's TextLoader to load; use the text segmenter CharacterTextSplitter to divide the text into blocks, and set appropriate chunk_sizeh and chunk_overlap to maintain the text continuity between blocks. After the segmentation is completed, use the embedding model to generate the corresponding vector embedding for each block and store it in the vector database; after the text document is processed, define a retriever component to provide additional context information based on the user's query and the semantic similarity between the embedded blocks; create a prompt template to specify how the large language model uses the retrieved context to answer questions, for example: "Use the retrieved text fragment to answer the question, if you don't know the answer, answer that you don't know", so as to avoid the large language model not answering questions based on the document; finally, build a RAG process chain to connect the retriever, prompt template and the specified LLM. After defining the RAG chain, you can call it for generation.

[0068] 3) For the knowledge graph, first initialize the Neo4j graph database and specify the large language model. Through the GraphCypherQAChain function of LangChain, combine the Neo4j graph database and the large language model (LLM) to allow the graph database to be queried through natural language. When the user asks a question, the natural language question is converted into a Cypher statement that Neo4j can recognize for query. After the query is completed, the results retrieved from the graph database are returned.

[0069] S33. Create an intelligent agent and integrate it with the AR training system to achieve keyword extraction and dynamic 3D feedback

[0070] Use Langchain to create an intelligent agent, and create a tool through LLM Agent. The large language model and the code that meets the corresponding functions are embedded in the tool. While outputting the answer to the user, the answer given by the large language model is extracted by keywords, and the keywords are transmitted to the training system in AR through the TCP protocol, matching the corresponding dynamic three-dimensional feedback, such as the highlighted flashing of the button, the arrow guidance of the required operation position, and the model flashing of the equipment composition cognition. For example: set the keyword extraction rule to "device name + step + object". When the user asks: "How to turn on the teach pendant", the answer given by the large language model is: "You need to press the red button to turn on the teach pendant". At the same time, LLM Agent will extract the keyword "teach pendant + power on + button" according to the answer and rules, and transmit this information to the AR training system, matching the corresponding dynamic three-dimensional arrow feedback pointing to the red button of the teach pendant, guiding the user to complete the training further.

[0071] S4. Create an augmented reality virtual training system that includes the cognitive, operational, and maintenance scenarios required for equipment training, and implement image transmission and voice interactive operations.

[0072] The specific implementation steps of S4 are:

[0073] S41. Use Unity3D to create cognitive, operational, and maintenance scenarios required for training;

[0074] Use 3D modeling software (such as Solidworks, Creo) to create 3D models of the equipment and parts included in the training. To ensure the fusion effect of the virtual 3D model and the physical entity, the modeling accuracy should retain its characteristics as much as possible. Import the built 3D model into Blender for format conversion, convert it into .FBX format, and import it into Unity3D, add buttons, nameplates and other content corresponding to the equipment to the model; Create a Canve panel and add image components and TextMeshPro components to it. The image component is used to preview the photos taken by the AR device and the result image after the semantic segmentation model segmented the image. The TextMeshPro component is used to receive the label name returned in the semantic segmentation image result. The system displays the corresponding 3D model according to the returned label name; At the same time, add the gesture interaction function in the MRTK mixed reality development kit, and write C# code to implement the corresponding functions, such as the on and off of the device button indicator light, the disassembly and assembly of the device, and the start and stop of the device; Create an Amination animation to reflect the composition and operation of the equipment, so that users can interact with the model and the buttons on the model; Through the above process, the effect of controlling and recognizing the device in an augmented reality environment can be achieved.

[0075] S42. Use Vuforia virtual-reality fusion method to achieve the fusion of virtual 3D models and real physical entities in augmented reality environment;

[0076] Use Vuforia's Model Target Generator (MTG) software to import the CAD 3D model of the equipment or parts designed with the 3D modeling software, create a virtual object of the Model Target in the software, specify the advanced model target database, select the model length, width, and height fields to ensure that they are consistent with the size of the real physical model, specify the material and color of the model, set the guide view to the Digital Eyewear corresponding to the AR glasses, and select the recognition range as needed, such as 360°, 180°, etc.; train the model on the Vuforia cloud, save the trained model on the cloud and export it in .unitypackage format.

[0077] Import the Vuforia SDK and the trained .unitypackage file into the augmented reality development environment in Unity3D, overlap the trained model outline with the 3D CAD model, select the preview outline to display in the augmented reality environment, set the trigger conditions, and deploy it to the AR device when the model outline coincides with the physical entity outline. By comparing the model preview outline with the real physical entity outline, further guide the user to complete the virtual-real fusion. After the comparison is successful, the 3D CAD model of the device appears and coincides with its physical entity. The user can operate this model that coincides with the real physical entity in the augmented reality environment to complete its recognition, operation, disassembly and assembly tasks.

[0078] The virtual-real fusion method is a feature detection method, and its detection formula is as follows:

[0079] Let the error of the i-th pair of points be e i , then e i =p i -(Rp' i +t), where p a is a point on the 3D model in the augmented reality environment, p i is a point on a real physical entity, R is the rotation matrix, and t is the transfer vector (x 1 ,y 1 ,z 1 ), generate a model with contours through MTG, match the extracted points with the real points, so as to realize the superposition of 3D CAD model and real physical entity, that is, virtual-real fusion; the superposition formula is:

[0080]

[0081] When virtual and real are integrated, the system will call the spatial information of the AR device to continuously track the model, allowing the AR device to maintain virtual and real integration from multiple angles.

[0082] S43. Create interactive panels and dynamic three-dimensional feedback to achieve interaction between various scenes in the augmented reality virtual training system and the safety training knowledge question and answer prediction model.

[0083] First, create an interactive panel: open Unity3D, check the Microphone option in Capabilities in the Player configuration, add a Canvas canvas to the scene, add the Mixed Reality InputModule component in MRTK to the Main Camera to call the MixedRealityKeyboard component that provides methods for starting and closing the system keyboard and interacting with text entered through the keyboard. This component can call the ShowKeyboard(string text="",bool multiLine=false)HideKeyboard() function to show and hide the keyboard, and handle the OnCommitText, OnShowKeyboard, OnHideKeyboard keyboard display, hiding, and Enter key events to be processed; create two TextMeshPro components in the scene, respectively for receiving user questions and answers from the safety training knowledge question and answer large language model; add the TMP_KeyboardInputField component in the MixedRealityKeyboard component to the TextMeshPro component used to receive user questions.

[0084] Secondly, construct dynamic three-dimensional feedback:

[0085] Dynamic 3D feedback includes model flashing, highlighted buttons on the model, and arrows pointing to the model. In Unity3D, scripts with different functions are added to the model to achieve dynamic 3D feedback. For model flashing, the transparency of the material is modified, the flashing time interval is controlled, and the alpha value of the material is changed to help users locate key components and their locations. For model button highlighting, C# code is used to control the material of GameObject to achieve it, and the highlight change of the button material is used as dynamic 3D feedback during training. The arrow pointing to the model uses the DirectionalIndicator Class direction solver, starting from the reference point of SolverHandler Tracked Target and pointing to the DirectionalTarget target point. Assume that the position of the target point is T(x T ,y T ,z T ), the arrow points to the position P(x P ,y P ,z P ), the direction vector can be calculated as:

[0086]

[0087] Then, the direction indicator SolverHandler is used to calculate the direction pointing to the target. When this object is not in the view, the indicator points to the location of the model to help users find the location of the equipment or parts. At the same time, C# code is added to each model. Dynamic three-dimensional feedback will only be displayed when the keywords transmitted by LLM Agent to the AR virtual training system are received.

[0088] After completing the above content, deploy the scene to the AR device. Users can use the voice input function on the system keyboard to convert voice into text, thereby improving the efficiency of users asking questions to the large language model. After receiving the question, the large language model searches the training knowledge base and knowledge graph, transmits the answer to the TextMeshPro component used to receive the answer, and displays it on the interactive panel in the AR environment. At the same time, the keywords corresponding to the answer given by the large language model will also match the corresponding dynamic three-dimensional feedback and be displayed in the virtual training system on the AR device, guiding users to further complete the training and master the training content.

[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

[0090] The above description is only a further embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical solutions and concepts of the present invention within the scope disclosed by the present invention, which belong to the protection scope of the present invention.

Claims

1. A virtual-reality fusion safety training method for manufacturing workshops combining machine vision and a large language model, characterized in that: The specific steps are as follows: S1. Create a training database containing basic information of workshop equipment, 3D models of equipment and training knowledge for each module to call; S2. Establish a machine vision-based production equipment semantic understanding module to segment the semantic information contained in the equipment and part images used in training; S3. Build a human-machine intelligent dialogue module based on the knowledge graph of RAG architecture and the large language model to answer the questions users encounter in training in real time and provide dynamic three-dimensional feedback in the augmented reality environment; S4. Create an augmented reality virtual training system that includes the cognitive, operational, and maintenance scenarios required for equipment training, and implement image transmission and voice interactive operations.

2. The virtual-reality fusion safety training method for manufacturing workshops combining machine vision and a large language model as described in claim 1 is characterized in that: Step S2 is specifically as follows: S21. Take images of equipment and parts and create annotated datasets; S22. Train the semantic segmentation model and integrate it into the AR training system.

3. The virtual-reality fusion safety training method for manufacturing workshops combining machine vision and a large language model as described in claim 1 is characterized in that: Step S3 is specifically as follows: S31. Build a safety training knowledge base in the form of text and knowledge graph; S32. Build a security training question-answering system and integrate the text knowledge base and knowledge graph; S33. Create an intelligent agent and integrate it with the AR training system to achieve keyword extraction and dynamic three-dimensional feedback.

4. The virtual-reality fusion safety training method for manufacturing workshops combining machine vision and a large language model as described in claim 1 is characterized in that: Step S4 is specifically as follows: S41. Use Unity3D to create cognitive, operational, and maintenance scenarios required for training; S42. Use Vuforia virtual-reality fusion method to achieve the fusion of virtual 3D models and real physical entities in augmented reality environment; S43. Create interactive panels and dynamic three-dimensional feedback to achieve interaction between various scenes in the augmented reality virtual training system and the safety training knowledge question and answer prediction model.

Citation Information

Cited By

  • Power plant equipment three-dimensional interactive training system based on structured knowledge base

    CN120912392A