Multi-modal data fusion VR scene automatic modeling method for electric power operation teaching

Through the multimodal data fusion VR scene automatic modeling method, the problem of insufficient multi-source data fusion in the existing VR modeling system is solved, and a three-dimensional VR scene that is closer to the real power operation environment is generated, realizing efficient and accurate immersive learning of teaching content.

CN120689519AInactive Publication Date: 2025-09-23蒋子晴
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510810608.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-23
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120689519A_ABST
    Figure CN120689519A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric power work teaching, in particular to an electric power work teaching-oriented multi-modal data fusion VR scene automatic modeling method, which can efficiently and accurately extract key semantic information related to electric power work from multi-modal data through a deep learning-driven semantic recognition and label generation method, so as to improve the teaching efficiency of the electric power work. Multi-modal fusion is carried out by introducing a graph neural network structure, the problems of data splitting, semantic missing, scene stiffness and the like in a traditional modeling mode are effectively solved, comprehensive transition from low-level perception to high-level understanding is achieved, a generated three-dimensional VR scene model is closer to a real electric power operation environment, and teaching efficiency is improved by automatically generating and embedding teaching scripts. According to the system, the virtual teaching content is highly consistent with the actual operation process, the problems of subjectivity and inconsistency in the traditional manual script input process are effectively avoided, and the visualization effect and the real-time response capability of a teaching scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric power operation teaching, and in particular to a multimodal data fusion VR scene automatic modeling method for electric power operation teaching. Background Art

[0002] With the development of intelligent power industry, the requirements for practical training for power workers are increasing. Traditional teaching methods such as paper textbooks, video explanations or on-site training have problems such as high resource costs, high risks and poor situational immersion. In recent years, virtual reality (VR) technology has been introduced into power education. By constructing three-dimensional interactive scenes, it can achieve low-cost and high-fidelity virtual training.

[0003] At present, existing VR modeling systems often lack the ability to deeply integrate and process multi-source data such as on-site images, voice, text, and laser point clouds, resulting in large differences between the constructed virtual scenes and the real power operation environment, thus affecting the accuracy of teaching. To solve the above problems, we propose a multimodal data fusion VR scene automatic modeling method for power operation teaching. Summary of the Invention

[0004] Technical problem to be solved: Existing VR modeling systems often lack the ability to deeply integrate and process multi-source data such as on-site images, voice, text, laser point clouds, etc., resulting in large differences between the constructed virtual scenes and the actual power operation environment, thus affecting the accuracy of teaching.

[0005] In response to the shortcomings of the existing technology, the present invention provides a multimodal data fusion VR scene automatic modeling method for power operation teaching, thereby solving the technical problems mentioned in the background technology.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0007] Includes the following:

[0008] S1: Data acquisition: multi-modal data collection at the power operation site to obtain raw data such as 3D point clouds, images, audio, and text;

[0009] S2: Data preprocessing, denoising, alignment, time synchronization and coordinate unification of the collected multimodal data;

[0010] S3: Semantic recognition and label generation: Based on deep learning models, it analyzes images, speech, and text to extract semantic labels such as device type, operation process, and operation scenario.

[0011] S4: Multimodal fusion and modeling, using graph neural network structures to fuse different modal information, and automatically generate 3D VR scene models based on the combined modal information data, and configure interactive components according to teaching needs;

[0012] S5: Teaching script generation and embedding: Generate standardized teaching scripts based on speech and text data analysis and embed them into VR scenes for training and evaluation;

[0013] S6: Real-time rendering and deployment: Rendering and deploying VR scenes on teaching terminals to achieve immersive operation guidance and practical training assessment.

[0014] In one possible implementation, the data acquisition uses lidar, RGB-D camera, and voice recorder to collect high-precision three-dimensional point cloud data in the operating environment, obtain geometric information, image and depth information of power equipment and spatial structure, and realize visual perception and cognition of equipment surface features and spatial position. It also records the voice explanations, operating instructions and environmental sounds of on-site personnel, performs time synchronization and coordinate unification, and collaboratively completes the original multimodal data acquisition of on-site images, sounds and spatial structures.

[0015] In one possible implementation, the data preprocessing uses a filtering algorithm to remove background stray points in the lidar point cloud, lighting interference in the image, environmental noise in the voice signal, etc. for the collected data. For spatial data, through coordinate transformation and calibration operations, the point cloud and image data are uniformly mapped to a unified coordinate system to achieve spatial alignment.

[0016] In one possible implementation, the semantic recognition and label generation identifies the structural features of the data after data preprocessing, and parses the content of the job document based on the pre-trained language model to identify semantic elements such as job processes and operation steps, and finally unifies the recognition results into structured label data and binds them to specific three-dimensional coordinates or time segments.

[0017] In one possible implementation, after completing the semantic recognition and label generation of multimodal data, the multimodal fusion and modeling uses a graph neural network to deeply fuse different modal information such as images, point clouds, text, and voice, and construct a graph structure with devices as nodes and spatial and semantic relationships as edges. The image recognition results, point cloud geometry, voice / text instructions, etc. are input into the fusion model as node attributes to learn the contextual dependencies and semantic commonalities between the modalities. The fused structural information is input into the modeling engine to automatically generate a three-dimensional VR scene model with semantic attributes.

[0018] In one possible implementation, the teaching script is generated and embedded. After completing the multimodal fusion and modeling of multimodal data, the system uses natural language processing technology (NLP) to extract and structure key information such as work processes, operating instructions, and precautions based on the collected voice explanation content and text materials. Through task decomposition and time sequence arrangement, a standardized teaching script containing step instructions, operating prompts, standard language, assessment points, etc. is automatically generated.

[0019] In one possible implementation, the real-time rendering and deployment utilizes a real-time rendering engine to perform graphics rendering processing on the modeling of the three-dimensional VR scene and the embedding of the teaching script, including light and shadow simulation, material mapping, dynamic interactive response, etc., to ensure the visual realism and operational smoothness of the VR environment.

[0020] Beneficial effects compared with existing technologies:

[0021] 1. This solution introduces multiple sensors to enable the simultaneous collection of heterogeneous information. This not only captures the geometric shape and spatial structure of power operation scenes, but also captures the operator's verbal expressions, environmental noise, and image features of key equipment. Systematically preprocessing multimodal data significantly improves data availability and fusion quality. De-noising effectively eliminates redundant and interfering information, ensuring data clarity and accuracy. Time synchronization ensures content consistency across modalities and avoids information misalignment. Coordinate unification enables the coordinated positioning of multi-source spatial data, providing a unified spatial foundation for subsequent fusion modeling.

[0022] 2. In this solution, through deep learning-driven semantic recognition and label generation methods, key semantic information related to power operations can be efficiently and accurately extracted from multimodal data. Compared with traditional manual annotation or single-modal analysis methods, the automation and coverage of semantic extraction are significantly improved. Not only is the recognition accuracy higher, but it can also form a structured and reusable label set. By introducing a graph neural network structure for multimodal fusion, it effectively overcomes the problems of "data fragmentation," "semantic loss," and "sclerotic scenes" in traditional modeling methods, achieving a comprehensive transition from "low-level perception" to "high-level understanding," making the generated three-dimensional VR scene model closer to the real power operation environment.

[0023] 3. In this solution, by automatically generating and embedding teaching scripts, the system achieves a high degree of consistency between virtual teaching content and actual work processes, effectively avoiding the subjectivity and inconsistency problems in the traditional manual script entry process. Through real-time rendering and deployment mechanisms, the system can quickly transform the teaching scenes generated by multimodal fusion into an interactive and experiential immersive learning platform, effectively improving the visualization effect and real-time response capabilities of the teaching scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention with reference to the accompanying drawings.

[0025] Figure 1 Schematic diagram of the method structure of the present invention. DETAILED DESCRIPTION

[0026] In the description of the present invention, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "axial", "radial", "circumferential", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention.

[0027] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature identified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0028] In the present invention, unless otherwise clearly specified and limited, terms such as "install", "connect", "connect", and "fix" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integrated connection; it can be a mechanical connection, an electrical connection, or a communication connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.

[0029] In the present invention, unless otherwise clearly specified and limited, a first feature "above" or "below" a second feature may be such that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. In the description of this specification, the descriptions with reference to the terms "one scheme", "some schemes", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials or characteristics described in conjunction with the scheme or example are included in at least one scheme or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same scheme or example. Moreover, the specific features, structures, materials or characteristics described may be combined in an appropriate manner in any one or more schemes or examples;

[0030] In order to more clearly and completely illustrate the technical solution of the present invention, the present invention is further described below with reference to the accompanying drawings:

[0031] Example 1

[0032] Please refer to Figure 1 As shown, this embodiment introduces a multimodal data fusion VR scene automatic modeling method for power operation teaching, including the following contents:

[0033] S1: Data Collection: Using LiDAR, RGB-D cameras, voice recorders and other equipment to collect multimodal data at the power operation site, obtaining raw data such as 3D point clouds, images, audio, and text;

[0034] S1.1: Use LiDAR to collect high-precision 3D point cloud data in the operating environment to obtain geometric information of power equipment and spatial structures;

[0035] S1.2: By using an RGB-D camera, both image and depth information can be collected simultaneously, enabling visual perception of the device's surface features and spatial position.

[0036] S1.3: Use a voice recorder to record on-site personnel's voice explanations, operating instructions, and environmental sounds. Through time synchronization and coordinate unification, collaborate to complete the acquisition of original multimodal data of on-site images, sounds, and spatial structures.

[0037] By introducing multiple sensors to achieve synchronous collection of heterogeneous information, it is possible to not only obtain the geometric shape and spatial structure of the power operation scene, but also capture the operator's language expression, environmental noise, and image characteristics of key equipment. Compared with traditional single image or text collection methods, this method greatly improves the richness and multidimensionality of data, providing a solid foundation for subsequent multimodal data fusion and semantic modeling. At the same time, through the automated collection of raw data, the workload of manual intervention is reduced, the efficiency and accuracy of modeling are improved, and the virtual teaching scene is ensured to be more realistic and practical.

[0038] S2: Data preprocessing: denoising, alignment, time synchronization and coordinate unification of the collected multimodal data;

[0039] S2.1: After completing multimodal data acquisition, the raw data is first denoised using filtering algorithms to remove background stray points in the lidar point cloud, lighting interference in the image, and environmental noise in the voice signal.

[0040] S2.2: For different modal data such as images, point clouds, voice, and text, time synchronization is performed using timestamp matching to ensure that all types of information can be aligned under the same timeline;

[0041] S2.3: For spatial data, coordinate transformation and calibration operations are used to uniformly map point cloud and image data into a unified coordinate system to achieve spatial alignment.

[0042] Through systematic preprocessing of multimodal data, data availability and fusion quality are significantly improved. Denoising effectively removes redundant and interfering information, ensuring data clarity and accuracy. Time synchronization ensures content consistency across modalities, avoiding information dislocation. Coordinate unification enables the collaborative positioning of multi-source spatial data, providing a unified spatial foundation for subsequent fusion modeling. This overall preprocessing process not only improves the system's adaptability to multi-source data but also enhances the stability and robustness of model training and scene reconstruction, helping to build more realistic, complete, and semantically consistent VR teaching scenarios.

[0043] S3: Semantic Recognition and Label Generation: This process uses deep learning models to identify and analyze images, speech, and text, extracting semantic labels such as device type, operating procedures, and work scenarios. After preprocessing multimodal data, the trained deep learning model is used to perform joint semantic recognition and analysis on image, speech, and text data.

[0044] S3.1: The image recognition module uses a convolutional neural network model to identify the appearance and structural features of equipment in images and RGB-D images, annotating equipment names and locations such as "high-voltage circuit breaker" and "lightning arrester." The speech recognition module uses speech recognition and natural language understanding models to convert on-site speech into text and extract key instructions and operational terms. The text analysis module parses the content of job documents based on a pre-trained language model, identifying semantic elements such as job processes and operational steps. Ultimately, the recognition results are unified into structured labeled data and bound to specific three-dimensional coordinates or time segments.

[0045] Through deep learning-driven semantic recognition and label generation methods, key semantic information related to power operations can be efficiently and accurately extracted from multimodal data. Compared with traditional manual annotation or single-modal analysis methods, the automation and coverage of semantic extraction are significantly improved. Not only is the recognition accuracy higher, but it also forms a structured, reusable label set. The generated semantic labels provide clear semantic anchors for subsequent VR modeling, enabling the model to correctly restore the equipment functions and operation logic in the scene, greatly improving the teaching integrity and interactive authenticity of the virtual teaching scene;

[0046] S4: Multimodal fusion and modeling: Utilize graph neural network structures to fuse information from different modalities, automatically generate 3D VR scene models based on the combined modal information data, and configure interactive components based on teaching needs;

[0047] S4.1: After completing semantic recognition and label generation for multimodal data, a graph neural network is used to deeply fuse information from different modalities, such as images, point clouds, text, and voice. This constructs a graph structure with devices as nodes and spatial and semantic relationships as edges. Image recognition results, point cloud geometry, and voice / text commands are input into the fusion model as node attributes. The model learns the contextual dependencies and semantic commonalities between the modalities. This fused structural information is then fed into the modeling engine to automatically generate a 3D VR scene model with semantic attributes.

[0048] S4.2: Each label generated during semantic recognition and labeling extracts and structurally annotates key information such as device type, operating procedures, and location relationships from modal data models such as images, voice, and text. This provides accurate semantic input for subsequent fusion modeling. Semantic labels not only serve as the basis for identifying various objects during the modeling process but also provide node attributes and boundary conditions for fusion algorithms such as graph neural networks. This ensures a high degree of synergy between multimodal data at the geometric and semantic levels, making the generated 3D VR scenes more accurate, realistic, and efficient in terms of structural restoration, interactive configuration, and teaching logic.

[0049] By introducing a graph neural network structure for multimodal fusion, the system effectively overcomes the problems of "data fragmentation," "semantic loss," and "sclerotic scenarios" in traditional modeling approaches. This enables a comprehensive transition from "low-level perception" to "high-level understanding," making the generated 3D VR scene model more realistic than a real-world power operation environment. Interactive devices tailored to teaching needs enhance the flexibility and adaptability of the teaching system, improving both immersion and operability while also improving learning efficiency and the quality of practical training for power operators.

[0050] S5: Teaching Script Generation and Embedding: Generate standardized teaching scripts based on speech and text data analysis and embed them into VR scenes for training and evaluation;

[0051] S5.1: After completing the multimodal fusion and modeling of multimodal data, the system uses natural language processing technology (NLP) to extract and structure key information such as work procedures, operating instructions, and precautions based on the collected voice explanation content and text materials.

[0052] S5.2: By decomposing and chronologically arranging tasks, a standardized teaching script is automatically generated, including step-by-step instructions, operation prompts, standard language, and assessment points. This script is then embedded in the constructed 3D VR scene and bound to the corresponding device object or operation node. Combined with voice broadcasts, text pop-ups, and interactive guidance, it is presented in real time during the virtual training process, guiding students in their operations or taking assessments.

[0053] By automatically generating and embedding teaching scripts, the system achieves a high degree of consistency between virtual teaching content and actual work processes, effectively avoiding the subjectivity and inconsistencies inherent in traditional manual script entry. The embedding of teaching scripts not only enhances the instructional and practical nature of VR scenarios, but also allows for flexible customization based on diverse teaching needs. Supporting phased training, multi-scenario reuse, and automated scoring, the system enhances the relevance of the learning process and the objectivity of assessments, significantly improving the efficiency, standardization, and sustainability of power operation training.

[0054] S6: Real-time rendering and deployment: Render and deploy VR scenes on teaching terminals to achieve immersive operation guidance and practical training assessment;

[0055] S6.1: After completing the modeling of the 3D VR scene and embedding the teaching script, the system uses a real-time rendering engine to perform graphics rendering on the scene, including light and shadow simulation, material mapping, dynamic interactive response, etc., to ensure the visual realism and operational smoothness of the VR environment;

[0056] S6.2: The rendered scene is distributed to teaching terminals through the deployment module, including head-mounted display devices (VR helmets), desktop terminals, or mobile devices, achieving cross-platform support. Students can enter the immersive scene in the teaching terminal and perform operation training according to the embedded teaching script prompts. The system also collects interactive behaviors and operation processes for automatic scoring, skill assessment, and feedback output;

[0057] Through real-time rendering and deployment mechanisms, the system can quickly transform teaching scenarios generated by multimodal fusion into interactive, experiential, and immersive learning platforms, effectively improving the visualization and real-time responsiveness of teaching scenarios. Once deployed to various teaching terminals, it not only meets the unified management needs of centralized training scenarios, but also supports remote distributed learning models, broadening the scope of teaching. Combined with automatic recording and evaluation functions, it not only enhances the sense of participation and operational authenticity of the training process, but also significantly improves teaching efficiency, learning motivation, and the quantifiable degree of skill mastery, providing an efficient and standardized teaching solution for power operation teaching.

[0058] Finally, it should be noted that the above embodiments are merely examples for the purpose of illustrating the present invention and are not intended to limit the embodiments. Those skilled in the art will readily appreciate that other variations or modifications based on the above description are possible. It is not necessary and impossible to provide an exhaustive list of all embodiments. However, obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A multimodal data fusion VR scene automated modeling method for power operation teaching, characterized by: Includes the following: S1: Data acquisition: multi-modal data collection at the power operation site to obtain raw data such as 3D point clouds, images, audio, and text; S2: Data preprocessing, denoising, alignment, time synchronization and coordinate unification of the collected multimodal data; S3: Semantic recognition and label generation: Based on deep learning models, it analyzes images, speech, and text to extract semantic labels such as device type, operation process, and operation scenario. S4: Multimodal fusion and modeling, using graph neural network structures to fuse different modal information, and automatically generate 3D VR scene models based on the combined modal information data, and configure interactive components according to teaching needs; S5: Teaching script generation and embedding: Generate standardized teaching scripts based on speech and text data analysis and embed them into VR scenes for training and evaluation; S6: Real-time rendering and deployment: Rendering and deploying VR scenes on teaching terminals to achieve immersive operation guidance and practical training assessment.

2. The multimodal data fusion VR scene automatic modeling method for electric power operation teaching according to claim 1 is characterized in that: The data acquisition uses lidar, RGB-D camera, and voice recorder to collect high-precision three-dimensional point cloud data in the operating environment, obtain geometric information, image and depth information of power equipment and spatial structure, realize visual perception and cognition of equipment surface features and spatial position, record the voice explanations, operating instructions and environmental sounds of on-site personnel, perform time synchronization and coordinate unification, and collaboratively complete the original multimodal data acquisition of on-site images, sounds and spatial structures.

3. The multimodal data fusion VR scene automatic modeling method for electric power operation teaching according to claim 1 is characterized in that: The data preprocessing uses a filtering algorithm to remove background stray points in the lidar point cloud, lighting interference in the image, environmental noise in the voice signal, etc. for the data after collection. For spatial data, coordinate conversion and calibration operations are performed to uniformly map the point cloud and image data to a unified coordinate system to achieve spatial alignment.

4. The multimodal data fusion VR scene automatic modeling method for electric power operation teaching according to claim 1 is characterized in that: The semantic recognition and label generation identifies the structural features of the data after data preprocessing, and parses the content of the job document based on the pre-trained language model to identify semantic elements such as job processes and operation steps, and finally unifies the recognition results into structured label data and binds them to specific three-dimensional coordinates or time segments.

5. The multimodal data fusion VR scene automatic modeling method for electric power operation teaching according to claim 1 is characterized in that: The multimodal fusion and modeling, after completing the semantic recognition and label generation of multimodal data, uses a graph neural network to deeply fuse different modal information such as images, point clouds, text, and voice, and constructs a graph structure with devices as nodes and spatial and semantic relationships as edges. Image recognition results, point cloud geometry, voice / text instructions, etc. are input into the fusion model as node attributes to learn the contextual dependencies and semantic commonalities between the modalities. The fused structural information is input into the modeling engine to automatically generate a three-dimensional VR scene model with semantic attributes.

6. The multimodal data fusion VR scene automatic modeling method for electric power operation teaching according to claim 1 is characterized in that: The teaching script is generated and embedded. After completing the multimodal fusion and modeling of multimodal data, the system uses natural language processing technology (NLP) to extract and structure key information such as work processes, operating instructions, and precautions based on the collected voice explanation content and text materials. Through task decomposition and time sequence arrangement, it automatically generates standardized teaching scripts containing step instructions, operating prompts, standard language, assessment points, etc.

7. The multimodal data fusion VR scene automatic modeling method for electric power operation teaching according to claim 1 is characterized in that: The real-time rendering and deployment utilizes a real-time rendering engine to perform graphics rendering processing on the modeling of the three-dimensional VR scene and the embedding of the teaching script, including light and shadow simulation, material mapping, dynamic interactive response, etc., to ensure the visual realism and operational smoothness of the VR environment.

Citation Information

Cited By

  • Primary school mathematics multi-modal mixed reality learning resource generation method, device and equipment oriented to body agent teaching and storage medium

    CN121328711A

  • Method and device for generating multi-modal mixed reality learning resources for embodied agent teaching of primary school mathematics, equipment and storage medium

    CN121328711B

  • Virtual reality-based power transmission dense channel visual management and control system

    CN122199835A