Interactive three-dimensional teaching resource generation system and method based on 3D gaussian splashing

The interactive 3D teaching resource generation system based on 3D Gaussian splashing solves the problems of high production cost and insufficient interactivity of 3D teaching resources, realizes intelligent interaction and structured knowledge annotation, and improves the intuitiveness and personalized service of teaching.

CN121170164BActive Publication Date: 2026-02-10HUAZHONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511714595.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-02-10
Estimated Expiration
2045-11-21

AI Technical Summary

Technical Problem

Existing 3D teaching resources are costly and time-consuming to produce, and lack intelligent interactive capabilities and structured knowledge annotations, failing to meet the intuitive and personalized needs of modern education.

Method used

An interactive 3D teaching resource generation system based on 3D Gaussian splashing is adopted, which includes a multi-view image acquisition module, an enhanced 3D reconstruction module, and an intelligent interaction module. The system extracts multi-view image sequences from teaching videos through the Gaussian splashing algorithm, reconstructs 3D teaching models, determines anchor point attributes, and realizes intelligent interaction and structured knowledge annotation.

Benefits of technology

An interactive 3D teaching model was generated, which can intelligently display according to user instructions, meet personalized needs, and improve the intuitiveness and interactivity of teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170164B_ABST
    Figure CN121170164B_ABST
Patent Text Reader

Abstract

The application discloses an interactive three-dimensional teaching resource generation system and method based on 3D Gaussian splash, and belongs to the technical field of three-dimensional reconstruction.The system comprises a multi-view image acquisition module, which is used for extracting a multi-view image sequence related to a teaching theme from a teaching video resource and determining the category of each structural region in each view image; an enhanced three-dimensional reconstruction module, which is used for reconstructing the multi-view image sequence into a three-dimensional teaching model and determining the anchor point attribute of each anchor point in the three-dimensional teaching model; an intelligent interaction module, which is used for acquiring a user instruction, determining a target anchor point according to the user instruction, and acquiring the anchor point attribute corresponding to the target anchor point; and a three-dimensional visualization module, which is used for rendering and displaying the three-dimensional teaching model and displaying the anchor point attribute corresponding to the target anchor point.The three-dimensional teaching model generated by the application has the interactive ability and has been subjected to structured knowledge annotation, and the technical problem that the current three-dimensional model lacks the intelligent interaction ability is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of three-dimensional reconstruction, and in particular to an interactive three-dimensional teaching resource generation system and method based on 3D Gaussian splatting. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute the prior art.

[0003] With the in-depth development of educational informatization, although traditional two-dimensional teaching resources are rich, they lack three-dimensional display capabilities, students have difficulty observing and understanding complex teaching objects from multiple angles, and have been difficult to meet the needs of modern education for intuitiveness, interactivity and immersion. Based on this, a technology for displaying knowledge using three-dimensional teaching resources has been developed.

[0004] However, the existing three-dimensional teaching resource production has high cost and long cycle, requires professional three-dimensional modeling personnel, and is difficult to promote and apply on a large scale; and the existing three-dimensional teaching resources use static three-dimensional models, lack intelligent interaction capabilities, and cannot provide targeted explanations and answers according to the personalized needs of students; in addition, the existing three-dimensional reconstruction technology does not consider the special needs of teaching scenes, lacks structured knowledge labeling and interactive guidance mechanisms. SUMMARY

[0005] To solve the above problems, the present application proposes an interactive three-dimensional teaching resource generation system and method based on 3D Gaussian splatting, which can generate three-dimensional teaching models with interactive capabilities and structured knowledge labeling.

[0006] To achieve the above purpose, the present application adopts the following technical solutions:

[0007] In a first aspect, an interactive three-dimensional teaching resource generation system based on 3D Gaussian splatting is proposed, comprising:

[0008] A multi-view image acquisition module for extracting a multi-view image sequence related to a teaching theme from a teaching video resource and determining the category of each structural region in each view image;

[0009] An enhanced three-dimensional reconstruction module for reconstructing the multi-view image sequence extracted by the multi-view image acquisition module into a three-dimensional teaching model using a Gaussian splatting algorithm according to the category of each structural region in the view image; and determining the anchor point attribute of each anchor point in the three-dimensional teaching model;

[0010] An intelligent interaction module for obtaining a user instruction; determining a target anchor point according to the user instruction, and obtaining the anchor point attribute corresponding to the target anchor point;

[0011] The 3D visualization module is used to render and display 3D teaching models and to display the anchor point attributes corresponding to the target anchor points.

[0012] Furthermore, the multi-view image acquisition module is used to extract teaching videos related to the teaching topic from multiple teaching video platforms; extract keyframe images from the teaching videos related to the teaching topic, and the extracted keyframe images are evenly distributed within a 360° range; extract the region where the teaching topic is located from each keyframe image, determine the category of each structural region in each region where the teaching topic is located, and use it as a multi-view image sequence related to the teaching topic.

[0013] Furthermore, the enhanced 3D reconstruction module is used to determine the camera parameters corresponding to each viewpoint image; based on the camera parameters and the multi-viewpoint image sequence, it determines the initialized 3D Gaussian model and the anchor point attributes of each anchor point in the initialized 3D Gaussian model; based on the camera pose and the anchor point attributes of each anchor point in the initialized 3D Gaussian model, it calculates photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss; based on the photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss, it optimizes and updates the initialized 3D Gaussian model to obtain a 3D teaching model.

[0014] Furthermore, the enhanced 3D reconstruction module is used to determine candidate anchor points and the influence domain of each candidate anchor point based on camera parameters and multi-view image sequences; and to establish a bidirectional mapping between the candidate anchor points and each Gaussian in their influence domains to obtain an initialized 3D Gaussian model.

[0015] Furthermore, the enhanced 3D reconstruction module is used to determine the Gaussian properties in the initialized 3D Gaussian model. The process includes: determining the sparse 3D point cloud of the multi-view image sequence; determining the initial Gaussian points based on camera parameters and the sparse 3D point cloud; determining the initial Gaussian points and the opacity of each structural region based on the category of each structural region in each view image; determining the color of each view image; determining the color attribute of each Gaussian through multi-view color weighted fusion; and using spherical harmonic function coefficients to represent view-related color changes.

[0016] Furthermore, the anchor point attributes determined by the enhanced 3D reconstruction module include spatial attributes, knowledge attributes, visual attributes, and interaction attributes.

[0017] Furthermore, the enhanced 3D reconstruction module is used to determine the feature matching points between adjacent view images in a multi-view image sequence with the goal of minimizing reprojection error; and to determine camera parameters based on the feature matching points between adjacent view images.

[0018] Furthermore, the enhanced 3D reconstruction module is used to optimize and update the initialized 3D Gaussian model through coarse-grained shape recovery, detail optimization, and anchor fine-tuning stages based on photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss, thereby obtaining a 3D teaching model.

[0019] The coarse-grained shape recovery stage only optimizes the position and opacity parameters of Gaussian points in the 3D Gaussian model;

[0020] The detailed optimization stage optimizes all parameters of the 3D Gaussian model;

[0021] The anchor point fine-tuning stage only optimizes the anchor point positions and related Gaussians of the 3D Gaussian model.

[0022] Furthermore, the 3D visualization module is used to perform Gaussian splash rendering on the 3D teaching model and ensure that the anchor points in the 3D teaching model always face the camera, displaying the rendered model.

[0023] Secondly, a method for generating interactive 3D teaching resources based on 3D Gaussian splashing is proposed, including:

[0024] Extract multi-view image sequences related to the teaching topic from teaching video resources and determine the category of each structural region in each view image;

[0025] Based on the category of each structural region in the viewpoint image, the Gaussian splashing algorithm is used to reconstruct the multi-view image sequence extracted by the multi-view image acquisition module into a three-dimensional teaching model; and the anchor point attributes of each anchor point in the three-dimensional teaching model are determined.

[0026] Render and display the 3D teaching model;

[0027] Obtain user instructions;

[0028] Based on user instructions, determine the target anchor point and obtain the anchor point attributes corresponding to the target anchor point;

[0029] Display the anchor point attributes corresponding to the target anchor point.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] This invention proposes an interactive 3D teaching resource generation system and method based on 3D Gaussian splashing. The system includes a multi-view image acquisition module, an enhanced 3D reconstruction module, an intelligent interaction module, and a 3D visualization module. The multi-view image acquisition module extracts multi-view image sequences related to the teaching topic from teaching video resources. The enhanced 3D reconstruction module uses the Gaussian splashing algorithm to reconstruct the multi-view image sequences into a 3D teaching model and determines the anchor point attributes of each anchor point in the 3D teaching model. The 3D visualization module renders and displays the 3D teaching model. The intelligent interaction module obtains user commands. Based on the user commands, the target anchor point is determined, and the corresponding anchor point attributes are obtained. Then, the 3D visualization module displays the anchor point attributes. Therefore, the system can generate an interactive 3D teaching model with structured knowledge annotation.

[0032] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0033] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an undue limitation of this application.

[0034] Figure 1 This is an example of an interactive 3D teaching resource generation system architecture based on 3D Gaussian splashing.

[0035] Figure 2 The flowchart below shows an example of an interactive 3D teaching resource generation method based on 3D Gaussian splashing. Detailed Implementation

[0036] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0037] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0038] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0039] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0040] Example 1

[0041] In this embodiment, an interactive 3D teaching resource generation system based on 3D Gaussian splashing is disclosed, such as... Figure 1 As shown, it includes:

[0042] The multi-view image acquisition module is used to extract multi-view image sequences related to the teaching topic from teaching video resources and determine the category of each structural region in each view image.

[0043] The enhanced 3D reconstruction module is used to reconstruct a 3D teaching model from the multi-view image sequence extracted by the multi-view image acquisition module based on the category of each structural region in the view image and using the Gaussian splashing algorithm; and to determine the anchor point attributes of each anchor point in the 3D teaching model.

[0044] The intelligent interaction module is used to obtain user commands; based on the user commands, determine the target anchor point and obtain the anchor point attributes corresponding to the target anchor point;

[0045] The 3D visualization module is used to render and display 3D teaching models and to display the anchor point attributes corresponding to the target anchor points.

[0046] By incorporating a multi-view image acquisition module, an enhanced 3D reconstruction module, an intelligent interaction module, and a 3D visualization module, this embodiment discloses an interactive 3D teaching resource generation system based on 3D Gaussian splashing. This system can acquire multi-view images of teaching resources from multiple teaching video platforms, reconstruct a 3D teaching model of the resources using an improved 3DGS (Gaussian splashing) algorithm, and automatically embed interactive anchor points. It associates the anchor points with relevant knowledge corresponding to the teaching resources, using this knowledge as the anchor point's attribute. Then, based on user commands, a target anchor point can be determined, and its attributes can be retrieved and displayed. This achieves intelligent interaction with the 3D teaching model.

[0047] The multi-view image acquisition module is used to extract teaching videos related to the teaching topic from various teaching video platforms; extract keyframe images from the teaching videos obtained by the video retrieval unit, and the extracted keyframe images are evenly distributed within a 360° range; perform semantic segmentation on each keyframe image to determine the category of each structural region in each keyframe image, as a multi-view image sequence related to the teaching topic. The specific process includes:

[0048] S1.1: Extract teaching videos related to the teaching topic from multiple teaching video platforms. This step also includes:

[0049] S1.1.1: Generate an expanded keyword set based on the teaching theme. Taking the teaching theme of heart in biology teaching as an example, the theme words are expanded from "heart" to construct the following keyword set: ["heart structure", "heart anatomy", "heart model", "heart teaching", "cardiac anatomy"];

[0050] S1.1.2: Access multiple video platforms such as Bilibili, YouTube, and China University MOOC in parallel via API. Obtain video BV number, duration, resolution, and play count through Bilibili API, obtain video_id, duration, and resolution through YouTube Data API, and crawl open course video resources through China University MOOC.

[0051] S1.1.3: From various video platforms, videos are selected based on resolution and content relevance scores, with priority given to videos with a resolution of 1080p or higher. The relevance of the titles, descriptions, and tags of the selected videos based on resolution to the teaching topic keywords is calculated. Videos with a relevance greater than a set value are selected as relevant videos. Videos with severe shaky footage are removed from the relevant videos, and videos that contain multi-angle displays of the target are confirmed.

[0052] S1.1.4: Standardize the format of each teaching video extracted in S1.1.3, convert the format of each teaching video to MP4, and standardize the frame rate of each teaching video to 30fps.

[0053] S1.1.5: Detect continuous motion segments in each teaching video through multi-level motion feature analysis and mark their start and end times. Specifically, calculate the dense optical flow field between adjacent frames to obtain the global motion vector distribution, and simultaneously extract the inter-frame grayscale histogram difference value and structural similarity index (SSIM). Weightedly fuse these three features to generate a motion intensity time-series curve. Perform sliding window analysis on this curve. When the motion intensity value of N consecutive frames (N≥15) exceeds the adaptive threshold T, it is determined as the start point of the continuous motion segment. When the motion intensity value of M consecutive frames (M≥10) is below the threshold T, it is marked as the end point. Use timestamp smoothing to eliminate false detections caused by short-term fluctuations, and finally output the precise start and end times of each continuous motion segment.

[0054] S1.1.6: Identify valid segments containing complete rotation displays in continuous motion dynamic segments with marked start and end times, remove non-model segments containing text explanations or PPT presentations, and retain valid segments with a duration of 30 seconds to 5 minutes as teaching videos related to the teaching theme.

[0055] S1.2: Extract keyframe images from teaching videos related to the teaching topic, and try to ensure that the extracted keyframe images are evenly distributed within a 360° range. This step includes:

[0056] S1.2.1: Divide the teaching video extracted in S1.1 into frames, evaluate the frame quality of each frame, and calculate the frame quality using the following formula, which is used to select high-quality frames in step S1.2.2:

[0057] quality_score=w1*sharpness+w2*contrast+w3* brightness_uniformity.

[0058] Where quality_score is the quality score of the frame, sharpness is the sharpness, contrast is the contrast of the histogram distribution range, brightness_uniformity is the average brightness of all frames, and w1, w2 and w3 are the weights of each item.

[0059] S1.2.2: Extract SIFT feature points from each high-quality frame and perform feature descriptor matching. Estimate the homography matrix H between adjacent frames using the RANSAC algorithm. Calculate the rotation angle θ and direction of the camera view change based on matrix H. Establish a polar coordinate system with the initial frame as a reference. Accumulate and calculate the view deflection angle of each frame relative to the reference frame. Map it to the range of 0° to 360°. Divide 360° into 24 sectors to ensure that each sector has at least 2 to 3 high-quality keyframes.

[0060] S1.2.3: Motion blur detection and filtering. The cepstral analysis method is used to detect the direction and length of motion blur. When the length of the blur kernel exceeds 5 pixels or the edge sharpness value is lower than the preset threshold τ (adaptively adjusted according to the overall video quality), the frame is marked as a motion blur frame. Motion blur frames are filtered out, and frames at static or slow motion moments are retained.

[0061] S1.2.4: Perform time-uniform sampling on the frames retained in S1.2.3 to ensure uniform time distribution. Prioritize frames with complementary viewpoints and finally output 30-60 high-quality keyframes as keyframe images.

[0062] S1.3: Perform semantic segmentation on each keyframe image to determine the relevant structural regions of each teaching topic in each keyframe image and determine the category of each structural region, as a multi-view image sequence related to the teaching topic. This includes object detection and localization, pixel-level semantic segmentation, part annotation, and background removal. This step includes:

[0063] S1.3.1: Use YOLOv12 to perform object detection on each keyframe image, determine the boundary of the teaching topic in each keyframe image, and use a multi-frame voting mechanism to determine a stable teaching topic region.

[0064] S1.3.2: Apply SAM (Segment Anything Model) to perform pixel-level segmentation of the teaching theme region determined in S1.3.1 to generate a high-quality foreground mask.

[0065] S1.3.3: Based on the component segmentation model, the categories of each structural region in the foreground mask generated in S1.3.2 are identified. For example, in the heart model, the structures such as atria, ventricles, valves, and aorta are identified, and a semantic label map is generated, with each pixel corresponding to a specific part.

[0066] S1.3.4: Apply a mask to remove the background of the semantic label image generated in S1.3.3, use an image inpainting algorithm to fill the occluded area, perform edge feathering to avoid hard boundaries, and obtain the final multi-view image sequence.

[0067] Specifically, the multi-view image acquisition module includes a video retrieval unit, a keyframe extraction unit, and a semantic segmentation unit.

[0068] The video retrieval unit includes a keyword expansion module, a multi-platform parallel retrieval module, and a video quality assessment module.

[0069] The keyword expansion module is used to generate expanded keyword sets based on teaching topics.

[0070] The multi-platform parallel retrieval module is used to access multiple video platforms such as Bilibili, YouTube, and China University MOOC in parallel via API.

[0071] The video quality assessment module is used to filter videos from various video platforms based on resolution and content relevance.

[0072] The keyframe extraction unit includes a scene change detection module, a frame quality assessment module, a viewpoint coverage analysis module, and a motion blur detection module.

[0073] The scene change detection module is used to detect continuous motion dynamic segments in each teaching video through multi-level motion feature analysis and mark the start and end times.

[0074] The frame quality assessment module is used to assess the frame quality of each frame in a dynamic clip of a teaching video related to the teaching topic.

[0075] The view coverage analysis module is used to estimate the view deflection angle of each frame relative to the reference frame to ensure that the extracted keyframes are distributed as evenly as possible within a 360-degree range.

[0076] The motion blur detection module is used to filter motion-blurred frames and retain frames that are still or moving slowly in dynamic clips of teaching videos that are relevant to the teaching topic.

[0077] The semantic segmentation unit includes an object detection and localization module, a pixel-level semantic segmentation module, a component annotation module, and a background removal module, etc.

[0078] The target detection and localization module is used to determine the region where the teaching topic is located in each keyframe image;

[0079] A pixel-level semantic segmentation module is used to generate a pixel-level foreground mask for the region where the teaching topic is located;

[0080] The component annotation module is used to determine the category of each structural region in the foreground mask;

[0081] The background removal module is used to remove the background area from each keyframe image to obtain the final multi-view image sequence related to the teaching topic.

[0082] The enhanced 3D reconstruction module in this embodiment is used to determine the camera parameters corresponding to each viewpoint image; based on the camera parameters and the multi-viewpoint image sequence, it determines the initialized 3D Gaussian model and the anchor point attributes of each anchor point in the initialized 3D Gaussian model; based on the camera pose and the anchor point attributes of each anchor point in the initialized 3D Gaussian model, it calculates the photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss; based on the photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss, it optimizes and updates the initialized 3D Gaussian model to obtain a 3D teaching model.

[0083] The enhanced 3D reconstruction module is used to determine candidate anchor points and the influence domain of each candidate anchor point based on camera parameters and multi-view image sequences; and to establish a bidirectional mapping between the candidate anchor points and each Gaussian in their influence domain to obtain an initialized 3D Gaussian model.

[0084] The enhanced 3D reconstruction module is used to determine the Gaussian properties in the initialized 3D Gaussian model. The process includes: determining the sparse 3D point cloud of the multi-view image sequence; determining the initial Gaussian points based on camera parameters and the sparse 3D point cloud; determining the initial Gaussian points and the opacity of each structural region based on the category of each structural region in each view image; determining the color of each view image; determining the color attribute of each Gaussian through multi-view color weighted fusion; and using spherical harmonic function coefficients to represent view-related color changes.

[0085] The anchor point attributes determined by the enhanced 3D reconstruction module include spatial attributes, knowledge attributes, visual attributes, and interaction attributes.

[0086] The enhanced 3D reconstruction module is used to determine the feature matching points between adjacent view images in a multi-view image sequence with the goal of minimizing reprojection error; and to determine camera parameters based on the feature matching points between adjacent view images.

[0087] The enhanced 3D reconstruction module is used to optimize and update the initialized 3D Gaussian model through coarse-grained shape recovery, detail optimization, and anchor fine-tuning stages based on photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss, thereby obtaining a 3D teaching model.

[0088] The coarse-grained shape recovery stage only optimizes the position and opacity parameters of Gaussian points in the 3D Gaussian model;

[0089] The detailed optimization stage optimizes all parameters of the 3D Gaussian model;

[0090] The anchor point fine-tuning stage only optimizes the anchor point positions and related Gaussians of the 3D Gaussian model.

[0091] Specifically, the enhanced 3D reconstruction module uses the Gaussian splashing algorithm to reconstruct a 3D teaching model from the multi-view image sequence extracted by the multi-view image acquisition module; and the process of determining the anchor point attributes of each anchor point in the 3D teaching model includes:

[0092] S2: Camera parameter estimation and scene initialization, including feature extraction and matching, initial camera parameter estimation and bundle adjustment optimization.

[0093] S2.1: Extract SIFT features and SuperPoint depth features from each viewpoint image, perform feature matching and validation on the SIFT features and SuperPoint depth features respectively, and remove abnormal matches. This step includes...

[0094] S2.1.1: Extract SIFT feature points and SuperPoint deep feature points from each viewpoint image. SIFT detects keypoints in multi-scale space by constructing a Gaussian difference pyramid and calculates a 128-dimensional descriptor vector for each keypoint. This descriptor is generated by statistically analyzing the gradient direction histogram in the neighborhood of the keypoint and has rotation invariance and scale invariance. The SuperPoint network outputs a 256-dimensional deep learning descriptor, which is learned by a convolutional neural network and can capture higher-level semantic information. Both feature descriptors are L2 normalized to a magnitude of 1 to eliminate the influence of illumination changes and improve the robustness of subsequent matching.

[0095] S2.1.2: SIFT and SuperPoint features are matched between adjacent viewpoint images. Potential corresponding points are found by calculating the Euclidean distance between feature descriptors under different viewpoints. Reliable matching pairs are initially screened using nearest neighbor matching combined with Lowe's ratio test. The RANSAC algorithm is used to estimate the fundamental matrix between viewpoints. The geometric consistency of matching points is verified by epipolar geometric constraints, and outliers that do not meet the constraints are eliminated. The verified matching points are extended to construct a feature trajectory chain across multiple viewpoints. The purpose is to establish the two-dimensional projection correspondence of the same 3D point under different viewpoints, providing stable feature correspondence for subsequent multi-view 3D reconstruction and camera parameter estimation.

[0096] S2.1.3: Perform global consistency analysis on the constructed cross-view feature trajectory chain, verify the transitivity of feature point matching across multiple views through three-view cyclic consistency check, and ensure that all feature points on the same trajectory do indeed correspond to the same 3D point in space; for incomplete trajectories caused by occlusion or detection failure, trajectory completion is performed through local affine transformation model and image patch matching; finally, abnormal matching points are eliminated based on the statistical distribution of reprojection error to ensure that high-precision and high-reliability feature correspondence is provided for camera parameter estimation.

[0097] S2.2: Initial estimation of camera parameters, calibration of camera intrinsic parameters, estimation of relative pose, and incremental reconstruction by gradually adding new viewpoints. This step includes:

[0098] S2.2.1: Perform initial estimation of camera intrinsic parameters for each viewpoint image. Assume that all viewpoints use the same camera equipment, the focal length is set to a fixed value, and the principal point coordinates are set at the pixel coordinate center position of each input image (cx=width / 2, cy=height / 2), that is, assume that the optical axis passes through the geometric center of the imaging plane; the radial distortion coefficients k1, k2 and the tangential distortion coefficients p1, p2 are initialized to zero, and the initial intrinsic parameter matrix K is constructed. These rough intrinsic parameter estimates will be used as the initial values ​​for the bundled adjustment in S2.3.

[0099] S2.2.2: Relative pose estimation. Analyze the verified feature matching points between adjacent viewpoint images obtained in S2.1, randomly select 5 pairs of corresponding points, and use the five-point method to solve for the essential matrix E, which encodes the relative geometric relationship between the two viewpoints. Decompose the essential matrix E into a relative rotation matrix R and a translation vector t through singular value decomposition (SVD), where there are four possible solutions. Determine the unique correct solution by triangulating the test points and checking their visibility in front of the two cameras. Since monocular vision cannot recover the absolute scale, the baseline length between the first pair of cameras is normalized to 1 as the scale benchmark for the entire reconstruction. These initial pose estimates will be further optimized in S2.3.

[0100] S2.2.3: Incremental reconstruction selects two viewpoint images with the most feature matching and a moderate baseline width (disparity angle between 5° and 60°) as the initial image pair. This image pair should have sufficient common feature points and a stable geometric configuration. The matching feature points in the initial image pair are triangulated, and their three-dimensional coordinates in the world coordinate system are calculated. These 3D points represent the actual spatial positions of the object surface. New viewpoints with the most common feature points with the reconstructed parts are added incrementally. The pose of the new viewpoints is estimated using the PnP algorithm. The newly observed feature points are triangulated to generate new 3D points. After every 3 to 5 new viewpoints are added, the local binding adjustment in S2.3.2 is triggered to ensure the stability of incremental reconstruction.

[0101] S2.3: Layered bundling adjustment and optimization, refining the initial camera parameters and 3D points obtained in S2.2, performing local optimization first and then global optimization. This step includes:

[0102] S2.3.1: Construct a nonlinear least squares optimization objective function. Based on the camera parameters and 3D point coordinates reconstructed in S2.2, construct an objective function that minimizes the reprojection error.

[0103] min Σ||π(Ki, Ri, ti, Xj) - xij|| 2

[0104] Where Ki is the camera intrinsic parameter, Ri, ti are the camera extrinsic parameters, Xj is the coordinates of the 3D point, π is the projection function, which projects the 3D point Xj in the world coordinate system onto the image plane through the camera intrinsic parameter Ki and the extrinsic parameter (Ri, ti), and xij is the coordinates of the 2D feature point actually observed by the 3D point in the i-th view.

[0105] S2.3.2: Implementation of multi-scale hierarchical optimization strategy. A hierarchical optimization strategy from local to global is adopted. First, in the incremental reconstruction process of S2.2.3, local bundled adjustment is triggered every 3-5 newly added views, optimizing only the poses of the newly added 5-10 cameras and their observed 3D points, while keeping the other parameters fixed. This local optimization has high computational efficiency and can correct accumulated errors in a timely manner. After all views have been added, global bundled adjustment is performed, optimizing the intrinsic and extrinsic parameters of all cameras and the coordinates of all 3D points, and performing 5-10 iterations until convergence. Finally, an alternating optimization strategy is adopted, first fixing the 3D points and optimizing only the camera parameters, then fixing the camera parameters and optimizing only the 3D points, alternating in this way for 2-3 rounds to further improve the optimization accuracy.

[0106] S2.3.3: Robust Optimization and Outlier Removal: Replacing the standard quadratic loss function with the Huber loss function ρ(r) = {r 2 / 2 if |r|≤δ; δ(|r|-δ / 2) if |r|>δ}, where r is the reprojection error and δ is the threshold (initially set to 2 pixels). This function maintains a quadratic penalty for small errors and a linear penalty for large errors, effectively reducing the impact of abnormal matching points. After each round of optimization, the median absolute deviation (MAD) of all reprojection errors is calculated, and the threshold δ=1.48×MAD×k (k decreases with the number of iterations) is dynamically updated. Observations with reprojection errors exceeding 3δ are marked as outliers and removed from the optimization. If a 3D point has fewer than 3 effective observations, the point is deleted. Through iterative optimization and removal processes, high-precision camera intrinsic and extrinsic parameters and sparse 3D point clouds are finally obtained. These optimized camera parameters will be directly used for viewpoint localization and projection calculation in subsequent 3D Gaussian reconstruction, while sparse 3D points will serve as initialization seed points for the Gaussian ellipsoid, providing a reliable geometric constraint basis for dense 3D Gaussian representation.

[0107] This embodiment sets differentiated point cloud densities for each structural region category in the image. Based on the sparse 3D point cloud data obtained in S2.3.3 and the categories of structural regions in the image, Gaussian attributes are initialized, and Gaussian density is added at the structural boundaries to protect teaching features.

[0108] The structural regions are categorized into critical structures, main structures, and other structures. High density (e.g., 5000) is set for critical structural regions, medium density (e.g., 2000) is set for main structural regions, and low density (e.g., 500) is set for other regions.

[0109] When constructing the initial 3D Gaussian model, the Gaussian properties in the initial 3D Gaussian model were also determined. The process included: determining the sparse 3D point cloud of the multi-view image sequence; determining the initial Gaussian points based on camera parameters and the sparse 3D point cloud; determining the initial Gaussian points and the opacity of each structural region based on the category of each structural region in each view image; determining the color of each view image; determining the color attribute of each Gaussian through multi-view color weighted fusion; and using spherical harmonic function coefficients to represent view-related color changes. Specifically, based on the S2-optimized sparse 3D point cloud and camera parameters, an initial 3D Gaussian representation of the scene was constructed. First, the sparse point cloud was used as the Gaussian center position, and then multi-view stereo matching or... Spatial interpolation performs dense sampling in key structural regions according to a preset density to generate initial Gaussian points. For each Gaussian center, the local covariance matrix is ​​calculated using PCA analysis of K nearest neighbors to determine the shape and orientation parameters of the Gaussian ellipsoid. Based on the semantic segmentation results, the scene is layered, with high opacity (0.9~1.0) set for key structural regions, medium opacity (0.6~0.8) set for main structural regions, and low opacity (0.3~0.5) set for other structural regions to ensure the visual priority of teaching content. The color attributes of each Gaussian are determined through multi-view color weighted fusion, and spherical harmonic function coefficients are used to represent view-related color changes to ensure that key teaching content is given priority in 3D reconstruction.

[0110] Gaussian density is increased at the boundaries of each structural region to reserve Gaussian clusters for important key features and establish hierarchical relationships between Gaussians.

[0111] This embodiment generates candidate anchor point positions based on the structural region center and Gaussian point cloud geometric features. It defines the spatial, knowledge, visual, and interactive attributes of each candidate anchor point, establishes a spatial index using a KD-tree, calculates the influence domain of each candidate anchor point, and establishes a bidirectional mapping between the candidate anchor point and each Gaussian point within its influence domain, thereby obtaining an initialized 3D Gaussian model. Specifically:

[0112] Based on the importance of the structural region to teaching, when the importance exceeds a certain threshold, the center of the structural region is set as the anchor point center. At the same time, based on the local geometric features of the Gaussian point cloud, local density extreme points (locations where the point cloud density changes by more than 30%), curvature change points (locations where the curvature value exceeds twice the standard deviation of the average curvature), and geometric branch points (such as the intersection of multiple planes) are identified. Key points are selected from all anchor point centers as candidate anchor points.

[0113] Define spatial attributes (location, orientation, radius of influence), knowledge attributes (knowledge ID, display level, associated anchors), visual attributes (icon type, color scheme), and interaction attributes (trigger distance, interaction mode, animation sequence) for each candidate anchor point.

[0114] To accelerate neighborhood queries, a spatial index KD-tree is constructed. The Gaussian set controlled by each candidate anchor point is determined, and this Gaussian set represents the influence domain of that candidate anchor point. A bidirectional mapping table is established between each Gaussian point and the anchor point within the influence domain. The forward mapping records the set of Gaussian points controlled by each candidate anchor point, and the reverse mapping records the candidate anchor points to which each Gaussian point belongs. The influence weight w = exp(-d) is calculated based on the distance. 2 / 2σ 2 ), where d is the distance from the Gaussian point to the candidate anchor point, and σ is 1 / 3 of the influence radius. This weighting mechanism ensures a smooth transition of the structural boundary, so that the anchor point attribute information can be accurately mapped to the 3D Gaussian structure.

[0115] This embodiment combines photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss to calculate a multi-constraint loss function, and optimizes the initialized 3D Gaussian model in the coarse-grained shape recovery stage, the detail optimization stage, and the anchor point fine-tuning stage.

[0116] The multi-constraint loss function is:

[0117] L_total = λ1 * L_photo + λ2 * L_semantic + λ3 * L_anchor + λ4 * L_structure

[0118] Where L_total is the multi-constraint loss, and λ1, λ2, λ3, and λ4 are weighting coefficients.

[0119] The photometric uniformity loss function L_photo is:

[0120] L_photo = Σ||I_rendered - I_ground_truth|| 2

[0121] Where L_photo is the photometric consistency loss, I_rendered is the model image obtained through differentiable rendering, I_ground_truth is the real image from the corresponding viewpoint, and N is the total number of pixels.

[0122] The semantic consistency loss function is:

[0123] L_semantic = Σ CE(S_rendered, S_ground_truth)

[0124] Where L_semantic is the semantic consistency loss, S_rendered and S_ground_truth are the rendered and ground semantic segmentation maps, respectively, CE is the cross-entropy loss, and M is the total number of pixels.

[0125] The anchor point stability loss function is:

[0126] L_anchor = Σ||anchor_pos_t - anchor_pos_t-1|| 2

[0127] Where L_anchor is the anchor stability loss, anchor_pos_t and anchor_pos_t-1 are the anchor positions, and t is the optimization iteration step.

[0128] The structure preservation loss function is:

[0129]

[0130] Where L_structure is the structure preservation loss, and the first term is the edge preservation loss (λ_edge=0.8), obtained through the gradient operator. The first term ensures the sharpness of the model outline and internal structure boundaries; the second term is the adaptive surface smoothing loss (λ_smooth=0.2), where G_k is the Gaussian point density field. It applies strong smoothing constraints in flat regions and reduces smoothing intensity in feature regions, thus preserving details while avoiding noise.

[0131] The optimization process consists of three stages: coarse-grained shape recovery, detail optimization, and anchor point fine-tuning. The coarse-grained shape recovery stage undergoes 1000 iterations with a large learning rate to quickly converge to the approximate shape, optimizing only the position and opacity parameters of the Gaussian points. The detail optimization stage undergoes 2000 iterations with a medium learning rate, optimizing all parameters and increasing the weights of semantic consistency loss, anchor point stability loss, and structure preservation loss. The anchor point fine-tuning stage undergoes 500 iterations with a small learning rate, fixing most parameters and optimizing only the anchor point position and related Gaussian points.

[0132] During the optimization of the 3D Gaussian model, adaptive density control is used to add Gaussians in high gradient regions, remove Gaussians with excessively low opacity, and exclude Gaussians around anchor points from pruning.

[0133] This embodiment also defines four LoD levels, and simplifies the levels through Gaussian clustering and importance sampling, and performs smooth transition processing between levels.

[0134] The Gaussian counts in the four LoD levels are 50,000, 10,000, 2,000, and 500, respectively.

[0135] Each level is simplified, similar Gaussians are merged, and Gaussians with high visual contribution are retained to ensure the functionality of anchor points at each level.

[0136] A progressive loading strategy is adopted to smoothly transition between layers. Layer switching is triggered by view distance, and hybrid rendering is used to avoid abrupt changes.

[0137] Next, the enhanced 3D reconstruction module proposed in this embodiment will be described in detail.

[0138] The enhanced 3D reconstruction module uses an improved 3D Gaussian splashing algorithm to reconstruct multi-view images into 3D teaching models. This module includes a camera pose estimation unit, a teaching-oriented 3DGS reconstruction unit, and an interactive anchor point embedding unit.

[0139] The camera pose estimation unit is used to estimate the camera parameters for each viewpoint image and generate an initial sparse 3D point cloud.

[0140] The instruction-oriented 3DGS reconstruction unit is used to determine the initial 3D Gaussian model based on camera parameters and the initial sparse 3D point cloud; calculate photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss based on the camera pose and the anchor point attributes of each anchor point in the initial 3D Gaussian model; and optimize and update the initial 3D Gaussian model based on the photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss to obtain the 3D instruction model.

[0141] The interactive anchor embedding unit is used to automatically embed knowledge anchors with spatial location, knowledge attributes, and interactive attributes into a 3D model.

[0142] The interactive anchor embedding unit includes an anchor definition module and an anchor mapping module;

[0143] The anchor point definition module is used to automatically generate candidate anchor points and define anchor point attributes for each candidate anchor point;

[0144] The anchor point mapping module is used to establish a bidirectional mapping relationship between anchor points and each Gaussian exciton in the influence domain. It supports dynamic rendering control of local areas and maintains the consistency between the anchor point position and the relevant Gaussian exciton through anchor point stability loss during the 3DGS optimization process.

[0145] The camera pose estimation unit includes a feature extraction and matching module, a camera intrinsic parameter calibration module, and a relative pose estimation module.

[0146] The feature extraction and matching module is used to extract SIFT features and SuperPoint depth features for each viewpoint image, perform feature matching and verification on the SIFT features and SuperPoint depth features respectively, and remove abnormal matches.

[0147] The camera intrinsic parameter calibration module is used for initial estimation of camera parameters.

[0148] The relative pose estimation module is used for hierarchical bundling adjustment and optimization.

[0149] The instruction-oriented 3DGS reconstruction unit includes a semantic awareness initialization module, a multi-constraint optimization module, and a multi-stage optimization strategy module.

[0150] The semantic awareness initialization module is used to determine the Gaussian density based on each structure category, such as using a high density (e.g., 5000 points / cm) for key structural regions. 3 Initialization, the main structural area uses a density of 2000 points / cm². 3 Initialize, using low density such as 500 dots / cm² in other structural areas. 3 initialization;

[0151] The multi-constraint optimization module is used to calculate photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss based on the camera pose and the anchor point attributes of each anchor point in the initialized 3D Gaussian model.

[0152] The multi-stage optimization strategy module is used to jointly optimize photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss, and to perform staged optimization on the initialized 3D model. The staged optimization strategy includes a coarse-grained shape recovery stage, a detail optimization stage, and an anchor point fine-tuning stage.

[0153] The 3D visualization module in this embodiment is used to perform Gaussian splash rendering on the 3D teaching model and ensure that the anchor points in the 3D teaching model always face the camera. It displays the rendered model and realizes real-time rendering of 3D teaching resources and interactive operation based on anchor points.

[0154] This embodiment compiles Gaussian rendering, anchor point rendering, and contour rendering shaders to achieve Gaussian splash rendering. It ensures that the anchor points in the 3D teaching model always face the camera, including view frustum clipping, depth sorting, and alpha blending. An anchor point rendering layer is implemented using Billboard technology and distance attenuation effects. The process includes:

[0155] Initialize the WebGL2.0 context, define and compile the Gaussian rendering, anchor rendering, and contour rendering shaders, and create the frame buffer and depth buffer for the rendering target;

[0156] Perform view frustum clipping to remove Gaussians outside the field of view, perform depth sorting to render from back to front, and perform alpha blending to achieve the correct transparency.

[0157] Billboard technology is used to ensure that anchor points always face the camera, distance decay is set to make distant anchor points fade out automatically, obscured anchor points are set to be semi-transparent, and animation effects such as pulse and rotation are set for anchor points.

[0158] This embodiment also constructs a spherical coordinate system, controls camera rotation based on the spherical coordinate system, achieves anchor point interactive detection through ray projection, and generates interactive feedback such as highlight effects, region isolation, and automatic viewpoint adjustment, including:

[0159] Basic camera control involves updating camera parameters by calculating the pitch angle and other rotations based on a spherical coordinate system, and calculating logarithmic scaling factors for scaling operations to ensure smooth scaling.

[0160] Anchor point interaction detection involves projecting rays from screen coordinates into 3D space, quickly eliminating non-intersecting anchor points through bounding box testing, determining the nearest anchor point through precise intersection testing, and selecting the most important anchor point when overlapping.

[0161] The interactive feedback effect highlights specific interactive areas, while fading out others and automatically adjusts to the best viewing angle.

[0162] This embodiment also includes a labeled UI component, which achieves intelligent layout by calculating screen space occupancy and occlusion, and supports content display with scrollbars, multimedia embedding, etc.

[0163] Label the UI components, including the name of the anchor point, a brief description of the structure corresponding to the anchor point, and the detailed area of ​​the corresponding structure on the 3D teaching model.

[0164] Intelligent adjustment of annotation display involves calculating screen space occupancy, performing overlap detection and avoidance, calculating automatic routing of the leader line from the model to the annotation, and performing responsive adaptation.

[0165] Content display control includes enabling hierarchical expansion and collapse, automatic scroll bar management, and embedded display of multimedia content.

[0166] The 3D visualization module includes a rendering engine unit, an interactive control unit, and an annotation display unit.

[0167] The rendering engine unit includes the basic Gaussian renderer, anchor point rendering layer, and multi-level detail manager, etc.

[0168] A basic Gaussian renderer is used to implement view frustum clipping, depth sorting, and alpha blending;

[0169] Anchor point rendering layer, used to use Billboard technology to ensure that anchor points always face the camera, achieving distance attenuation and occlusion processing;

[0170] A multi-level detail manager is used to automatically switch between different precision levels of the model based on the viewing distance.

[0171] The interaction control unit includes a camera controller, an anchor point interaction detector, and an interaction feedback generator.

[0172] The camera controller supports spherical coordinate system rotation and logarithmic scaling; the anchor point interaction detector determines the user-selected anchor point through ray casting and bounding box testing; and the interaction feedback generator enables highlighting effects, area isolation, smooth transitions, and automatic viewpoint adjustment.

[0173] The annotation display unit includes an intelligent layout module and a content display controller.

[0174] The intelligent layout module uses intelligent layout algorithms to perform screen space occupancy calculations, overlap detection and avoidance, and automatic routing of lead lines; the content display controller supports hierarchical expansion / collapse, multimedia content embedding, and related knowledge recommendations.

[0175] The intelligent interaction module in this embodiment utilizes a large language model to provide context-aware question-and-answer functionality associated with anchor points, including a knowledge graph construction unit, a content generation unit, and a multimodal perception question-and-answer unit.

[0176] The knowledge graph construction unit collects knowledge points related to anchor points from multiple sources such as textbooks, encyclopedias, and papers, unifies synonyms and aliases for entity alignment, defines relation type ontology, extracts attributes such as function, features, and parameters, and constructs a hierarchical tree of teaching resource knowledge structure, i.e., a knowledge graph. The mapping relationship between entities and anchor points in the knowledge graph includes attributes such as name, function, and location.

[0177] Taking the heart structure as an example, the mapping is as follows:

[0178] knowledge_mapping = {

[0179] 'anchor_id': 'mitral_valve_001',

[0180] 'entity_id': 'KB_mitral_valve',

[0181] 'properties': {

[0182] 'name': ['mitral valve', 'mitral valve'],

[0183] 'function': 'Prevent blood backflow',

[0184] 'location': 'between the left atrium and left ventricle'

[0185] 'diseases': ['mitral stenosis', 'mitral regurgitation']

[0186] }

[0187] }

[0188] The intelligent interaction module in this embodiment also connects to the large language model API interface, designs multiple types of prompt word templates, maintains the conversation history and the context state of the current focus, and calls the large language model to generate resource introduction content based on the prompt word templates according to different knowledge points.

[0189] Design a unified API call interface for a large language model, and build prompt word templates, such as:

[0190] prompt_template = """

[0191] You are a professional teacher using 3D models to explain things to your students.

[0192] The current student is viewing: {current_anchor_name}

[0193] Relevant background knowledge: {knowledge_context}

[0194] Student question: {user_question}

[0195] Please provide:

[0196] 1. Answer the question directly (within 100 words).

[0197] 2. In-depth explanation

[0198] 3. Recommended Related Knowledge

[0199] Note: Maintain professionalism while keeping it easy to understand.

[0200] """

[0201] Perform conversation history tracking, maintain current focus, investigate prompt words for the construction of the large language model, generate explanatory content, and update the knowledge status.

[0202] Real-time detection (VAD) of user commands, which are user voice commands; Automatic Recognition (ASR) of user voice command content and noise suppression; Identification of user voice command content to determine prompt words; Invocation of large language model API interface based on prompt words to generate corresponding text; Text-to-Speech (TTS) of generated text.

[0203] The knowledge graph construction unit includes a multi-source knowledge collector, a knowledge structuring processor, and an anchor-knowledge mapper. The multi-source knowledge collector extracts relevant knowledge points from sources such as textbooks, encyclopedias, and academic papers. The knowledge structuring processor performs entity alignment, relationship standardization, attribute extraction, and hierarchical construction. The anchor-knowledge mapper establishes the correspondence between anchor IDs and knowledge entities.

[0204] The content generation unit includes a prompt word design module and a knowledge generation module. The prompt word design module designs corresponding prompt words for different types of knowledge points; the knowledge generation module calls a large language model and combines the corresponding prompt words to generate the explanation content for the knowledge point based on the knowledge entity corresponding to the anchor point.

[0205] The multimodal perceptual question-answering unit includes a multimodal input module, a conversation state tracker, and a question-answering response module. The multimodal input module performs real-time voice activity detection and converts the input speech into text content through speech recognition. The conversation state tracker records the user's interaction history and the current anchor point. The question-answering response module generates response content by calling a large language model based on the user input, combined with the historical dialogue and the current anchor point, and synthesizes speech.

[0206] The interactive 3D teaching resource generation system based on 3D Gaussian splashing proposed in this embodiment utilizes multiple teaching video platforms to acquire multi-view images of teaching resources, uses an improved 3DGS algorithm to reconstruct the 3D teaching resources and automatically embeds interactive anchor points, and uses a large language model to realize intelligent explanation and question-and-answer of teaching resources by associating anchor points with relevant knowledge graphs corresponding to teaching resources; thus providing a new path for constructing intelligent interactive 3D teaching resources.

[0207] Example 2

[0208] In this embodiment, a method for generating interactive 3D teaching resources based on 3D Gaussian splashing is disclosed, such as... Figure 2 As shown, it includes:

[0209] Extract multi-view image sequences related to the teaching topic from teaching video resources and determine the category of each structural region in each view image;

[0210] Based on the category of each structural region in the viewpoint image, the Gaussian splashing algorithm is used to reconstruct the multi-view image sequence extracted by the multi-view image acquisition module into a three-dimensional teaching model; and the anchor point attributes of each anchor point in the three-dimensional teaching model are determined.

[0211] Render and display the 3D teaching model;

[0212] Obtain user instructions;

[0213] Based on user instructions, determine the target anchor point and obtain the anchor point attributes corresponding to the target anchor point;

[0214] Display the anchor point attributes corresponding to the target anchor point.

[0215] The method disclosed in Example 2 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.

[0216] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0217] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. An interactive 3D teaching resource generation system based on 3D Gaussian splashing, characterized in that, include: The multi-view image acquisition module is used to extract multi-view image sequences related to the teaching topic from teaching video resources and determine the category of each structural region in each view image. The enhanced 3D reconstruction module is used to reconstruct a 3D teaching model from the multi-view image sequence extracted by the multi-view image acquisition module based on the category of each structural region in the view image and using the Gaussian splashing algorithm; and to determine the anchor point attributes of each anchor point in the 3D teaching model. The enhanced 3D reconstruction module is used to determine the camera parameters corresponding to each viewpoint image; based on the camera parameters and multi-viewpoint image sequences, it determines the initialized 3D Gaussian model and the anchor point attributes of each anchor point in the initialized 3D Gaussian model; the anchor point attributes of each anchor point determined by the enhanced 3D reconstruction module include spatial attributes, knowledge attributes, visual attributes, and interaction attributes; based on the camera pose and the anchor point attributes of each anchor point in the initialized 3D Gaussian model, it calculates photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss; based on the photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss, it optimizes and updates the initialized 3D Gaussian model through coarse-grained shape recovery, detail optimization, and anchor point fine-tuning stages to obtain a 3D teaching model; The coarse-grained shape recovery stage only optimizes the position and opacity parameters of Gaussian points in the 3D Gaussian model; The detailed optimization stage optimizes all parameters of the 3D Gaussian model; The anchor point fine-tuning stage only optimizes the anchor point positions and related Gaussians of the 3D Gaussian model; The intelligent interaction module is used to obtain user commands; based on the user commands, determine the target anchor point and obtain the anchor point attributes corresponding to the target anchor point; The 3D visualization module is used to render and display 3D teaching models and to display the anchor point attributes corresponding to the target anchor points.

2. The interactive 3D teaching resource generation system based on 3D Gaussian splashing as described in claim 1, characterized in that, The multi-view image acquisition module is used to extract teaching videos related to the teaching topic from multiple teaching video platforms; extract keyframe images from the teaching videos related to the teaching topic, and the extracted keyframe images are evenly distributed within a 360° range; extract the region where the teaching topic is located from each keyframe image, determine the category of each structural region in each region where the teaching topic is located, and use it as a multi-view image sequence related to the teaching topic.

3. The interactive three-dimensional teaching resource generation system based on 3D Gaussian splashing as described in claim 1, characterized in that, The enhanced 3D reconstruction module is used to determine the feature matching points between adjacent view images in a multi-view image sequence with the goal of minimizing reprojection error; and to determine camera parameters based on the feature matching points between adjacent view images.

4. The interactive 3D teaching resource generation system based on 3D Gaussian splashing as described in claim 1, characterized in that, The enhanced 3D reconstruction module is used to determine candidate anchor points and the influence domain of each candidate anchor point based on camera parameters and multi-view image sequences; and to establish a bidirectional mapping between the candidate anchor points and each Gaussian in their influence domains to obtain an initialized 3D Gaussian model.

5. The interactive 3D teaching resource generation system based on 3D Gaussian splashing as described in claim 4, characterized in that, The enhanced 3D reconstruction module is used to determine the Gaussian properties in the initialized 3D Gaussian model. The process includes: determining the sparse 3D point cloud of the multi-view image sequence; determining the initial Gaussian points based on camera parameters and the sparse 3D point cloud; determining the initial Gaussian points and the opacity of each structural region based on the category of each structural region in each view image; determining the color of each view image; determining the color attribute of each Gaussian through multi-view color weighted fusion; and using spherical harmonic function coefficients to represent view-related color changes.

6. The interactive 3D teaching resource generation system based on 3D Gaussian splashing as described in claim 1, characterized in that, The 3D visualization module is used to perform Gaussian splash rendering on 3D teaching models and ensure that the anchor points in the 3D teaching models always face the camera, displaying the rendered models.

7. A method for generating interactive 3D teaching resources based on 3D Gaussian splashing, characterized in that, include: Extract multi-view image sequences related to the teaching topic from teaching video resources and determine the category of each structural region in each view image; Based on the category of each structural region in the viewpoint image, the Gaussian splashing algorithm is used to reconstruct the multi-view image sequence extracted by the multi-view image acquisition module into a three-dimensional teaching model; and the anchor point attributes of each anchor point in the three-dimensional teaching model are determined. The enhanced 3D reconstruction module is used to determine the camera parameters corresponding to each viewpoint image; based on the camera parameters and multi-viewpoint image sequences, it determines the initialized 3D Gaussian model and the anchor point attributes of each anchor point in the initialized 3D Gaussian model; the anchor point attributes of each anchor point determined by the enhanced 3D reconstruction module include spatial attributes, knowledge attributes, visual attributes, and interaction attributes; based on the camera pose and the anchor point attributes of each anchor point in the initialized 3D Gaussian model, it calculates photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss; based on the photometric consistency loss, semantic consistency loss, anchor point stability loss, and structure preservation loss, it optimizes and updates the initialized 3D Gaussian model through coarse-grained shape recovery, detail optimization, and anchor point fine-tuning stages to obtain a 3D teaching model; The coarse-grained shape recovery stage only optimizes the position and opacity parameters of Gaussian points in the 3D Gaussian model; The detailed optimization stage optimizes all parameters of the 3D Gaussian model; The anchor point fine-tuning stage only optimizes the anchor point positions and related Gaussians of the 3D Gaussian model; Render and display the 3D teaching model; Obtain user instructions; Based on user instructions, determine the target anchor point and obtain the anchor point attributes corresponding to the target anchor point; Display the anchor point attributes corresponding to the target anchor point.

Citation Information

Patent Citations

  • Virtual reality-based teaching knowledge point display method and system

    CN113724399A

  • Three-dimensional model reconstruction method based on 3D Gaussian Splitting

    CN119091051A