Automobile fault diagnosis method and system based on 3DGS modeling and human-computer interaction
Patent Information
- Application Number
- CN202611073030.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-08-18
AI Technical Summary
[0005]针对现有技术存在的不足,本发明的目的在于提供基于3DGS建模和人机交互的汽车故障诊断方法及系统,以解决传统汽车故障诊断中故障定位直观有限、交互即时性差及故障诊断效率低的技术问题;通过3DGS算法构建可实时渲染且轻量化的3DGS数字孪生体,将多源数据与加权故障概率转化为可视化的三维故障概率热力图,实现故障部位与车辆几何结构的绑定,通过AR虚实叠加、眼动-手势融合交互完成故障快速定位,同时构建低延迟的多用户三维空间同步通道,支持远程专家实时标注与3D维修动画指导,从可视化、交互性、协同性三方面提升故障诊断效率,降低对人工经验与专业能力的依赖,实现高效且沉浸式的汽车故障诊断与远程协同维修
[0060] 1. This invention addresses the shortcomings of traditional diagnostic methods based on text, two-dimensional charts, or single sensor data, which cannot intuitively display the three-dimensional spatial distribution of faults and require manual correlation between data and component locations. It employs a 3DGS algorithm to generate a lightweight 3DGS digital twin, combined with a Gaussian point-by-point attribute assignment mechanism, mapping the fault probability distribution matrix into a three-dimensional fault probability heatmap bound to the vehicle's geometry. This transforms fault risk from abstract data into a visualized three-dimensional space. A millisecond-level synchronous update link is established, and combined with virtual-real fusion overlay technology, fault locations are mapped one-to-one with real vehicle components. Repair personnel can intuitively view the severity and spatial distribution of faults. By calculating weighted fault probabilities and filtering invalid data, the invention eliminates the need for manual analysis of massive amounts of sensor data, shortening fault location time, reducing reliance on repair personnel experience, and improving fault diagnosis efficiency and accuracy.
Smart Images

Figure CN122597675A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automotive fault detection technology, specifically to an automotive fault diagnosis method and system based on 3DGS modeling and human-computer interaction. Background Technology
[0002] Current automotive fault diagnosis typically relies on fault code reading, experience-based judgment, and 2D image detection, which struggles to intuitively present the spatial relationships between the vehicle's 3D structure and internal components. Furthermore, it suffers from insufficient fault location accuracy and poor user experience. Traditional 3D reconstruction methods can reconstruct vehicle structures, but their slow modeling speed and sensitivity to scene lighting prevent them from meeting the needs of rapid on-site diagnosis. Simultaneously, conventional diagnostic systems lack effective human-computer interaction mechanisms, making it difficult for repair personnel to intuitively annotate fault areas, adjust perspectives, and perform spatial measurements, leading to the potential for missed or misdiagnosed complex and hidden faults. 3D Gaussian sputtering (3DGS), as an emerging radiation field modeling method, can rapidly construct high-quality 3D scenes from sparse images and achieve real-time rendering. It boasts advantages such as fast reconstruction speed, high fault location accuracy, and strong real-time rendering capabilities, providing a new approach for visualized fault diagnosis.
[0003] Existing patent application CN119672221A discloses a fault display method, vehicle, device, and storage medium. The method includes: acquiring vehicle fault information; acquiring fault sub-models of faulty components in the vehicle based on the fault information; combining the fault sub-models with a three-dimensional frame model of the vehicle's overall frame, and displaying the combined three-dimensional frame model. The three-dimensional frame model includes sub-models corresponding to each vehicle component, and the display effects of the fault sub-models and the three-dimensional frame model differ. This technical solution can display faulty components of a vehicle to users more efficiently and intuitively based on three-dimensional technology.
[0004] Although existing technologies have enabled preliminary three-dimensional visualization of faulty vehicle components, there are still problems such as limited intuitiveness in fault location, poor real-time interactivity, and low efficiency in fault diagnosis. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention aims to provide a vehicle fault diagnosis method and system based on 3DGS modeling and human-computer interaction. This addresses the technical problems of limited intuitive fault location, poor real-time interaction, and low efficiency in traditional vehicle fault diagnosis. By constructing a lightweight, real-time renderable 3DGS digital twin using 3DGS algorithms, multi-source data and weighted fault probabilities are transformed into a visualized 3D fault probability heatmap, binding the fault location to the vehicle's geometry. Rapid fault location is achieved through AR virtual-real overlay and eye-tracking / gesture fusion interaction. Simultaneously, a low-latency multi-user 3D spatial synchronization channel is built, supporting real-time annotation by remote experts and 3D repair animation guidance. This improves fault diagnosis efficiency in terms of visualization, interactivity, and collaboration, reducing reliance on human experience and expertise, and achieving efficient and immersive vehicle fault diagnosis and remote collaborative repair.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] Automotive fault diagnosis methods based on 3DGS modeling and human-computer interaction include:
[0008] Collect multi-view RGB images, point cloud data, real-time data streams, and historical fault records, and preprocess them to generate a multimodal fusion dataset;
[0009] A component tag library is constructed. Based on a multimodal fusion dataset and 3DGS algorithm, a 3D model of the vehicle is reconstructed. After semantic segmentation, the components are bound to the component tag library, and a 3DGS digital twin is generated in an optimized manner.
[0010] Based on the 3DGS digital twin, the weighted fault probability is calculated by combining real-time data stream and historical fault records, a fault probability distribution matrix is constructed, and a three-dimensional fault probability heat map is generated by using the point-by-point attribute assignment mechanism of 3DGS Gaussian points, and a three-dimensional fault diagnosis scenario is output.
[0011] A spatial positioning coordinate system is established, and the three-dimensional fault diagnosis scene is superimposed on the real vehicle to analyze the operation instructions and locate the fault location. A multi-user three-dimensional spatial synchronization channel is constructed to synchronize the three-dimensional fault diagnosis scene to the remote terminal in real time, and provide step-by-step guidance with 3D animation locally to generate a fault diagnosis feedback report.
[0012] Specifically, the steps for reconstructing a 3D model of a vehicle include:
[0013] A component label library is constructed, and sparse point clouds are generated from multi-view RGB images using motion recovery structures as initialization seeds.
[0014] The vehicle scene is represented as a set of anisotropic three-dimensional Gaussian functions, and the position mean, covariance matrix, opacity and spherical harmonic coefficient of the Gaussian points are defined.
[0015] The Gaussian points are projected onto the imaging plane of the corresponding multi-view RGB images from each camera. The rendered image generated by the projection is compared pixel by pixel with the corresponding multi-view RGB image. The projection error is obtained by weighted summation of the mean absolute value error and the structural similarity loss.
[0016] An adaptive density control strategy is adopted for Gaussian point splitting and pruning, and the original LiDAR point cloud is used as a geometric prior constraint to iteratively optimize and reconstruct the vehicle's 3D model.
[0017] Specifically, the steps for generating a 3DGS digital twin include:
[0018] A graph convolutional neural network is used to perform pixel-level component semantic segmentation on point cloud data to obtain segmented point clouds, which are then associated with Gaussian points and bound to a component label library.
[0019] Determine the boundary region and perform bilateral weighted smoothing on the Gaussian points in the boundary region to generate an enhanced 3DGS model;
[0020] After performing hierarchical pruning on the enhanced 3DGS model, a three-level detail hierarchy structure of high, medium and low is constructed. The number of Gaussian points in each detail level is dynamically adjusted according to the viewing distance to generate a 3DGS digital twin.
[0021] Specifically, the steps for constructing the fault probability distribution matrix include:
[0022] Associate real-time data streams with corresponding components and establish a real-time parameter cache to store key parameters of each component.
[0023] Iterate through each fault rule in the historical fault records. If the key parameters meet the conditions of the fault rule, determine the fault probability contribution base based on the degree of deviation of the key parameters from the rule threshold and the actual proportion of the duration specified by the fault rule.
[0024] When multiple fault rules for the same component are triggered, if they are from the same fault source, the largest fault probability contribution base is selected as the initial fault probability; otherwise, the fault probability bases of multiple fault rules are weighted and fused to obtain the initial fault probability.
[0025] Determine the normal range and count the cumulative number of times key parameters exceed the normal range. Determine the correction factor and calculate the weighted failure probability based on the preliminary failure probability to construct the failure probability distribution matrix.
[0026] Specifically, the steps for generating a three-dimensional fault probability heatmap include:
[0027] Traverse all Gaussian points corresponding to each component in the 3DGS digital twin, obtain the RGB three color components through linear interpolation based on the weighted failure probability, and adjust the transparency.
[0028] Attributes are assigned through ID mapping, and the spatial coordinates of Gaussian points mapped by ID are calibrated based on the global spatial coordinate system of the 3DGS digital twin.
[0029] By using rendering pipeline level blending and same-channel texture synthesis mechanism, the fault probability thermal rendering layer and the vehicle body structure rendering layer are blended in the same rendering channel to obtain the blended Gaussian point rendering attribute set.
[0030] Based on semantic boundary constraints of components, the thermal transition region between Gaussian points of different components is smoothed to generate a three-dimensional fault probability heat map.
[0031] Specifically, the steps for outputting a 3D fault diagnosis scenario include:
[0032] Construct a first buffer and a second buffer for the Gaussian point rendering attribute set, which are used for rendering the current frame and receiving the attributes after the fault probability distribution matrix is updated, respectively;
[0033] Perform attribute updates and color mappings on the Gaussian points in the high, medium, and low detail levels at the current camera view distance;
[0034] Monitor the real-time data stream, recalculate the weighted failure probability of components whose parameters have changed, and update the color, transparency, and size parameters of the corresponding Gaussian point in the second buffer.
[0035] At the end of the rendering cycle, a millisecond-level synchronous update link between the 3D fault probability heatmap and the real-time data stream is established through buffer swapping.
[0036] The encapsulated interactive interface supports clicking on any thermal area to retrieve and display relevant information about components, retains the view control function to adapt to the real-time rendering frame rate requirements, and outputs a 3D fault diagnosis scene.
[0037] Specifically, the steps for establishing a spatial positioning coordinate system include:
[0038] Extract the set of key geometric feature points of the vehicle from the 3DGS digital twin and construct a prior database of vehicle feature points;
[0039] The real-time video stream is processed frame by frame. Key geometric feature points in the current frame are detected and descriptors are calculated. Hamming distance is calculated by combining the corresponding descriptors in the vehicle feature point prior database and a preliminary matching pair is established. The RANSAC algorithm is used to retain the successfully matched feature point pairs.
[0040] The initial camera pose is obtained by solving the extrinsic parameters of the mobile device relative to the three-dimensional coordinate system defined by the 3DGS digital twin based on the feature point pairs. The optimized camera pose is obtained by optimizing the initial camera pose using the projection error.
[0041] A local coordinate system for the vehicle is established, and the optimized camera pose is transformed to the local coordinate system for the establishment of a spatial positioning coordinate system.
[0042] Specifically, the steps for locating the faulty part include:
[0043] The coordinates of Gaussian points in the 3D fault diagnosis scene are transformed to the spatial positioning coordinate system. The projection matrix is calculated based on the real-time pose data of the mobile device to generate a virtual rendering image.
[0044] A virtual-real fusion rendering pipeline is constructed, using the acquired real vehicle video stream as the background layer and the virtual rendered image as the foreground layer for layer fusion.
[0045] Calculate the actual depth value corresponding to each virtual pixel, and perform parallax offset correction on the virtual rendered image in combination with the display parameters of the mobile device;
[0046] The system acquires and parses user operation commands, converts them into component selection commands, view control commands, transparency commands, and sectioning commands, and outputs the user operation command set and the target component.
[0047] Based on the fault distribution data of the 3D fault probability heatmap and the user operation command set, the target component selected by the user is matched to locate the fault location.
[0048] Specifically, the steps for generating a fault diagnosis feedback report include:
[0049] Construct a multi-user three-dimensional spatial synchronization channel and divide the data corresponding to the three-dimensional fault diagnosis scenario into static data and dynamic data;
[0050] Static data is transmitted during the initial connection, and only dynamic data is transmitted thereafter. The dynamic attributes of each Gaussian point are encoded to generate synchronization data packets and transmitted via the UDP protocol.
[0051] Perform full operation on the 3D fault diagnosis scene and transmit operation commands to the local terminal in real time; add 3D annotations on the remote terminal, encode them as Gaussian point attributes, transmit them to the local terminal, and display them persistently.
[0052] AR devices instantly sense and overlay remote guidance information, and retrieve standard maintenance procedures based on the target component and the fault probability distribution matrix.
[0053] 3D animations of maintenance operations are generated using 3DGS digital twins, and then overlaid on AR devices through a virtual-real fusion rendering pipeline to form a fault diagnosis feedback report.
[0054] The automotive fault diagnosis system based on 3DGS modeling and human-computer interaction includes: a multi-source data fusion module, a digital twin modeling module, a 3D fault diagnosis module, and a collaborative interaction module.
[0055] The multi-source data fusion module is used to collect multi-view RGB images, point cloud data, real-time data streams and historical fault records, and preprocess them to generate a multimodal fusion dataset.
[0056] The digital twin modeling module is used to build a component tag library, reconstruct the vehicle's three-dimensional model based on a multimodal fusion dataset and the 3DGS algorithm, and bind the components to the component tag library after semantic segmentation to optimize the generation of a 3DGS digital twin.
[0057] The three-dimensional fault diagnosis module is used to calculate the weighted fault probability based on the 3DGS digital twin, combined with real-time data stream and historical fault records, to construct a fault probability distribution matrix, and to generate a three-dimensional fault probability heat map using the point-by-point attribute assignment mechanism of 3DGS Gaussian points, and output a three-dimensional fault diagnosis scenario.
[0058] The collaborative interaction module is used to establish a spatial positioning coordinate system, overlay the three-dimensional fault diagnosis scene onto the real vehicle, parse the operation instructions and locate the fault location; construct a multi-user three-dimensional spatial synchronization channel to synchronize the three-dimensional fault diagnosis scene to the remote terminal in real time, provide step-by-step guidance locally with 3D animation, and generate a fault diagnosis feedback report.
[0059] The beneficial effects of this invention are:
[0060] 1. This invention addresses the shortcomings of traditional diagnostic methods based on text, two-dimensional charts, or single sensor data, which cannot intuitively display the three-dimensional spatial distribution of faults and require manual correlation between data and component locations. It employs a 3DGS algorithm to generate a lightweight 3DGS digital twin, combined with a Gaussian point-by-point attribute assignment mechanism, mapping the fault probability distribution matrix into a three-dimensional fault probability heatmap bound to the vehicle's geometry. This transforms fault risk from abstract data into a visualized three-dimensional space. A millisecond-level synchronous update link is established, and combined with virtual-real fusion overlay technology, fault locations are mapped one-to-one with real vehicle components. Repair personnel can intuitively view the severity and spatial distribution of faults. By calculating weighted fault probabilities and filtering invalid data, the invention eliminates the need for manual analysis of massive amounts of sensor data, shortening fault location time, reducing reliance on repair personnel experience, and improving fault diagnosis efficiency and accuracy.
[0061] 2. This invention addresses the limitations of traditional video call remote guidance, such as limited perspective, abstract descriptions, and lag in interaction. It constructs a multi-user three-dimensional spatial synchronization channel based on 3DGS digital twins. Through an incremental transmission strategy that separates static and dynamic data and Gaussian point dynamic attribute encoding technology, it achieves real-time synchronization of three-dimensional fault diagnosis scenarios in low-bandwidth environments. Remote experts can perform full-dimensional operations on the scene, such as rotation, sectioning, and scaling. The added three-dimensional annotations are directly encoded as Gaussian point attributes. Local AR devices instantly perceive remote guidance information and overlay it for display, achieving WYSIWYG remote guidance. Combined with a natural interaction method that integrates eye-tracking and gestures and the generated step-by-step repair 3D animation, it improves the accuracy and efficiency of collaborative diagnosis and remote repair guidance. Attached Figure Description
[0062] Figure 1 This is a schematic diagram of a vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction.
[0063] Figure 2 This is a flowchart of the process for generating a 3DGS digital twin in this invention;
[0064] Figure 3 This is a flowchart of the process for generating a three-dimensional fault probability heatmap in this invention;
[0065] Figure 4 This is a flowchart illustrating the process of locating the faulty part in this invention;
[0066] Figure 5 This is a structural diagram of an automotive fault diagnosis system based on 3DGS modeling and human-computer interaction. Detailed Implementation
[0067] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0068] Example 1
[0069] refer to Figures 1 to 4 As shown, this embodiment introduces a vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction, including the following steps:
[0070] Multi-view RGB images and point cloud data of the vehicle are acquired through a visual camera array and LiDAR deployed around the vehicle. Simultaneously, real-time data streams covering the powertrain, chassis, and electrical systems, as well as historical fault records from a maintenance history database, are collected via the onboard OBD-II interface. The acquired multi-source data undergoes preprocessing via a timestamp synchronization mechanism. This includes unifying the multi-source data to the same timestamp and spatial coordinate system through spatiotemporal alignment, correcting lens distortion in the multi-view RGB images using distortion correction, and eliminating the impact of lighting variations on the multi-view RGB images through white balance. Sensor physical attributes and historical fault labels are processed into structured data packets to generate a spatiotemporally consistent multimodal fusion dataset. The multi-source data includes multi-view RGB images, point cloud data, real-time data streams, and historical fault records.
[0071] A component tag library is constructed. Based on multi-view RGB images and point cloud data from a multimodal fusion dataset, a 3DGS algorithm is used to reconstruct the vehicle's 3D model. A graph convolutional neural network is used to perform pixel-level component semantic segmentation on the point cloud data. Each segmented component is bound to a unique identifier, geometric constraint, and physical parameter from the component tag library. Through hierarchical pruning and detail rendering optimization of Gaussian point clouds, redundant background Gaussian points are removed and the rendering accuracy is dynamically adjusted according to the viewing distance to generate a lightweight 3DGS digital twin that can be rendered in real time. This is used to dynamically map the real-time physical state of the vehicle's actual components.
[0072] Based on a 3DGS digital twin, and combining sensor parameters from real-time data streams with fault rules from historical fault records, the weighted fault probability of each component is calculated in real-time through comparison of preset fault feature thresholds and frequency weighting. A fault probability distribution matrix is constructed. Using the point-by-point attribute assignment mechanism of 3DGS Gaussian points, the weighted fault probabilities in the fault probability distribution matrix are mapped to visual parameters, including the RGB color, transparency, and size parameters of the corresponding Gaussian points of the components. A 3D fault probability heatmap is generated, aligned with the coordinates of the 3DGS digital twin and fused with the rendering pipeline. A millisecond-level synchronous update link between the 3D fault probability heatmap and the real-time data stream is established, outputting a dynamic and interactive 3D fault diagnosis scene. The sensor parameters include voltage, temperature, and rotational speed.
[0073] A centimeter-level spatial positioning coordinate system is established based on visual SLAM. The 3D fault diagnosis scene is overlaid on the real vehicle through a mobile device. Eye-tracking ray detection and gesture recognition engine analyze the user's operation commands for specific faulty components, triggering transparent masking of the vehicle shell and one-shot perspective switching to locate the faulty part. At the same time, a multi-user 3D spatial synchronization channel is built to synchronize the complete local 3D fault diagnosis scene to the remote terminal in real time. Remote experts can view and annotate the scene in 3D by rotating, sectioning, and zooming. The 3D annotation is configured with persistent arrows, annotations, and voice labels in 3D space. Local staff can perceive remote guidance information in real time through AR devices and see the step-by-step repair operation demonstrated by 3D animation. Finally, a fault diagnosis feedback report is generated, which fundamentally overcomes the problems of abstract flat video communication and poor real-time interaction in traditional remote guidance.
[0074] Specifically, the steps for generating a lightweight 3DGS digital twin that can be rendered in real time include:
[0075] A component tag library is constructed based on the models and physical property parameters of various automotive parts. Specifically, a component list is extracted from vehicle design drawings, a unique identifier is assigned to each component, physical parameters are entered, geometric constraints are defined, and a version number is set to form the component tag library. Based on multi-view RGB images and point cloud data from a multimodal fusion dataset, a 3DGS algorithm is used to reconstruct the vehicle's 3D model. The component tag library adopts a relational data table structure. Each record contains the component's unique identifier, component name, vehicle model adaptation code, and version number. Physical parameters include, but are not limited to, material and rated operating parameters. For example, materials include cast iron and aluminum alloy, and rated operating parameters include temperature threshold and speed range. Geometric constraints are defined in JSON format, including the component's bounding box in the vehicle's local coordinate system, a list of key feature point coordinates, and the allowable assembly tolerance range.
[0076] The 3DGS algorithm reconstruction process includes: generating sparse point clouds from multi-view RGB images using motion recovery structures as initialization seeds; representing the vehicle scene as an anisotropic set of three-dimensional Gaussian functions; defining the position mean, covariance matrix, opacity, and spherical harmonic coefficients for each Gaussian point; projecting the Gaussian points onto the imaging plane of the corresponding multi-view RGB images from each camera using a differentiable rasterizer; comparing the generated rendered image with the corresponding multi-view RGB image pixel by pixel; calculating the mean absolute error and structural similarity loss respectively; setting equal weights based on the principle of balancing rendering fidelity and structural consistency; weighted summing of the mean absolute error and structural similarity loss to obtain the projection error; and employing an adaptive density control strategy for Gaussian point splitting. The process involves pruning and using the original LiDAR point cloud as a geometric prior. Specifically, it includes: constructing a KD tree from the original LiDAR point cloud; calculating the Euclidean distance from each Gaussian point to its nearest neighbor LiDAR point; sorting all Euclidean distances in ascending order and selecting the minimum Euclidean distance; setting a distance threshold based on the average point spacing of the LiDAR point cloud, with a value of twice the average spacing (e.g., 2 cm); if the minimum Euclidean distance is greater than the distance threshold, the corresponding Gaussian point is pruned and deleted; otherwise, the corresponding Gaussian point is retained and allowed to participate in subsequent iterative optimization. This limits the location of the Gaussian points to the spatial distribution range of the LiDAR point cloud and constrains the Gaussian points to be within the actual vehicle geometric contour. After iterative optimization, a 3D vehicle model containing position, covariance, opacity, and color features is reconstructed.
[0077] The specific calculation process of mean absolute error and structural similarity loss includes: for each rendered image and the corresponding multi-view RGB image, extract the pixel values of the two in the three color channels of red, green and blue for each pixel; for each color channel of each pixel, calculate the absolute difference between the pixel value of the rendered image and the pixel value of the corresponding multi-view RGB image; accumulate the absolute differences of all channels of all pixels, and then divide by the product of the total number of pixels and the number of color channels to obtain the mean absolute error under the current view; take the arithmetic mean of the mean absolute errors of all views to obtain the global mean absolute error.
[0078] Using a sliding window approach, the rendered image and the corresponding multi-view RGB image are divided into multiple partially overlapping local image blocks. The size of each local image block is set to n×n pixels, where n can be 10. The mean brightness, variance brightness, and covariance of each pair of local image blocks are calculated. The brightness similarity of each pair of local image blocks is calculated based on the mean brightness, the contrast similarity of each pair of local image blocks is calculated based on the variance brightness, and the structural similarity of each pair of local image blocks is calculated based on the covariance. According to the principle that brightness, contrast, and structure are equally important to the perceived quality of the image, three weight coefficients with equal values are set. For example, the weight coefficients of brightness similarity, contrast similarity, and structural similarity are all set to 1 / 3, and the sum of the three weight coefficients is 1, ensuring that the contribution of the three dimensions to the final similarity evaluation is completely equal.
[0079] Combining the aforementioned weighting coefficients, brightness similarity, contrast similarity, and structural similarity are weighted and fused to obtain the structural similarity index of local image patches. The arithmetic mean of the structural similarity indices of all local image patches is then taken to obtain the overall structural similarity index from the current viewpoint. Subtracting the overall structural similarity index from 1 yields the structural similarity loss from the current viewpoint. The arithmetic mean of the structural similarity losses from all viewpoints is then taken to obtain the global structural similarity loss. Brightness similarity is calculated by comparing the mean brightness of each pair of local image patches, with a value ranging from 0 to 1. The brightness similarity reaches its maximum value of 1 when the two mean brightness values are equal; the greater the difference between the two mean brightness values, the smaller the brightness similarity value. Contrast similarity is calculated by comparing the luminance variance of each pair of local image blocks, with a value ranging from 0 to 1. The luminance variance reflects the contrast level of pixel brightness changes within a local image block. When the two luminance variances are exactly equal, the contrast similarity reaches its maximum value of 1. The larger the difference between the two luminance variances, the smaller the contrast similarity value. Structural similarity is calculated by the covariance between each pair of local image blocks, with a value ranging from 0 to 1. The covariance reflects the correlation of pixel brightness change patterns in each pair of local image blocks, i.e., the degree of structural similarity. When the pixel brightness change patterns of each pair of local image blocks are consistent, the structural similarity reaches its maximum value of 1. The larger the difference between the two covariances, the smaller the structural similarity value.
[0080] The adaptive density control strategy specifically includes: setting a gradient threshold based on the global distribution statistics of the multi-view projection gradient, with the value being the 90th quantile of the multi-view projection gradient distribution, such as 0.05; splitting Gaussian points whose cumulative gradient is greater than the gradient threshold; copying the split Gaussian points and reducing their scale to half of the original scale; and setting a judgment threshold based on the statistical distribution of opacity, with the value being the 10th quantile of the statistical distribution, such as 0.005; and directly pruning and removing Gaussian points whose opacity is less than the judgment threshold.
[0081] A known graph convolutional neural network (Graph Convolutional Neural Network) is used to perform pixel-level component semantic segmentation on point cloud data in a multimodal fusion dataset, resulting in segmented point clouds. The Graph Convolutional Neural Network employs three layers of graph convolutions, each with a 3×3 kernel size and ReLU activation function. The training set is a multimodal fusion dataset of vehicle parts, and the annotation standard follows the general annotation specifications for semantic segmentation of industrial parts. Component categories, boundaries, and attribute labels are annotated point-by-point to ensure annotation accuracy and consistency. The training process uses a cross-entropy loss function combined with an Adam optimizer with an initial learning rate of 1e-3 and a weight decay coefficient of 1e-5, training until the loss decreases by less than 1% for 10 consecutive rounds. Nearest neighbor search is used to associate Gaussian points with the segmented point clouds, binding them to unique identifiers and geometric constraints from the component label library. And physical parameters; determine the boundary region based on the neighborhood consistency feature of semantic tags, that is, the intersection region of point clouds with different semantic tags, and perform bilateral weighted smoothing on the Gaussian points in the boundary region to eliminate semantic jaggedness, and generate an enhanced 3DGS model containing semantic tags and physical attributes; the specific process of nearest neighbor search includes: based on the three-dimensional spatial coordinates of Gaussian points and segmented point clouds, calculate the difference of the three coordinate components x, y, z respectively, square the difference, sum and then take the square root to obtain the Euclidean distance between Gaussian points and segmented point clouds, calculate the Euclidean distance between each Gaussian point and all segmented point clouds through the above method, select the point cloud with the closest Euclidean distance, assign the semantic tag of the point cloud with the closest Euclidean distance to the corresponding Gaussian point, and realize the association and binding of Gaussian points with component semantics;
[0082] The specific process of bilateral weighted smoothing includes: for each Gaussian point in the boundary region, selecting Gaussian points in the neighborhood, constructing a bilateral weight matrix containing spatial distance weights and semantic consistency weights, and using the Gaussian points in the neighborhood as rows and the spatial distance weights and semantic consistency weights as columns, performing a weighted average on the color and opacity parameters of the Gaussian points to weaken parameter abrupt changes in different semantic regions and eliminate semantic jaggedness; wherein, the weighted average is performed by traversing each Gaussian point in the neighborhood, multiplying the color and opacity parameters of each Gaussian point by the corresponding bilateral weights in the bilateral weight matrix, summing all the product results and dividing by the total weight to obtain the smoothed color and opacity parameters of the current Gaussian point;
[0083] A hierarchical pruning strategy is applied to the enhanced 3DGS model, employing a three-level pruning approach: The first level is background removal, deleting Gaussian points located outside the vehicle bounding box extension region or those semantically labeled as "ground" or "sky," representing background categories. The second level is structural redundancy pruning, extracting the maximum and minimum eigenvalues of the covariance matrix of local Gaussian points within the same component, calculating the ratio of the maximum to minimum eigenvalues to obtain the sphericity within the same component, and setting a similarity threshold of 0.85 based on the spatial distribution uniformity of Gaussian points within the component, merging similar Gaussian points with sphericity greater than the similarity threshold. The third level is visual importance pruning, statistically analyzing the gradient magnitude of each pixel in multi-view RGB images, setting an amplitude threshold of 30% of the global average gradient magnitude based on the image texture discrimination accuracy standard, and pruning the gradient magnitude... Regions with a gradient value less than the gradient threshold are identified as low-texture regions. The reconstruction loss contribution is calculated based on the change in projection error before and after removing the target Gaussian point within the low-texture region. Specifically, the total projection error before and after removing the target Gaussian point is calculated separately, and the reconstruction loss contribution of the target Gaussian point is obtained by subtracting the total projection error before removal from the total projection error after removal. A contribution threshold of 0.01 is set based on the reconstruction error tolerance range of the rendered image. If the position of the Gaussian point projected onto the multi-view RGB image is in the transition region of different semantic labels, i.e., there are two or more different semantic labels in the neighborhood window, it is identified as a key structure Gaussian point and retained. Otherwise, redundant Gaussian points in the low-texture region whose reconstruction loss contribution is less than the contribution threshold are deleted, and key geometric features of the vehicle are retained, including but not limited to component contours, surface curvature, and edge corners.
[0084] Based on the camera's viewing distance range and the importance weight of components, detail rendering is optimized, constructing a three-level detail hierarchy structure of high, medium, and low. The viewing distance is determined by the 3D spatial straight-line distance between the camera's imaging center point and the center of the component's bounding box. The number of Gaussian points at each level is dynamically adjusted according to the viewing distance, including: setting a first viewing distance threshold and a second viewing distance threshold based on human visual sensitivity and typical vehicle component dimensions, with values of 5 meters and 15 meters respectively; when the viewing distance is less than the first viewing distance threshold, 90% of Gaussian points are retained for the high detail level; when the viewing distance is between the first and second viewing distance thresholds, the Gaussian points at the medium detail level are clustered, reducing the number of Gaussian points by half; when the viewing distance is greater than the second viewing distance threshold, 20% of Gaussian points are retained for the low detail level. Finally, a lightweight 3DGS digital twin that can be rendered in real time is generated, used to dynamically map the real-time physical state of real vehicle components, including parameters such as temperature, vibration acceleration, and operating voltage.
[0085] Specifically, the steps for outputting a dynamic and interactive 3D fault diagnosis scenario include:
[0086] According to vehicle diagnostic standards, real-time data streams are associated with one or more corresponding components, and a real-time parameter cache is established to store the key parameters of each component at the current moment. The key parameters are real-time quantitative monitoring indicators that can characterize the operating status of the components.
[0087] The system iterates through each fault rule in the historical fault records of the maintenance history database. Each fault rule is constructed using the Apriori algorithm by performing feature association mining on the historical fault samples corresponding to the historical fault records. This includes a temperature threshold of 95 degrees Celsius, a time limit of 5 seconds, a voltage rating of 20%, a fluctuation threshold of 15%, and logical combinations and duration constraints. The system determines whether the current key parameters meet the conditions of the fault rule. These conditions include, but are not limited to, temperatures exceeding the 95-degree Celsius threshold and durations exceeding the 5-second time limit, and voltages falling below the 20% voltage rating and speed fluctuations exceeding the 15% fluctuation threshold. If the conditions are met, piecewise linear interpolation normalization is used to determine the fault probability contribution base based on the deviation of the current key parameters from the rule thresholds and the actual proportion of the duration specified by the fault rule. The number, ranging from 0 to 1, and the corresponding components are recorded. The process of determining the base for the contribution of fault probability includes: calculating the degree of deviation of each key parameter from the rule threshold and the deviation ratio of the corresponding rule threshold, and taking the maximum value of the deviation ratio as the comprehensive deviation degree; calculating the ratio of the actual duration of the current key parameter exceeding the corresponding rule threshold to the duration specified by the fault rule as the duration ratio; taking the arithmetic mean of the comprehensive deviation degree and the duration ratio to obtain the comprehensive triggering amount; using the rule triggering critical point as the segment node, using two-segment linear interpolation to normalize and map the comprehensive triggering amount to the 0~1 interval to obtain the base for the contribution of fault probability; wherein, the rule threshold includes the temperature threshold, the voltage rating value, and the fluctuation threshold; the degree of deviation includes the temperature difference between the temperature and the temperature threshold, the voltage difference between the voltage and the voltage rating value, and the fluctuation difference between the speed fluctuation and the fluctuation threshold;
[0088] For the same component, when multiple fault rules are triggered, if the multiple triggering rules are different observation dimensions of the same fault source, the largest fault probability contribution base is selected as the preliminary fault probability to avoid overestimation caused by the superposition of multiple dimensions of the same fault source; if the multiple triggering rules correspond to different fault sources and reflect composite fault information, the fault probability contribution bases of the multiple fault rules are weighted and fused to obtain the preliminary fault probability, preserving multi-dimensional fault characteristics and avoiding fault omissions and misjudgments; the weight of each fault rule is determined according to the frequency of fault occurrence of each fault rule in historical fault records; the selection criterion for the largest fault probability contribution base is that multiple fault rules of the same component originate from the same fault, and each fault probability contribution base is a reflection of the same fault source in different observation dimensions, and they are highly correlated rather than independent;
[0089] The system counts the cumulative number of times each component's key parameters exceed the normal range within a fixed time period. A higher number of such occurrences indicates repeated abnormalities in the component, thus increasing the confidence level of the fault. The normal range is determined based on the normal fluctuation range of historical statistical data, specifically by calculating the average value of the corresponding parameters for each component under historical normal operating conditions. with standard deviation ,definition[ The normal range is defined as follows: a correction coefficient is set based on the cumulative number of occurrences within the normal range. For example, no weighting is applied when the cumulative number of occurrences is 0, 10% weighting is applied when the cumulative number of occurrences is 1 to 3, and 20% weighting is applied when the cumulative number of occurrences is 4 or more. The value 1 is added to the correction coefficient to obtain the correction factor. The product of the correction factor and the initial failure probability is calculated to obtain the weighted failure probability, and the weighted failure probability is ensured not to exceed 100%. In this embodiment, it is assumed that the fixed time is the most recent hour.
[0090] A failure probability distribution matrix is constructed based on the weighted failure probabilities of all components, with component IDs as rows and weighted failure probabilities as columns. The elements of the failure probability distribution matrix correspond one-to-one with the components and are accompanied by timestamps for subsequent heatmap mapping.
[0091] Iterate through all Gaussian points corresponding to each component in the 3DGS digital twin. Based on the weighted failure probability of each component in the failure probability distribution matrix, obtain the three color components RGB through linear interpolation. When the weighted failure probability of a component is close to 0, it is greenish; when it is close to 0.5, it is yellowish; and when it is close to 1, it is reddish. Adjust the transparency, including setting the opacity of each Gaussian point to the weighted failure probability. Keep the basic vehicle structure opaque. When the weighted failure probability is 0, the thermal layer is completely transparent. When the weighted failure probability is 1, the thermal layer is opaque. When the weighted failure probability is between 0 and 1, the opacity transitions linearly.
[0092] Since the failure probability distribution matrix is indexed according to the component ID, and each Gaussian point in the 3DGS digital twin has been pre-bound with the unique identifier of its component, attribute values are assigned through ID mapping to ensure that thermal color is aligned with geometric position.
[0093] Using the global spatial coordinate system of the 3DGS digital twin as a reference, the Gaussian points mapped by ID are calibrated to ensure that the thermal color pixels are highly consistent with the geometric positions of the real vehicle parts. Through rendering pipeline layer blending and same-channel texture synthesis mechanism, the fault probability thermal rendering layer and the vehicle structure rendering layer are fused in the same rendering channel to obtain the fused Gaussian point rendering attribute set without adding a new independent rendering channel, thus maximizing the model's lightweight nature and real-time rendering performance. Based on the semantic boundary constraints of the parts, the thermal transition areas between Gaussian points of different parts are smoothed to avoid abrupt changes in thermal color or rendering jagged edges across parts, ensuring the continuity of the visualization. Finally, a 3D fault probability heatmap that is aligned with the coordinates of the 3DGS digital twin and fused with the rendering pipeline is generated.
[0094] To avoid screen tearing or attribute flickering during the update of the 3D fault probability heatmap, two sets of attribute buffers are constructed for the Gaussian point rendering attribute set. The first buffer is used for rendering the current frame, and the second buffer is used to receive the attributes after the fault probability distribution matrix is updated.
[0095] The above attribute update and color mapping are performed on the Gaussian points in the high, medium and low detail levels actually called under the current camera view distance. The update operation directly modifies the Gaussian point rendering attribute set in the attribute buffer that is currently in the writing state. Only the Gaussian points corresponding to the parts whose weighted failure probability has changed are updated. For Gaussian points in the low detail level, the attribute update frequency is reduced, for example, once every 10 frames, in order to balance the computing load.
[0096] Based on a 3D fault probability heatmap, the system continuously monitors the real-time data stream output at 100Hz from the vehicle's OBD-II interface. Upon receiving a new data frame, the system recalculates the weighted fault probability for components whose parameters have changed, and updates the color, transparency, and size parameters of the corresponding Gaussian point in the second buffer. At the end of each rendering cycle, the roles of the two attribute buffers are swapped to ensure that the new weighted fault probability takes effect in the next rendering and that the heatmap update delay remains stable within 10 milliseconds. This establishes a millisecond-level synchronous update link between the 3D fault probability heatmap and the real-time data stream. Based on the dynamically updated heatmap, an interactive interface based on component ID and bounding box detection is further encapsulated. This allows users to automatically retrieve and display component-related information when clicking on any heatmap area, including but not limited to component parameters, real-time sensor values, weighted fault probability, and historical fault information. The system also retains scene scaling, rotation, and translation control functions, adapting to the real-time rendering frame rate requirements of AR devices and mobile devices, with a minimum real-time rendering frame rate of 30 frames per second.
[0097] After integrating dynamic heat map updates, millisecond-level synchronization links, and interactive functions, the final output is a dynamic and interactive 3D fault diagnosis scene.
[0098] Specifically, the steps for generating a fault diagnosis feedback report include:
[0099] A set of key geometric feature points for the vehicle is extracted from the 3DGS digital twin, including component contours, surface curvature, and edge corners. Each feature point is accompanied by three-dimensional spatial coordinates and a corresponding Gaussian point index. The descriptor uses a 256-bit binary ORB descriptor. The database storage structure is a hash table, with the feature point index as the key and the three-dimensional spatial coordinates, Gaussian point index, and descriptor as the value, constructing a priori database of vehicle feature points. Real-time video streams captured by mobile devices are processed frame by frame, and the ORB feature extraction algorithm is used to detect key geometric feature points in the current frame, calculating the descriptor for each key geometric feature point. The extraction criteria for key geometric feature points include: selecting vertices with surface curvature greater than a curvature threshold, corner points at edge corners, and concave and convex points on component contours. The curvature threshold is determined by adding twice the standard deviation to the mean curvature of all vertices of the vehicle components, with a value of 0.1.
[0100] The system matches the currently detected key geometric feature points with the vehicle feature point prior database. Hamming distance is calculated using the descriptors of the key geometric feature points in the current frame and their corresponding descriptors in the vehicle feature point prior database. Preliminary matching pairs are established using the FLANN matching algorithm, and mismatched points are eliminated using the RANSAC algorithm, retaining only successfully matched feature point pairs. Based on the successfully matched feature point pairs, the PnP algorithm is used to solve for the extrinsic parameters of the mobile device relative to the 3D coordinate system defined by the 3DGS digital twin, including the rotation matrix and translation vector, to obtain the initial camera pose. The initial camera pose is then optimized using the Gaussian point projection error of the 3DGS digital twin to obtain the optimized camera pose. The optimization process involves projecting Gaussian points from the 3DGS digital twin onto the current frame's imaging plane, calculating the pixel error between the projected points and the corresponding image feature points, constructing a nonlinear optimization objective function, and iteratively optimizing the initial camera pose using the Levenberg-Marquardt method until the projection error is less than the error threshold. The error threshold is determined based on the 95th quantile of the statistical distribution of the feature point reprojection residuals, and is set to 1.5 pixels.
[0101] Establish a local coordinate system for the vehicle with the vehicle's center of mass as the origin, the longitudinal direction of the vehicle body as the X-axis, the lateral direction as the Y-axis, and the vertical upward direction as the Z-axis. Transform the optimized camera pose to the local coordinate system for the vehicle to form a spatial positioning coordinate system at the centimeter level.
[0102] The mobile device drives the 3D fault diagnosis scene to be aligned and superimposed with the real vehicle coordinates. The origin and coordinate axis directions of the global coordinate system of the 3D fault diagnosis scene are extracted and aligned with the spatial positioning coordinate system. The coordinate transformation matrix transforms the coordinates of all Gaussian points in the 3D fault diagnosis scene to the spatial positioning coordinate system. Based on the real-time pose data of the mobile device, the projection matrix of the 3D fault diagnosis scene on the camera imaging plane is calculated, including the camera intrinsic and extrinsic parameters. The coordinates of Gaussian points in 3D space are converted into pixel coordinates on the 2D imaging plane, realizing the projection mapping of the 3D fault diagnosis scene to the display screen of the mobile device.
[0103] A differentiable rasterizer is used to project Gaussian points in the 3D fault diagnosis scene, transformed to a spatial positioning coordinate system, onto the display screen of a mobile device to generate a virtual rendering image. A virtual-real fusion rendering pipeline is constructed, using the real vehicle video stream captured by the mobile device as the background layer and the virtual rendering image as the foreground layer, and alpha blending technology is used for layer fusion. In particular, the transparency of the vehicle structure part in the virtual rendering image is set to 0.3, and the transparency of the 3D fault probability heatmap part is set to a value proportional to the weighted fault probability.
[0104] To eliminate perspective deviation during the virtual-real overlay process, parallax correction is performed on the virtual rendered image. Based on the optimized camera pose and spatial positioning coordinate system, the real depth value corresponding to each virtual pixel is calculated. Combined with the display parameters of the mobile device, the parallax offset correction is performed on the virtual rendered image pixel by pixel.
[0105] This study utilizes an AR device to acquire user operation commands. The commands are then fused and analyzed using the AR device's built-in eye-tracking ray detection and gesture recognition engine. The 120Hz sampled eye-tracking ray data is used for pupil center localization, eye-tracking ray direction calculation based on the pupil-corneal reflection vector method, and blink detection. A gaze point clustering algorithm is employed to cluster eye-tracking ray points within a continuous 100ms as gaze points, which are then converted into three-dimensional spatial coordinates through inverse projection transformation. After preprocessing 60Hz sampled hand depth images, deep learning-based hand keypoint detection technology is used to detect the three-dimensional coordinates of hand keypoints. A gesture classification algorithm based on hand keypoints is used to identify common operation gestures such as clicking, swiping, zooming, rotating, clenching fists, and slicing. The gesture classification algorithm uses a lightweight 3D convolutional neural network to classify the spatiotemporal sequence of hand keypoints.
[0106] An eye-tracking ray and gesture fusion interaction mechanism is established, introducing a dual-modal confidence-weighted fusion. Eye-tracking gaze confidence is obtained by mapping gaze duration, while gesture recognition confidence is the classification output probability. Based on the principle that gaze intent is more dominant than gesture action, a gaze weight of 0.6 is set. Since the sum of the gaze weight and gesture weight is 1, the gesture weight is 0.4. The weighted sum is used to obtain the fusion score. When the fusion score is greater than a threshold and the following behavioral rules are met simultaneously, a corresponding instruction is triggered: when the user's eye-tracking ray points to the surface of a target component for more than 300 milliseconds and simultaneously performs a click gesture, it is determined as component selection; when pointing to the background space and performing a swipe gesture, it is determined as viewpoint switching; when making a fist gesture, it is determined as viewpoint locking; and when performing a cutting gesture... The system determines that a sectioning operation is to be performed on the current gaze region. The identified commands are semantically parsed and converted into component selection commands, viewpoint control commands, transparency commands, and sectioning commands. A command filtering mechanism is then employed: a command buffer queue is set up, and the command types and fusion scores of three consecutive frames are checked for consistency. Only when the commands in the command buffer queue are completely consistent and the average fusion score is greater than the minimum execution threshold, the verified user operation command set and target component are output. The score threshold is determined based on the statistical mean of offline interaction tests and is set to 0.7. The minimum execution threshold is determined based on the results of multiple rounds of output stability tests and is set to 0.65. The average fusion score is obtained by calculating the arithmetic mean of the fusion scores of three consecutive frames in the command buffer queue.
[0107] Based on the fault distribution data from the 3D fault probability heatmap and the user operation command set, the system matches the target component selected by the user to locate the fault location, i.e., to lock down the vehicle parts with potential faults. This includes: when a component selection command is received, triggering intelligent transparent masking of the vehicle shell and switching between one-shot perspectives; retrieving the bounding box information of the target component in the 3DGS digital twin to determine the 3D spatial range of the target component; constructing a semantic segmentation mask for the vehicle shell, marking all Gaussian points with the semantic label "vehicle shell" as areas to be processed; calculating the shortest Euclidean distance from each Gaussian point of the vehicle shell to the surface of the bounding box of the target component using a 3D spatial distance calculation method, and calculating the dynamic transparency value of each Gaussian point of the vehicle shell using a distance-weighted transparency calculation method; using the current camera pose and the spatial position information of the target component, calculating the path from the current camera pose to the optimal observation pose 2 meters away from the center of the bounding box of the target component, with the line of sight pointing to the center and forming a 45-degree angle with the longitudinal direction of the vehicle body; using the Bezier curve interpolation method to generate a switching path including acceleration, constant speed, and deceleration segments, with a total switching time of 1 second, achieving a seamless one-shot perspective switching effect.
[0108] To support remote collaborative diagnosis, a multi-user 3D spatial synchronization channel supporting full remote expert operation is constructed. The data corresponding to the 3D fault diagnosis scenario is divided into static and dynamic data. Static data includes Gaussian point positions, covariance, and semantic labels, while dynamic data includes color, transparency, weighted fault probability, operation instructions, and annotation information. Initially, static data is transmitted in its entirety; subsequently, only the changed dynamic data is transmitted. The dynamic attributes of each Gaussian point are encoded as a 32-bit integer, including 8 bits for red, 8 bits for green, 8 bits for blue, 4 bits for transparency, and 4 bits for weighted fault probability. A synchronization data packet containing a timestamp, the index of the changed Gaussian point, encoded attributes, and operation instructions is generated every 100 milliseconds and transmitted using the UDP protocol. The remote terminal receives the data. After synchronizing data packets, the local Gaussian point attributes are sorted and updated according to the timestamp. This allows remote experts to perform full operations such as rotation, sectioning, and scaling on the 3D fault diagnosis scene from a remote terminal. The expert operation commands are transmitted to the local terminal in real time through a multi-user 3D space synchronization channel. When multiple users issue operation commands to the same Gaussian point or the same scene viewpoint in the same time window, it is determined to be a multi-user operation conflict. A priority scheduling strategy based on user roles is adopted to resolve the multi-user operation conflict. The operation priority of remote experts is higher than that of local staff. High-priority operations cover low-priority operations. For the covered low-priority operations, a scene state rollback mechanism is executed to restore the parameters corresponding to the scene state to the state before the low-priority operation was executed, and the operation log is recorded.
[0109] Remote experts can add persistent 3D annotations in the form of arrows, text annotations, and voice tags on remote terminals. The system encodes the 3D annotation information into Gaussian point attributes: arrow annotations generate red, high-opacity, clearly visible Gaussian points at the start and end points, storing direction and length information; text annotations generate white, high-opacity 3D contour Gaussian points; voice tags are converted into text and stored in the custom attribute fields of the corresponding Gaussian points; the encoded 3D annotation information is transmitted to the local terminal through a multi-user 3D space synchronization channel, added to the local 3DGS digital twin according to the Gaussian point index, and rendered together with the original Gaussian points to achieve persistent display of 3D annotations; the AR device worn by the local staff perceives and overlays all remote guidance information in real time through the virtual-real fusion rendering pipeline; when the user gazes at the 3D annotation for more than a fixed time, the corresponding voice tag is automatically played and the text annotation is displayed; in this embodiment, the fixed time is set to 500 milliseconds;
[0110] Based on the target component and fault probability distribution matrix, a matching standard maintenance process is retrieved from the maintenance history database. Each maintenance step is broken down into basic actions such as disassembling bolts, removing cover plates, and replacing components. A maintenance animation scene is constructed using a 3DGS digital twin, with the target component and related components set as animation objects and others as background objects. Keyframe animation technology is used to generate 3D animations of maintenance operations, setting keyframes containing position, rotation, and scaling information for the animation objects, and generating transition animations through interpolation calculations. Flashing red arrows, text prompts, and voice narration are added to the animations, and users can pause, play, fast forward, and rewind the animations using gestures. The 3D animations of maintenance operations are overlaid on the AR devices worn by local staff through a virtual-real fusion rendering pipeline, intuitively showing the specific location and method of the maintenance operation.
[0111] Integrate data from the entire fault diagnosis process, including but not limited to component IDs, weighted fault probabilities, historical fault records, operating instructions, and 3D annotation information, to ultimately generate a structured fault diagnosis feedback report.
[0112] Example 2
[0113] Please see Figure 5 Another embodiment of the present invention provides: an automotive fault diagnosis system based on 3DGS modeling and human-computer interaction, comprising: a multi-source data fusion module, a digital twin modeling module, a three-dimensional fault diagnosis module, and a collaborative interaction module;
[0114] The multi-source data fusion module is used to acquire multi-view RGB images and point cloud data of the vehicle through a visual camera array and a lidar, respectively. It also uses the vehicle OBD-II interface to acquire real-time data streams and historical fault records. After the multi-source data is preprocessed through a timestamp synchronization mechanism, a spatiotemporally consistent multimodal fusion dataset is generated.
[0115] The digital twin modeling module is used to build a component label library. Based on a multimodal fusion dataset, it uses the 3DGS algorithm to reconstruct the vehicle's 3D model. It uses a graph convolutional neural network to perform pixel-level component semantic segmentation on the point cloud data, binding unique identifiers, geometric constraints, and physical parameters to each component. Through hierarchical pruning and detail rendering optimization of Gaussian point clouds, redundant background Gaussian points are removed and rendering accuracy is dynamically adjusted to generate a 3DGS digital twin that dynamically maps the real-time physical state of the vehicle's actual components.
[0116] The 3D fault diagnosis module is used to calculate the weighted fault probability of each component in real time and construct a fault probability distribution matrix based on the 3DGS digital twin, combined with sensor parameters in real-time data stream and fault rules in historical fault records. It uses the point-by-point attribute assignment mechanism of 3DGS Gaussian points to map the weighted fault probability in the fault probability distribution matrix into visualization parameters, generate a 3D fault probability heat map, establish a millisecond-level synchronous update link between the 3D fault probability heat map and the real-time data stream, and output a 3D fault diagnosis scenario.
[0117] The collaborative interaction module is used to establish a spatial positioning coordinate system based on visual SLAM, overlay the 3D fault diagnosis scene with the real vehicle, and analyze the user's operation commands through eye-tracking ray detection and gesture recognition engine to locate the fault location; it also builds a multi-user 3D spatial synchronization channel to synchronize the 3D fault diagnosis scene to the remote terminal in real time, allowing remote experts to view and annotate the scene in 3D by rotating, sectioning, and zooming, while local staff can instantly perceive remote guidance information through AR devices and see step-by-step repair operations demonstrated by 3D animation, generating a fault diagnosis feedback report.
[0118] Working principle and effects:
[0119] Multi-view RGB images and point cloud data are acquired by a visual camera array and LiDAR deployed around the vehicle. Combined with real-time data streams from the vehicle's OBD-II interface and historical fault records, the spatiotemporal alignment and preprocessing of multi-source data are completed through a timestamp synchronization mechanism to generate a spatiotemporally consistent multimodal fusion dataset. This achieves standardized acquisition and unified representation of multi-source heterogeneous vehicle data, eliminating information gaps caused by asynchronous data and heterogeneous formats, and providing high-quality, highly synchronous data support for subsequent digital twin modeling and fault diagnosis.
[0120] A component tag library was constructed. Based on a multimodal fusion dataset, a 3D vehicle model was reconstructed using the 3DGS algorithm. Point cloud semantic segmentation was used to bind unique identifiers, geometric constraints, and physical parameters to each component. Through hierarchical pruning and rendering optimization of Gaussian point clouds, a lightweight and real-time renderable 3DGS digital twin was generated, dynamically mapping the real-time physical state of the vehicle's actual components. This solves the problems of insufficient accuracy, low rendering efficiency, and inability to synchronize physical states in traditional vehicle modeling. It constructs a 3D digital foundation synchronized with the full state of the real vehicle, providing a real-time, interactive 3D platform for fault visualization and diagnosis.
[0121] Based on a 3DGS digital twin, and combining real-time data stream sensor parameters with fault rules from historical fault records, the weighted fault probability of each component is calculated, and a fault probability distribution matrix is constructed. Utilizing the 3DGS Gaussian point-by-point attribute assignment mechanism, the weighted fault probability is mapped to visual parameters, generating a 3D fault probability heatmap. A synchronous update link with the real-time data stream is established, outputting a dynamic and interactive 3D fault diagnosis scenario. This transformation of weighted fault probability into intuitive 3D visualization information solves the problems of unintuitive results and lagging dynamic updates in traditional fault diagnosis, achieving a comprehensive, dynamic, and visual representation of vehicle faults, and improving the accuracy of fault location and the efficiency of fault diagnosis.
[0122] A spatial positioning coordinate system is established based on visual SLAM, overlaying the 3D fault diagnosis scene with the real vehicle; eye-tracking ray detection and gesture interaction are used to analyze user operation commands and quickly locate the fault location; a multi-user 3D spatial synchronous channel is constructed to support remote expert remote 3D annotation and local AR repair guidance, generating a fault diagnosis feedback report. This overcomes the limitations of traditional remote guidance planar video interaction, solving the problems of abstract remote communication, delayed interaction, and unintuitive repair guidance, achieving immersive collaborative interaction between fault diagnosis and repair guidance, fundamentally solving the problem of poor interaction immediacy.
[0123] Overall, through a four-layer modular architecture of multi-source data fusion, digital twin modeling, 3D fault diagnosis, and collaborative interaction, a vehicle fault intelligent diagnosis and remote collaborative maintenance system has been constructed. This system enables visualized diagnosis, full-domain dynamic representation, and immersive collaborative maintenance of vehicle faults, solving the technical problems of low efficiency, unintuitive fault location, and poor interactivity and immediacy in traditional vehicle fault diagnosis. It provides reliable and traceable full-process technical support for vehicle repair and maintenance.
[0124] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction, characterized in that, include: Collect multi-view RGB images, point cloud data, real-time data streams, and historical fault records, and preprocess them to generate a multimodal fusion dataset; A component tag library is constructed. Based on a multimodal fusion dataset and 3DGS algorithm, a 3D model of the vehicle is reconstructed. After semantic segmentation, the components are bound to the component tag library, and a 3DGS digital twin is generated in an optimized manner. Based on the 3DGS digital twin, the weighted fault probability is calculated by combining real-time data stream and historical fault records, a fault probability distribution matrix is constructed, and a three-dimensional fault probability heat map is generated by using the point-by-point attribute assignment mechanism of 3DGS Gaussian points, and a three-dimensional fault diagnosis scenario is output. A spatial positioning coordinate system is established, and the three-dimensional fault diagnosis scene is superimposed on the real vehicle to analyze the operation instructions and locate the fault location. A multi-user three-dimensional spatial synchronization channel is constructed to synchronize the three-dimensional fault diagnosis scene to the remote terminal in real time, and provide step-by-step guidance with 3D animation locally to generate a fault diagnosis feedback report.
2. The vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction according to claim 1, characterized in that, The specific steps for reconstructing a vehicle's 3D model include: A component label library is constructed, and sparse point clouds are generated from multi-view RGB images using motion recovery structures as initialization seeds. The vehicle scene is represented as a set of anisotropic three-dimensional Gaussian functions, and the position mean, covariance matrix, opacity and spherical harmonic coefficient of the Gaussian points are defined. The Gaussian points are projected onto the imaging plane of the corresponding multi-view RGB images from each camera. The rendered image generated by the projection is compared pixel by pixel with the corresponding multi-view RGB image. The projection error is obtained by weighted summation of the mean absolute value error and the structural similarity loss. An adaptive density control strategy is adopted for Gaussian point splitting and pruning, and the original LiDAR point cloud is used as a geometric prior constraint to iteratively optimize and reconstruct the vehicle's 3D model.
3. The vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction according to claim 2, characterized in that, The specific steps for generating a 3DGS digital twin include: A graph convolutional neural network is used to perform pixel-level component semantic segmentation on point cloud data to obtain segmented point clouds, which are then associated with Gaussian points and bound to a component label library. Determine the boundary region and perform bilateral weighted smoothing on the Gaussian points in the boundary region to generate an enhanced 3DGS model; After performing hierarchical pruning on the enhanced 3DGS model, a three-level detail hierarchy structure of high, medium and low is constructed. The number of Gaussian points in each detail level is dynamically adjusted according to the viewing distance to generate a 3DGS digital twin.
4. The vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction according to claim 3, characterized in that, The specific steps for constructing the failure probability distribution matrix include: Associate real-time data streams with corresponding components and establish a real-time parameter cache to store key parameters of each component. Iterate through each fault rule in the historical fault records. If the key parameters meet the conditions of the fault rule, determine the fault probability contribution base based on the degree of deviation of the key parameters from the rule threshold and the actual proportion of the duration specified by the fault rule. When multiple fault rules for the same component are triggered, if they are from the same fault source, the largest fault probability contribution base is selected as the initial fault probability; otherwise, the fault probability bases of multiple fault rules are weighted and fused to obtain the initial fault probability. Determine the normal range and count the cumulative number of times key parameters exceed the normal range. Determine the correction factor and calculate the weighted failure probability based on the preliminary failure probability to construct the failure probability distribution matrix.
5. The vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction according to claim 4, characterized in that, The specific steps for generating a 3D fault probability heatmap include: Traverse all Gaussian points corresponding to each component in the 3DGS digital twin, obtain the RGB three color components through linear interpolation based on the weighted failure probability, and adjust the transparency. Attributes are assigned through ID mapping, and the spatial coordinates of Gaussian points mapped by ID are calibrated based on the global spatial coordinate system of the 3DGS digital twin. By using rendering pipeline level blending and same-channel texture synthesis mechanism, the fault probability thermal rendering layer and the vehicle body structure rendering layer are blended in the same rendering channel to obtain the blended Gaussian point rendering attribute set. Based on semantic boundary constraints of components, the thermal transition region between Gaussian points of different components is smoothed to generate a three-dimensional fault probability heat map.
6. The vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction according to claim 5, characterized in that, The specific steps for outputting a 3D fault diagnosis scenario include: Construct a first buffer and a second buffer for the Gaussian point rendering attribute set, which are used for rendering the current frame and receiving the attributes after the fault probability distribution matrix is updated, respectively; Perform attribute updates and color mappings on the Gaussian points in the high, medium, and low detail levels at the current camera view distance; Monitor the real-time data stream, recalculate the weighted failure probability of components whose parameters have changed, and update the color, transparency, and size parameters of the corresponding Gaussian point in the second buffer. At the end of the rendering cycle, a millisecond-level synchronous update link between the 3D fault probability heatmap and the real-time data stream is established through buffer swapping. The encapsulated interactive interface supports clicking on any thermal area to retrieve and display relevant information about components, retains the view control function to adapt to the real-time rendering frame rate requirements, and outputs a 3D fault diagnosis scene.
7. The vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction according to claim 6, characterized in that, The specific steps for establishing a spatial positioning coordinate system include: Extract the set of key geometric feature points of the vehicle from the 3DGS digital twin and construct a prior database of vehicle feature points; The real-time video stream is processed frame by frame. Key geometric feature points in the current frame are detected and descriptors are calculated. Hamming distance is calculated by combining the corresponding descriptors in the vehicle feature point prior database and a preliminary matching pair is established. The RANSAC algorithm is used to retain the successfully matched feature point pairs. The initial camera pose is obtained by solving the extrinsic parameters of the mobile device relative to the three-dimensional coordinate system defined by the 3DGS digital twin based on the feature point pairs. The optimized camera pose is obtained by optimizing the initial camera pose using the projection error. A local coordinate system for the vehicle is established, and the optimized camera pose is transformed to the local coordinate system for the establishment of a spatial positioning coordinate system.
8. The vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction according to claim 7, characterized in that, The specific steps for locating the faulty part include: The coordinates of Gaussian points in the 3D fault diagnosis scene are transformed to the spatial positioning coordinate system. The projection matrix is calculated based on the real-time pose data of the mobile device to generate a virtual rendering image. A virtual-real fusion rendering pipeline is constructed, using the acquired real vehicle video stream as the background layer and the virtual rendered image as the foreground layer for layer fusion. Calculate the actual depth value corresponding to each virtual pixel, and perform parallax offset correction on the virtual rendered image in combination with the display parameters of the mobile device; The system acquires and parses user operation commands, converts them into component selection commands, view control commands, transparency commands, and sectioning commands, and outputs the user operation command set and the target component. Based on the fault distribution data of the 3D fault probability heatmap and the user operation command set, the target component selected by the user is matched to locate the fault location.
9. The vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction according to claim 8, characterized in that, The specific steps for generating a fault diagnosis feedback report include: Construct a multi-user three-dimensional spatial synchronization channel and divide the data corresponding to the three-dimensional fault diagnosis scenario into static data and dynamic data; Static data is transmitted during the initial connection, and only dynamic data is transmitted thereafter. The dynamic attributes of each Gaussian point are encoded to generate synchronization data packets and transmitted via the UDP protocol. Perform full operation on the 3D fault diagnosis scene and transmit operation commands to the local terminal in real time; add 3D annotations on the remote terminal, encode them as Gaussian point attributes, transmit them to the local terminal, and display them persistently. AR devices instantly sense and overlay remote guidance information, and retrieve standard maintenance procedures based on the target component and the fault probability distribution matrix. 3D animations of maintenance operations are generated using 3DGS digital twins, and then overlaid on AR devices through a virtual-real fusion rendering pipeline to form a fault diagnosis feedback report.
10. A vehicle fault diagnosis system based on 3DGS modeling and human-computer interaction, used to implement the vehicle fault diagnosis method based on 3DGS modeling and human-computer interaction as described in any one of claims 1-9, characterized in that, include: Multi-source data fusion module, digital twin modeling module, 3D fault diagnosis module, and collaborative interaction module; The multi-source data fusion module is used to collect multi-view RGB images, point cloud data, real-time data streams and historical fault records, and preprocess them to generate a multimodal fusion dataset. The digital twin modeling module is used to build a component tag library, reconstruct the vehicle's three-dimensional model based on a multimodal fusion dataset and the 3DGS algorithm, and bind the components to the component tag library after semantic segmentation to optimize the generation of a 3DGS digital twin. The three-dimensional fault diagnosis module is used to calculate the weighted fault probability based on the 3DGS digital twin, combined with real-time data stream and historical fault records, to construct a fault probability distribution matrix, and to generate a three-dimensional fault probability heat map using the point-by-point attribute assignment mechanism of 3DGS Gaussian points, and output a three-dimensional fault diagnosis scenario. The collaborative interaction module is used to establish a spatial positioning coordinate system, overlay the three-dimensional fault diagnosis scene onto the real vehicle, parse the operation instructions and locate the fault location; construct a multi-user three-dimensional spatial synchronization channel to synchronize the three-dimensional fault diagnosis scene to the remote terminal in real time, provide step-by-step guidance locally with 3D animation, and generate a fault diagnosis feedback report.
Citation Information
Patent Citations
Fault display method, vehicle, equipment and storage medium
CN119672221A