Intelligent auditing auxiliary system based on mixed reality technology

The intelligent audit assistance system, which utilizes mixed reality technology, leverages multimodal sensors and edge computing modules, combined with a cloud-based knowledge base, to solve the problems of low efficiency and poor accuracy in traditional audit methods, achieving efficient and accurate automated auditing and natural interaction.

CN120848722APending Publication Date: 2025-10-28CCSC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510932683.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Traditional review methods are inefficient, prone to omissions, unable to collect and analyze audio and video data in real time, and rely on manual judgment, which leads to delays in problem detection.

Method used

The intelligent audit assistance system based on mixed reality technology includes a multimodal sensor array, an edge computing module, a mixed reality interactive interface, and a cloud knowledge base. Through multimodal data fusion and intelligent comparison, combined with federated learning and dynamic threshold determination, it achieves efficient matching and accuracy.

Benefits of technology

It improves review efficiency and accuracy, reduces the risk of missed reviews, enables real-time data analysis and natural human-computer interaction, and supports instant wake-up across operation stages and efficient data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848722A_ABST
    Figure CN120848722A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent auditing auxiliary head-mounted equipment, in particular to an intelligent auditing auxiliary system based on a mixed reality technology. According to the technical scheme, the system comprises a hardware architecture and a software system, the hardware architecture comprises a multi-modal sensor array, an edge computing module, a mixed reality interaction interface and a cloud knowledge base, and the software system comprises an intelligent comparison engine, a mixed reality interaction system and a data storage and management system. According to the invention, the intelligent comparison engine rapidly matches the auditing terms without manual searching and comparison, so that the auditing speed is greatly improved. In the aspect of accuracy, a multi-modal data acquisition and fusion technology is combined with a dynamic threshold judgment mechanism of an intelligent comparison engine, so that the difference between an auditing object and terms can be identified more accurately, and subjectivity and errors of manual judgment and missed auditing terms are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent audit assistance head-mounted devices, and more particularly to an intelligent audit assistance system based on mixed reality technology. Background Art

[0002] Traditional audits rely on paper documents, requiring auditors to manually review and compare each clause, which easily leads to omissions or duplicate reviews, and is extremely inefficient, especially in complex scenarios. For example, in an audit of an automobile production line, manually checking hundreds of standards can take several hours, with a risk of omissions as high as 15%.

[0003] Existing devices such as tablets and mobile phones require frequent interface switching (e.g., switching from viewing terms to the record interface), disrupting the review process. Traditional methods cannot collect and analyze audio and video data in real time. For example, abnormal equipment noise or worker violations in industrial environments cannot be automatically identified, and the subjectivity of human judgment leads to a delay of more than 30 minutes in problem detection. Summary of the Invention

[0004] The purpose of this invention is to address the problems existing in the background technology by proposing an intelligent review assistance system based on mixed reality technology that avoids the disconnect in electronic device interaction and improves review efficiency.

[0005] The technical solution of this invention: an intelligent audit assistance system based on mixed reality technology, comprising a multimodal sensor array, an edge computing module, a mixed reality interactive interface, and a cloud knowledge base connected in sequence, wherein:

[0006] The multimodal sensor array is used to collect visual, audio, and motion data in the auditing scenario in real time. The multimodal sensor array includes:

[0007] The binocular camera generates 3D reconstruction data of the scene through stereo vision algorithms and transmits it to the edge computing module;

[0008] The microphone array uses beamforming technology to directionally collect speech signals, which are then processed for noise reduction and output to the speech recognition unit.

[0009] The eye-tracking module captures the user's gaze point information and triggers the dynamic display of the corresponding review terms on the interactive interface;

[0010] A nine-axis motion sensor senses changes in the device's posture in real time and generates motion command signals;

[0011] The edge computing module is used to process multimodal data and perform intelligent comparison;

[0012] The mixed reality interactive interface is used to display review information and receive user instructions;

[0013] The cloud-based knowledge base is connected to the edge computing module via wireless communication, stores the feature library of audit terms and historical audit data, and updates the local comparison model through a federated learning algorithm.

[0014] Optionally, the edge computing module includes:

[0015] The vision processing unit is used to receive 3D reconstruction data and identify target objects in the scene through real-time semantic segmentation algorithms.

[0016] The speech recognition unit receives noise-reduced speech signals and converts them into text information using a locally deployed NLP engine;

[0017] The data fusion unit performs multimodal fusion of visual features, speech text, and action command signals to generate a joint feature vector;

[0018] The intelligent comparison engine calculates the similarity between the joint feature vector and the review clause features in the cloud knowledge base, and outputs the matching result by determining the dynamic threshold.

[0019] Optionally, the mixed reality interactive interface includes:

[0020] The holographic projection unit maps virtual review clauses to corresponding locations in the physical scene using spatial anchoring technology;

[0021] The gesture recognition module receives motion command signals from the nine-axis motion sensor and parses them into virtual interface operation commands.

[0022] The voice command module responds to preset wake words and executes the corresponding review operations.

[0023] Optionally, the data fusion unit adopts a federated learning framework to collaboratively train the locally generated joint feature vector with the global feature vector of the cloud knowledge base, thereby dynamically optimizing the feature matching model.

[0024] Optionally, the similarity calculation of the intelligent comparison engine includes the following steps:

[0025] Matching joint feature vectors with review clause features based on cosine similarity algorithm;

[0026] Dynamic judgment thresholds are generated based on the distribution of historical audit data.

[0027] When the similarity exceeds the threshold, the review result is automatically generated and stored in the structured database;

[0028] When the similarity does not reach the threshold, a manual review process is triggered and the user is prompted through the interactive interface.

[0029] Optionally, the holographic projection unit covers a virtual screen with a field of view of 30°-110° in the user's field of vision, and adjusts the position and hierarchy of the displayed content in real time based on the gaze point information of the eye-tracking module.

[0030] Optionally, the gesture recognition module parses the user's gestures using the optimized MediaPipe Hands algorithm, and combines head posture changes with gestures to generate composite operation commands.

[0031] Optionally, the data interaction between the cloud-based knowledge base and the edge computing module includes:

[0032] Regularly download the latest audit terms feature library to local storage.

[0033] Upload local audit data to the cloud for structured archiving and analysis;

[0034] The semantic understanding model of the local NLP engine is updated through federated learning.

[0035] Optionally, the voice signal processing flow of the voice command module includes:

[0036] User voice is captured directionally using a microphone array;

[0037] A deep learning-based noise reduction algorithm is used to filter environmental noise.

[0038] The processed speech data is input into the local NLP engine for intent recognition;

[0039] The identification results will trigger the generation of audit records or the retrieval of terms.

[0040] Optionally, the structured database stores audit data including timestamps, scene 3D reconstruction models, voice and text records, and audit result tags, supporting data traceability and statistical analysis based on multi-dimensional conditions.

[0041] In summary, this application includes at least one of the following beneficial technical effects:

[0042] 1. This invention achieves efficient matching of audio and video data with review terms through multimodal fusion and intelligent comparison, and improves the accuracy and adaptability of the comparison by using an intelligent comparison engine.

[0043] 2. Using mixed reality technology for interaction, digital information is overlaid on real-world scenes through methods such as perspective optical display. Combined with holographic menu projection, gesture recognition, and voice command systems, this achieves accurate mapping between virtual information and physical scenes, ensuring stable display of virtual content and highly realistic interaction in real-world scenes.

[0044] 3. By leveraging edge computing and data processing capabilities, heterogeneous computing architecture, and locally deployed NLP engines and video processing acceleration technologies, the system ensures efficient data processing and analysis on-premises, reducing reliance on the cloud and improving system real-time performance and stability.

[0045] 4. The present invention employs a technology in the intelligent comparison engine that automatically adjusts the matching threshold based on factors such as real-time data distribution and historical data distribution to adapt to the complexity of different review scenarios.

[0046] 5. By collaboratively training models between local devices and the cloud, the feature library is updated while protecting data privacy, thus preventing local models from becoming outdated;

[0047] 6. The instantaneous mapping of eye-tracking coordinates to the operation area achieved by the focus-gesture binding unit eliminates the audit interruption caused by traditional device interface switching. The voice standby maintenance unit supports instant wake-up across operation stages, allowing auditors to record voice simultaneously when marking defects with gestures, improving the smoothness of human-computer collaboration. The interface dynamic calibration unit compensates for virtual interface position offset through posture data, ensuring the stability of holographic display. The visual-voice data spatiotemporal alignment mechanism established by the hardware synchronization unit solves the problem of misjudgment caused by audio and video asynchrony. Attached Figure Description

[0048] Figure 1 This is a flowchart illustrating the hardware-software collaboration process of the present invention.

[0049] Figure 2 This is a system architecture diagram of the present invention;

[0050] Figure 3 Here is the algorithm flowchart;

[0051] Figure 4 For data flow diagrams;

[0052] Figure 5 This is an interactive flowchart;

[0053] Figure 6 Workflow diagram for multimodal interaction assurance and collaborative analysis. Detailed Implementation

[0054] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0055] Example 1, as Figures 1 to 5 As shown, the intelligent audit assistance system based on mixed reality technology proposed in this invention can be worn using VR headsets or VR glasses. The system includes hardware and software components, wherein the hardware component includes a multimodal sensor array, an edge computing module, a mixed reality interactive interface, and a cloud knowledge base.

[0056] Specifically, the multimodal sensor array includes:

[0057] The binocular 4K camera utilizes existing camera systems found in devices such as the PICO 4Ultra, Pimax Crystal, or VR glasses. It employs stereoscopic vision algorithms to reconstruct 3D scenes, supporting centimeter-level positioning and acquiring broader scene information. This provides auditors with clear and accurate visual data, such as the location, shape, and status of devices, helping them gain a more comprehensive understanding of the audit site.

[0058] The 360° microphone array supports sound source localization and noise reduction. Its integrated beamforming technology can accurately capture speech even in an 85dB noise environment, with a signal-to-noise ratio (SNR) greater than 20dB. It can capture sound information from all directions, accurately determine the direction of the sound source, effectively reduce environmental noise interference, and obtain clear speech data, providing a reliable basis for speech recognition and analysis.

[0059] The eye-tracking module accurately tracks the reviewer's eye movements with a high sampling rate of over 500Hz, obtaining information such as gaze points, providing the interactive system with a more accurate basis for judging user intent. For example, when the reviewer gazes at a certain device or area, the system can automatically pop up relevant review terms or prompts.

[0060] Nine-axis motion sensor (IMU): Real-time sensing of the device's motion status, including acceleration, angular velocity, and magnetic field information, for attitude tracking and motion recognition, enabling the device to respond accordingly to the auditor's actions, such as the auditor nodding, shaking their head, waving their hand, etc., which can be used to operate the virtual interface.

[0061] This system further deploys interactive continuity assurance devices and a multimodal collaborative analysis engine to enhance core functions:

[0062] The interactive continuity assurance device consists of the following units:

[0063] The focus-gesture binding unit generates a gesture operation area activation signal through the focus coordinates output by the eye-tracking module, and transmits it to the gesture recognition module, so that the virtual interface elements automatically enter the operable state when the user looks at them.

[0064] The voice standby maintenance unit continuously outputs a low-power monitoring signal to the voice command module during gesture recognition, supporting real-time response to voice commands at any stage of operation.

[0065] The interface dynamic calibration unit generates virtual interface position compensation parameters in real time based on the attitude data of the nine-axis motion sensor, and transmits them to the holographic projection unit to eliminate display offset caused by slight head movements.

[0066] The multimodal collaborative analysis engine executes within the edge computing module:

[0067] The hardware synchronization unit establishes a hardware-level pipeline between the vision processing unit and the speech recognition unit, and adds a synchronization timestamp (accuracy ±10ms) to the 3D reconstruction data acquired by the binocular camera and the noise-reduced speech signal of the microphone array.

[0068] The cross-modal coding unit receives timestamped visual features and speech text, generates a joint feature vector through a spatiotemporal alignment algorithm, and transmits it to the intelligent comparison engine.

[0069] The confidence arbitration unit monitors the difference in collaborative confidence weights between visual and speech data. When the difference exceeds a dynamic threshold, it outputs a multimodal review trigger command to the mixed reality interactive interface.

[0070] Through zero-delay switching between eye-tracking, gesture, and voice operations, such as synchronous voice marking of defects when swiping gestures during vehicle inspections;

[0071] Through precise spatiotemporal matching of multimodal data, such as synchronously associating visual states when identifying abnormal noise in equipment;

[0072] The confidence level arbitration can reduce the single-modal false positive rate, such as increasing the confidence level of stain recognition on conveyor belts in food factories to 98%.

[0073] As one implementation, the edge computing module of the system includes:

[0074] Heterogeneous computing architecture (CPU+NPU+GPU): The CPU is responsible for general computing tasks, such as system control and data transmission; the NPU accelerates deep learning tasks, such as speech recognition and image classification; the GPU handles graphics and image-related computing, such as rendering virtual images and processing video streams. The three work together to improve overall data processing efficiency.

[0075] Locally deployed NLP engine (BERT-base model quantization): Utilizing Natural Language Processing (NLP) technology to process speech and text data, the BERT-base model quantization reduces model size and computational load while maintaining performance, achieving efficient speech recognition and text understanding. It can accurately recognize the auditor's voice instructions and convert the speech content into processable text information.

[0076] Real-time video stream processing capability (TensorRT acceleration): With the help of TensorRT, video streams are processed by hardware acceleration to achieve real-time video analysis and feature extraction, meeting the real-time requirements of intelligent comparison and scene analysis, such as quickly identifying key information such as equipment and personnel actions in the video.

[0077] AI text generation coprocessor and local cache of auditing knowledge graph.

[0078] As one implementation method, the mixed reality interface of the system includes:

[0079] It adopts a perspective optical display solution with an ambient light transmittance of over 80%. Through spatial anchoring technology, it achieves accurate mapping between virtual information and physical scenes. While preserving the auditor's real field of vision, it displays virtual information through holographic menu projection (FOV40° virtual screen). It supports gesture recognition interaction (Optimized MediaPipeHands) and voice command system (Customized wake word recognition) to achieve natural interaction with the auditor.

[0080] As one implementation method, the system's cloud-based knowledge base includes:

[0081] It stores a wealth of information, including audit terms, standard documents, and historical audit data. It interacts with edge computing modules via wireless transmission to provide data support for the audit process, such as providing the latest audit standard terms and reference cases.

[0082] The software component in this embodiment includes: an intelligent comparison engine, a mixed reality interaction system, and a data storage and management system. Among them,

[0083] The intelligent comparison engine fuses audio text and visual features into a unified feature vector through multimodal feature fusion. It uses federated learning to update the clause feature library from the cloud, maintaining the accuracy and timeliness of the model. It uses cosine similarity to calculate the similarity with the feature vector of the review clause and adopts a dynamic threshold judgment mechanism based on historical data distribution. When the similarity exceeds the threshold, it outputs the matching clause ID and confidence score; otherwise, it triggers manual review.

[0084] Mixed reality interactive systems include:

[0085] 1. Holographic menu projection (FOV 30°-110° virtual screen): Displays review-related information, such as review terms and prompts, in the form of a virtual screen within the reviewer's field of vision. The 40° field of view ensures the integrity of the information display without excessively interfering with the real field of vision.

[0086] 2. Gesture Recognition Interaction: Utilizing the optimized MediaPipe Hands library to recognize the auditor's gestures, enabling operations on the virtual interface such as clicking, swiping, and zooming, thereby improving the convenience and naturalness of the interaction.

[0087] 3. Voice command system: The voice recognition function is activated by a preset wake word, which recognizes the auditor's voice commands and executes the corresponding operations, further improving interaction efficiency.

[0088] Data storage and management systems include:

[0089] Structured storage of audio and video data, audit results, audit paths, and other information during the audit process facilitates subsequent traceability and analysis, providing data support for continuous improvement of audit quality. For example, historical audit data can be used to statistically analyze the frequency of occurrence and problem types of different audit clauses, providing a reference for the formulation of audit plans.

[0090] In terms of review efficiency, the intelligent comparison engine quickly matches review clauses, eliminating the need for manual searching and comparison, thus significantly improving review speed. Regarding accuracy, multimodal data collection and fusion technology, combined with the intelligent comparison engine's dynamic threshold determination mechanism, can more accurately identify differences between the review object and the clauses, avoiding the subjectivity and errors of manual judgment, as well as missed clauses.

[0091] In terms of interactive experience, the mixed reality interactive system's holographic menu projection, gesture recognition, and voice command functions enable natural and smooth human-computer interaction. Auditors can operate the virtual interface through simple gestures and voice commands without manual input or device switching, making the operation convenient and efficient, and realizing the intelligentization of the audit process. For example, in a hygiene audit of a food processing plant, the auditor can mark problem areas with gestures while recording the audit results via voice. The entire process is smooth and natural, greatly improving work efficiency.

[0092] In terms of speech recognition, the noise reduction function of the 360° microphone array and the optimized voice command system effectively improve the speech recognition accuracy in industrial scenarios. In an 85dB noise environment, the speech recognition accuracy of this device can reach 92%, while traditional devices only achieve 78%, ensuring the accurate collection and processing of voice information and avoiding auditing errors caused by speech recognition mistakes.

[0093] In terms of data management, structured storage facilitates the traceability and analysis of audit data. By mining historical audit data, companies can identify potential problems, optimize audit processes, and improve overall management. For example, by analyzing historical audit data, a manufacturing company discovered frequent non-conformities in specific production processes, and subsequently made targeted improvements to production processes and management measures.

[0094] Example 2: Quality audit of automobile production line, such as Figures 1 to 5 As shown in the figure, based on Embodiment 1, the intelligent audit assistance system based on mixed reality technology proposed in this invention has the following usage process:

[0095] 1. Data Collection

[0096] Auditors wear equipment to enter the workshop. Binocular cameras capture video of the welding stations on the production line, microphone arrays capture equipment operating noise and worker conversations, and eye-tracking modules record gaze points (such as focusing on a specific weld point).

[0097] 2. Intelligent Analysis

[0098] The edge computing module performs real-time semantic segmentation on the video stream to identify the solder joint morphology; the NLP engine analyzes the voice command "View welding standards" and retrieves the corresponding ISO15614 clause. The intelligent comparison engine fuses the visual features (solder joint size) with the clause parameters (standard diameter ±0.2 mm), and the matching degree reaches 92%, determining it to be qualified.

[0099] 3. Interaction and Recording

[0100] The auditor switches to the next clause (bolt torque detection) by gesture sliding, and the voice command "Mark torque wrench" triggers virtual tag anchoring. The system automatically records the time, location, and audit results.

[0101] 4. Data Storage

[0102] After the audit is completed, all audio and video clips, operation logs, and matching results are structured and stored locally, and synchronized to the cloud to generate an audit report, including a heat map of problem distribution and improvement suggestions.

[0103] Example 3, Abnormal Recognition in Food Factory Hygiene Audit, such as Figures 1 to 5 As shown, based on Examples 1 and 2, this example includes the following steps:

[0104] 1. Abnormal Recognition

[0105] During the audit of the packaging workshop, the camera identifies stains on the conveyor belt surface through real-time semantic segmentation (confidence level 98%), and the microphone array detects abnormal noise from the equipment (frequency 4500 Hz, outside the standard range).

[0106] 2. Mixed Reality Interaction

[0107] The system automatically pops up the relevant clauses of the General Hygiene Standard for Food Production GB14881. The auditor circles the stain area by gesture and inputs "The cleaning is not thorough and needs to be processed immediately" by voice. The system generates a rectification work order and pushes it to the workshop supervisor.

[0108] 3. Dynamic Path Optimization

[0109] Based on historical audit data, the system recommends prioritizing the audit of high-risk areas (such as the raw material storage area). The time taken for the audit path this time is shortened by 25% compared to the traditional method, and the clause coverage rate is increased to 100%.

[0110] It should be noted that:

[0111] Mixed Reality (MR): Integrates virtual information with the real scene, supports two-way interaction, and this device achieves virtual-real integration through perspective display and spatial anchoring.

[0112] Edge computing: Adopting a "local-first + cloud-collaboration" architecture, NPU / GPU is responsible for real-time computing to reduce network latency (local processing accounts for ≥80%).

[0113] Federated learning: A distributed model training technology where local devices and the cloud only synchronize parameter updates, and the original data does not leave the device, ensuring privacy and security.

[0114] This invention constructs an efficient, accurate, and traceable intelligent auditing system through multimodal hardware collaboration, intelligent algorithm optimization, and mixed reality interaction innovation. It fills the technical gaps in real-time performance, interactivity, and data management in traditional auditing and is applicable to auditing scenarios in various industries undergoing digital transformation.

[0115] The above specific embodiments are merely several optional embodiments of the present invention. Based on the technical solutions of the present invention and the relevant teachings of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. An intelligent auditing assistance system based on mixed reality technology, characterized in that, It includes a multimodal sensor array, an edge computing module, a mixed reality interactive interface, and a cloud-based knowledge base connected in sequence, wherein: The multimodal sensor array is used to collect visual, audio, and motion data in the auditing scenario in real time. The multimodal sensor array includes: The binocular camera generates 3D reconstruction data of the scene through stereo vision algorithms and transmits it to the edge computing module; The microphone array uses beamforming technology to directionally collect speech signals, which are then processed for noise reduction and output to the speech recognition unit. The eye-tracking module captures the user's gaze point information and generates focus coordinates, triggering the interactive interface to dynamically display the corresponding review terms while maintaining the continuity of the main field of vision; A nine-axis motion sensor senses changes in the device's posture in real time and generates motion command signals; The edge computing module is used to process multimodal data and perform intelligent comparison; The mixed reality interactive interface is used to display review information and receive user instructions; The cloud-based knowledge base is connected to the edge computing module via wireless communication, stores the feature library of audit terms and historical audit data, and updates the local comparison model through a federated learning algorithm.

2. The intelligent audit assistance system based on mixed reality technology according to claim 1, characterized in that, The system includes an interaction continuity guarantee device, which includes: The focus-gesture binding module generates a gesture operation area activation signal based on the focus coordinates output by the eye-tracking module, and sends it to the gesture recognition module. The voice standby module continuously outputs a low-power monitoring signal to the voice command module while the gesture recognition module receives operation commands. The interface dynamic calibration module generates virtual interface position compensation parameters through attitude data output by the nine-axis motion sensor and sends them to the holographic projection unit.

3. The intelligent audit assistance system based on mixed reality technology according to claim 2, characterized in that, The edge computing module includes a multimodal collaborative analysis engine, which comprises: The hardware synchronizer establishes a hardware-level pipeline between the vision processing unit and the speech recognition unit, and adds a synchronization timestamp to the 3D reconstruction data and the noise-reduced speech signal. The cross-modal encoder receives time-stamped visual features and speech text, generates a spatiotemporally aligned joint feature vector, and feeds it into the intelligent comparison engine. The confidence arbitrator outputs a multimodal review trigger command to the mixed reality interactive interface when the difference in the collaborative confidence weights of visual and voice data exceeds a preset threshold.

4. The intelligent audit assistance system based on mixed reality technology according to claim 1, characterized in that, The edge computing module includes: The vision processing unit is used to receive 3D reconstruction data and identify target objects in the scene through real-time semantic segmentation algorithms. The speech recognition unit receives noise-reduced speech signals and converts them into text information using a locally deployed NLP engine; The data fusion unit performs multimodal fusion of visual features, speech text, and action command signals to generate a joint feature vector; The intelligent comparison engine calculates the similarity between the joint feature vector and the review clause features in the cloud knowledge base, and outputs the matching result by determining the dynamic threshold.

5. The intelligent audit assistance system based on mixed reality technology according to claim 1, characterized in that, The mixed reality interactive interface includes: The holographic projection unit maps virtual review clauses to corresponding locations in the physical scene using spatial anchoring technology; The gesture recognition module receives motion command signals from the nine-axis motion sensor and parses them into virtual interface operation commands. The voice command module responds to preset wake words and executes the review operation corresponding to the voice command; The voice signal processing flow of the voice command module includes: User voice is captured directionally using a microphone array; A deep learning-based noise reduction algorithm is used to filter environmental noise. The processed speech data is input into the local NLP engine for intent recognition; The identification results will trigger the generation of audit records or the retrieval of terms.

6. The intelligent audit assistance system based on mixed reality technology according to claim 5, characterized in that, The data fusion unit adopts a federated learning framework, which collaboratively trains the locally generated joint feature vector with the global feature vector of the cloud knowledge base to dynamically optimize the feature matching model.

7. The intelligent audit assistance system based on mixed reality technology according to claim 4, characterized in that, The similarity calculation of the intelligent comparison engine includes the following steps: Matching joint feature vectors with review clause features based on cosine similarity algorithm; Dynamic judgment thresholds are generated based on the distribution of historical audit data. When the similarity exceeds the threshold, the review result is automatically generated and stored in the structured database; When the similarity score does not reach the threshold, a manual review process is triggered and the user is prompted through the interactive interface. The structured database stores audit data including timestamps, 3D scene reconstruction models, voice and text records, and audit result tags, supporting data traceability and statistical analysis based on multi-dimensional conditions.

8. The intelligent audit assistance system based on mixed reality technology according to claim 5, characterized in that, The holographic projection unit covers a virtual screen with a field of view of 30°-110° in the user's field of vision, and adjusts the position and level of the displayed content in real time based on the gaze point information of the eye-tracking module.

9. The intelligent audit assistance system based on mixed reality technology according to claim 5, characterized in that, The gesture recognition module analyzes user gestures using the optimized MediaPipe Hands algorithm, and combines head posture changes with gestures to generate compound operation commands.

10. The intelligent audit assistance system based on mixed reality technology according to claim 1, characterized in that, The data interaction between the cloud-based knowledge base and the edge computing module includes: Regularly download the latest audit terms feature library to local storage. Upload local audit data to the cloud for structured archiving and analysis; The semantic understanding model of the local NLP engine is updated through federated learning.

Citation Information

Patent Citations

  • Intelligent head-mounted device and intelligent wearing system

    CN106257356A

  • Multi-modal data auditing method and device, storage medium and electronic equipment

    CN118378066A

  • Bridge crack real-time monitoring method based on edge calculation

    CN119904451A

  • Repair and maintenance method, system and equipment based on AR glasses and storage medium

    CN119992011A

  • Augmented reality cooperation system of cloud desktop

    CN120196205A