AI large model + OS deep fusion AR intelligent glasses system device and method
By using a lightweight AI large model engine, a deeply integrated operating system layer, a multimodal interaction system, and a holographic projection display module, the problems of insufficient AI processing capabilities, unnatural interaction, and poor display effects in existing AR smart glasses have been solved. This has enabled a high-performance, low-power immersive user experience, improving the efficiency and accuracy of industrial inspection, medical assistance, and education.
Patent Information
- Application Number
- CN202510980312.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing AR smart glasses have many shortcomings in terms of AI processing capabilities, multimodal interaction, display effects, and system architecture, making it difficult to meet the requirements of high performance, low power consumption, and immersive user experience.
It adopts a lightweight AI large model engine, a deeply integrated operating system layer, a multimodal interaction system, a holographic projection display module, and an edge computing collaboration module. Combined with pruning optimization, deep kernel modification, multimodal data processing, intelligent resource scheduling, and cloud collaboration optimization, it achieves improved AI processing capabilities, natural and smooth interaction, high-definition display, and low-latency collaborative computing.
It significantly improves the computing power, interactive experience, and display effect of AR smart glasses, extends the device's battery life, improves the accuracy of fault diagnosis and surgical precision, and optimizes resource allocation and learning efficiency.
Smart Images

Figure CN120821401A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of augmented reality smart glasses, and in particular to an AR smart glasses system device and method that deeply integrates an AI large model with an operating system. Background Art
[0002] In recent years, with the development of artificial intelligence and augmented reality technologies, AR smart glasses have shown broad application prospects in industries such as industry, healthcare, and education. However, existing AR smart glasses systems still have many shortcomings in terms of AI processing capabilities, multimodal interaction, display quality, and system architecture.
[0003] For example, in terms of AI processing capabilities, existing technologies usually adopt cloud-based AI processing or edge computing assistance. For example, US2020 / 0319174A1 discloses an AR system based on cloud-based AI processing, which transmits data to the cloud for processing through a high-bandwidth network, but this solution has obvious network dependence. When the network connection is unstable, the AI task processing delay can reach 60-120ms, seriously affecting the real-time interactive experience. In order to reduce dependence on the cloud, CN112363287B proposes an edge computing-assisted AR solution, but the amount of parameters of its locally deployed AI model still requires more than 1.2GB, which not only takes up a lot of storage space, but also the multimodal recognition accuracy in complex scenarios is only 85-90%, which cannot meet the needs of high-precision interaction.
[0004] In terms of display systems, the optical design of traditional AR glasses has significant flaws. The optical waveguide display system described in US10,921,788B2 has a field of view (FOV) of only 30-40° and suffers from 2-3% edge image distortion, significantly limiting the user's immersive experience. The multi-plane display technology disclosed in CN113466710A, while capable of presenting images at varying depths, only has a depth resolution of 1-3cm, significantly different from real-world depth perception and can easily lead to visual fatigue with prolonged use.
[0005] In terms of multimodal interaction technology, existing solutions have not yet completely solved the problems of signal synchronization and conflict resolution. The multimodal interaction framework proposed by US2021 / 0150728A1 combines multiple interaction methods such as vision, voice, and gestures, but does not disclose an effective multimodal signal synchronization mechanism, and its conflict resolution accuracy is less than 90%. The interaction method combining eye tracking and gesture recognition described in CN114564083A only involves basic interaction channels and does not introduce advanced interaction methods such as EEG signals. The error rate is high in complex environments.
[0006] At the system architecture level, traditional AR glasses employ a two-tier architecture: "application layer + operating system," resulting in a separation between AI reasoning and system resource scheduling. For example, the AR system architecture disclosed in US11,016,198B2 features an AI reasoning engine that operates independently from the operating system kernel, resulting in task response latencies exceeding 60ms and an inability to dynamically optimize system resource allocation based on real-time reasoning requirements. The AI model lightweighting method proposed in CN115373146A, while capable of reducing model parameter count, lacks deep integration with the operating system. Consequently, frequent memory swapping still occurs on memory-constrained AR devices, leading to significant fluctuations in AI processing performance.
[0007] In summary, existing AR smart glasses technology has obvious defects in multiple dimensions, making it difficult to meet users' needs for high-performance, low-power, and immersive AR experience. Summary of the Invention
[0008] The purpose of this invention is to provide an AR smart glasses system device and method that deeply integrates a large AI model with an operating system. Through innovative architecture design and algorithm optimization, this system significantly improves overall performance and user experience. This addresses the problems of existing AR devices, such as insufficient computing power, unnatural interactions, and poor display quality.
[0009] Technical solution:
[0010] The present invention proposes an AR smart glasses system device with deep integration of AI large model and OS, which mainly includes:
[0011] 1. Lightweight AI Large Model Engine: Integrated into the local system of AR smart glasses, this lightweight AI large model engine adapts to the computing power of mobile devices through pruning optimization. This pruning optimization includes unstructured pruning of model weight parameters, structured pruning of model structured units, and post-pruning model fine-tuning. Using technologies such as parameter pruning and quantization compression, it optimizes large models with hundreds of billions of parameters to a scale of hundreds of millions while retaining core capabilities.
[0012] 2. Deeply integrated operating system layer: This layer includes specialized driver components and an AI scheduling module, which deeply transforms the kernel and enables dynamic resource allocation.
[0013] 3. Multimodal interaction system: Integrates multiple interaction methods such as vision, voice, and gestures, and uses AI algorithms to achieve natural and smooth human-computer interaction.
[0014] 4. Holographic projection display module: uses advanced optical design and micro-projection technology to achieve high-resolution, large-field-of-view holographic display effects.
[0015] 5. Edge computing collaboration module: Collaborates with edge servers through low-latency networks to achieve dynamic allocation of computing tasks.
[0016] The present invention also proposes an implementation method based on the above device, which mainly includes the following steps:
[0017] 1. System startup optimization: Layered loading and prefetching technologies are used to ensure that the system starts quickly and enters a usable state.
[0018] 2. Multimodal data processing: Efficient collection and processing of sensor data is achieved through a unified data bus and memory management.
[0019] 3. Intelligent resource scheduling: Dynamically adjusts CPU / GPU / memory resource allocation based on task priority and real-time system status.
[0020] 4. Augmented reality rendering: Combines AI semantic understanding results to generate augmented reality images that match the user's current scene.
[0021] 5. Cloud-based collaborative optimization: Enable collaborative training and optimization of local and cloud models through edge computing.
[0022] Beneficial effects
[0023] 1. Improved computing power: Through model compression and edge collaboration, cloud-level AI processing capabilities can be achieved on limited hardware resources.
[0024] 2. Optimized interactive experience: Multimodal fusion interaction makes operation more natural and smooth, and conforms to ergonomic design.
[0025] 3. Enhanced display effect: Large field of view and high-resolution holographic display technology significantly enhance the sense of immersion.
[0026] 4. Improved battery life: Intelligent resource scheduling and energy efficiency optimization algorithms extend device usage time by more than 30%. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a schematic diagram of the overall architecture of the system device of the present invention, in which the core fusion layer realizes the deep collaboration between the AI big model and the operating system, which is the innovative core of the present invention.
[0028] Figure 2 This is a multimodal interaction processing flow chart of the present invention, which realizes efficient collaborative processing of multimodal signals through unified coding and attention mechanism.
[0029] Figure 3 This is the flow chart of the intelligent resource scheduling algorithm of the present invention, which dynamically optimizes the allocation strategy through AI task priority and real-time resource status.
[0030] Figure 4 This is a schematic diagram of the cloud-based collaborative optimization of the present invention, which achieves efficient resource utilization through three-level node division of labor and feedback closed loop. DETAILED DESCRIPTION
[0031] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0032] General system device embodiment:
[0033] like Figure 1 As shown, the AR smart glasses system device of the present invention mainly consists of the following parts:
[0034] 1. Lightweight AI large model engine 10
[0035] Integrated into the local system of AR smart glasses, the lightweight AI large model engine 10 performs pruning optimization to adapt to the computing power of mobile devices. The pruning optimization includes: unstructured pruning of model weight parameters, structured pruning of model structured units, and fine-tuning of the pruned model;
[0036] 2. Deep integration operating system layer 20
[0037] The Linux kernel has been extensively reworked, adding an AI scheduling module and hardware acceleration driver. The AI scheduling module utilizes real-time operating system (RTOS) technology to ensure response latency for key AI tasks is less than 10ms. The hardware acceleration driver is optimized for dedicated AI chips, providing efficient computing instruction set support.
[0038] 3. Multimodal Interaction System
[0039] It includes multiple sensors, including cameras, microphones, and inertial sensors, and uses data fusion algorithms to achieve multimodal interaction. The vision module uses the YOLOv8 lightweight object detection algorithm, the speech module uses the Wav2Vec2.0 speech recognition model, and gesture recognition uses MediaPipe gesture tracking technology.
[0040] 4. Holographic projection display module
[0041] The device uses MEMS micro-vibration mirror scanning technology combined with optical waveguide lenses to achieve a 50° field of view and a binocular display with a resolution of 1920×1080. The display system supports HDR10+ and a high refresh rate of 120Hz, with a brightness of up to 1200nits, meeting the requirements of use in bright environments.
[0042] 5. Edge computing collaboration module
[0043] Establish low-latency connections to edge servers via Wi-Fi 7 or 5G millimeter wave networks. Use a task splitting algorithm to offload computationally intensive tasks to edge servers while ensuring local data privacy and security.
[0044] Example General Method
[0045] like Figure 2 As shown, the implementation method of the present invention mainly includes the following steps:
[0046] 1. System startup optimization
[0047] The system uses a layered loading mechanism, first loading the core OS components and AI engine basic modules, ensuring the system is ready for use within 5 seconds. The remaining modules are continuously loaded and optimized in the background, using incremental update technology to reduce startup time.
[0048] 2. Multimodal Data Processing
[0049] Sensor data is transmitted to the data processing center using a zero-copy mechanism, employing memory pool reuse to reduce memory allocation overhead. Data preprocessing involves noise reduction and normalization, and then is transmitted to the AI engine via a dedicated data bus.
[0050] 3. Intelligent resource scheduling
[0051] Create AI task heat maps to monitor the computing load and priority of each task in real time. Dynamically adjust CPU / GPU / NNP resource allocation based on reinforcement learning algorithms, allowing high-priority tasks to have exclusive access to computing cores, ensuring the smooth operation of critical functions.
[0052] 4. Augmented Reality Rendering
[0053] The rendering engine combines AI semantic analysis results to generate augmented reality images with interactive prompts. It uses deferred rendering and instancing techniques to improve rendering efficiency while achieving advanced visual effects such as shadows and lighting.
[0054] 5. Cloud-based collaborative optimization
[0055] Through the federated learning mechanism, bidirectional parameter updates between the cloud and local models are achieved while ensuring user data privacy. The edge server collects model update gradients from multiple devices, aggregates them, and generates a global update package that is sent to each device.
[0056] like Figure 3 As shown in the figure, the AR smart glasses system with deep integration of AI large model + OS is applied to industrial inspection scenarios. The specific implementation method is as follows:
[0057] 1. Hardware Configuration and Customization
[0058] Perception hardware enhancement
[0059] AR smart glasses integrate a high-precision infrared thermal imaging camera, a vibration sensor, and a gas detector. The infrared thermal imaging camera measures the surface temperature distribution of the equipment with an accuracy of ±1°C. The vibration sensor, with a sampling frequency of 10kHz, monitors the vibration frequency and amplitude of the equipment in real time. The gas detector detects the concentration of six industrial gases, including methane and carbon monoxide, with a response time of less than 2 seconds. These sensor data are transmitted directly to the system core via a multimodal data bus, ensuring real-time and accurate data collection.
[0060] Display and interaction adaptation
[0061] The holographic projection display module utilizes high-light enhancement technology, maintaining clear images even in bright sunlight of 10,000 lux. Its field of view is expanded to 55°, allowing inspectors to quickly access device information. In terms of interaction, in addition to voice and gesture control, an industrial-grade touchpad has been added, supporting operation with gloves to meet the needs of complex industrial site operating environments.
[0062] 2. Software Function Customization
[0063] Industry-specific AI model deployment
[0064] Pre-trained industrial equipment fault diagnosis models (such as motor bearing fault diagnosis models and pipeline leak detection models) in the lightweight AI large model engine are deployed locally. Using progressive knowledge distillation technology, the model parameters are compressed to 150MB. This achieves 97% accuracy in abnormal motor sound recognition and 96% accuracy in pipeline crack detection on the local NPU, and can complete analysis of a single infrared thermal image in under 300ms.
[0065] Intelligent inspection workflow system
[0066] Deeply integrated with the operating system layer, a dedicated workflow engine for industrial inspections has been developed. Upon system startup, 3D models and historical data for corresponding equipment are automatically loaded according to pre-set inspection routes. During inspections, an AI-powered scheduler dynamically allocates system resources based on real-time task priorities (e.g., automatically raising the priority of fault diagnosis tasks when equipment anomalies are detected), ensuring that critical tasks are prioritized.
[0067] 3. Specific inspection operation process
[0068] Preparation before inspection
[0069] Inspectors wearing AR smart glasses activate the system using the voice command "Start industrial inspection mode." The system automatically downloads relevant data for the current inspection task from the cloud (including equipment location, historical fault records, inspection standards, etc.) and completes incremental loading of the lightweight AI model within 3 seconds.
[0070] Inspection execution phase
[0071] Automatic data collection: When inspectors approach equipment, smart glasses automatically trigger cameras and sensors to collect data. For example, within 5 meters of a motor, an infrared thermal imaging camera begins capturing the motor's surface temperature, while a vibration sensor simultaneously collects vibration data.
[0072] Real-time analysis and early warning: Collected data undergoes multimodal co-encoding and is then fed into a local lightweight AI model for analysis. If an abnormally high motor temperature or vibration frequency exceeds a threshold is detected, the smart glasses project a holographic image at the corresponding location on the device, highlighting the warning in red and announcing, "Possible motor bearing failure, please check immediately."
[0073] Expert Remote Assistance: For complex faults, inspectors can initiate remote assistance requests through the edge computing collaboration module. Upon receiving the request, cloud experts can view the inspector's perspective in real time, circle the fault point on the screen using AR annotation, and provide text or voice guidance. The entire process is kept within 200ms of latency.
[0074] Inspection end processing
[0075] After the inspection is complete, the system automatically generates an inspection report summarizing the inspection data, anomalies, and recommended actions. This report is uploaded to the cloud management platform via the edge node, while a lightweight version is retained locally for easy access. Based on the inspection data, the system fine-tunes and optimizes the local AI model through a federated learning mechanism to improve diagnostic accuracy for the next inspection.
[0076] 4. Technical advantages and effects
[0077] Inspection efficiency is greatly improved
[0078] Compared with traditional manual inspections, this system reduces the inspection time of a single device from 15 minutes to 3 minutes through automatic data collection and real-time analysis, and improves overall inspection efficiency by 80%.
[0079] Fault diagnosis accuracy is significantly improved
[0080] By combining multimodal data and lightweight AI models, the accuracy of equipment fault diagnosis has increased from 75% of manual inspections to over 96%, effectively avoiding missed detections and false detections.
[0081] Reduce operation and maintenance costs
[0082] By detecting potential equipment failures in advance and reducing sudden downtime accidents, it is expected that the company's equipment maintenance costs can be reduced by more than 30%, while improving the safety and stability of equipment operation.
[0083] like Figure 4 As shown, the specific implementation of AR smart glasses with deep integration of AI big model + OS in medical assistance
[0084] 1. Hardware configuration upgrade
[0085] Professional medical sensor integration
[0086] Medical-grade vital sign sensors are integrated into AR smart glasses, including a wearable ECG sensor, a non-invasive blood glucose monitoring module, and a blood oxygen saturation probe. The ECG sensor has a sampling rate of 500Hz, capturing detailed ECG waveforms in real time. The non-invasive blood glucose monitoring module uses near-infrared spectroscopy to automatically measure blood glucose levels every five minutes, with an error rate of ±5%. This data is synchronized with visual and audio data via a multimodal data bus in milliseconds, providing multi-dimensional information for medical decision-making.
[0087] Display and security optimization
[0088] The holographic projection display module utilizes medical-grade optical materials and features anti-fog and antibacterial coatings, making it suitable for use in environments such as operating rooms. The display supports independent adjustment of red, green, and blue channels, allowing doctors to highlight different types of medical images as needed. Additionally, an emergency touch button allows doctors to call the emergency team or access a patient's emergency medical records with a single click in emergencies.
[0089] 2. Development of Medical-Specific Software Functions
[0090] Medical AI model deployment
[0091] Medical imaging diagnostic models (such as lung CT image nodule detection models and brain MRI tumor recognition models) and surgical navigation models are deployed in the lightweight AI large model engine. Through model compression technology, the parameters of the lung CT image analysis model are compressed to 180MB, achieving a lung nodule detection accuracy of 95% on the local NPU, and processing a single CT image in just 200ms. At the same time, the integrated medical knowledge base model can provide real-time information such as drug interactions and disease diagnosis and treatment guidelines. Intelligent Medical Workflow System
[0092] Deeply integrate the operating system layer to develop a medical-specific workflow engine. In the surgical scenario, after the system is started, it automatically retrieves patient medical records, preoperative images and other information from the hospital HIS system, and connects data with operating room equipment (such as anesthesia machines and monitors). The AI perception scheduler dynamically allocates resources according to the surgical process. For example, when performing delicate surgical operations, it prioritizes the computing power requirements of the surgical navigation model.
[0093] Preoperative preparation stage
[0094] Wearing AR smart glasses, a doctor can use a voice command, "Retrieve patient Zhang San's preoperative data." The system quickly retrieves the patient's CT and MRI images and medical history from the cloud, and completes a 3D reconstruction within one second. Using gestures, the doctor can rotate and dissect the virtual model, gaining a comprehensive understanding of the patient's lesion location and surrounding tissue structure, enabling the development of a personalized surgical plan.
[0095] Intraoperative auxiliary stage
[0096] Real-time image navigation: During surgery, smart glasses accurately integrate the patient's real-time surgical image with the preoperative 3D model. Holographic projection is used to annotate the lesion boundaries, key blood vessels, and nerve locations in the surgical field of view in real time. For example, during brain tumor resection, the system automatically marks the boundary between the tumor and normal brain tissue, assisting the surgeon in precisely removing the lesion.
[0097] Multimodal data monitoring and early warning: Continuously monitors multimodal information such as the patient's vital signs and surgical instrument operation data. If a sudden abnormality in the patient's heart rate is detected or a surgical instrument is operating near a dangerous area, the smart glasses will issue a multi-dimensional warning through vibration, sound, and a highlighted warning box, and pop up a prompt on the display interface to provide treatment suggestions.
[0098] Remote Expert Collaboration: When encountering complex conditions, the surgeon can initiate a remote consultation request through the edge computing collaboration module. Remote experts can view surgical footage and patient data in real time, circle key procedures on the screen using AR annotation, and provide voice guidance, enabling face-to-face remote surgical guidance with latency under 150ms.
[0099] Postoperative recovery stage
[0100] Patients are equipped with home-use AR smart glasses for post-operative rehabilitation monitoring. These glasses continuously collect vital signs and activity data, analyzing their recovery progress using AI models. If any abnormalities are detected, the system automatically alerts the patient to seek medical attention and synchronizes the data with the hospital system. Furthermore, personalized rehabilitation training videos are pushed to patients, allowing them to interact with a virtual trainer using gestures.
[0101] 4. Technical advantages and application effects
[0102] Improve surgical precision and safety
[0103] Through AR real-time image navigation and multimodal data monitoring, surgical operation accuracy is improved by 40%, surgical risks are reduced by 30%, and surgical complications caused by human misjudgment are reduced.
[0104] Optimizing allocation of medical resources
[0105] Remote expert collaboration breaks down geographical restrictions, allowing patients in primary care hospitals to receive guidance from top medical teams and improving access to high-quality medical resources. Furthermore, AI-assisted diagnosis reduces repetitive work for doctors and improves diagnosis and treatment efficiency.
[0106] Improve patient recovery experience
[0107] Personalized rehabilitation monitoring and training plans shorten patients' rehabilitation period by an average of 20%, and improve patients' rehabilitation enthusiasm and compliance through virtual interactive training.
[0108] like Figure 4 As shown, the specific implementation of AR smart glasses with deep integration of AI big model + OS in education and teaching
[0109] 1. Hardware Function Adaptation
[0110] Interaction and display enhancements
[0111] The AR smart glasses are equipped with highly sensitive eye-tracking sensors and gesture recognition modules, allowing students to select and turn pages through gaze and gestures, with a response time of less than 50ms. The holographic projection display module optimizes color reproduction and contrast, ensuring clear visibility of virtual models and text in classroom lighting conditions. An eye protection mode is also included to reduce visual fatigue after prolonged use.
[0112] Audio and data extensions
[0113] An integrated noise-canceling microphone and surround-sound bone conduction headphones accurately capture student voice commands and reduce ambient noise in multi-person classrooms. Built-in large-capacity local storage (256GB) and high-speed Wi-Fi 6 support fast downloading of teaching resources and real-time upload of learning data.
[0114] 2. Development of educational software functions
[0115] Intelligent teaching content generation
[0116] Education-specific models, such as knowledge graph construction models, problem-solving models, and personalized learning recommendation models, are deployed within the lightweight AI large-scale model engine. Leveraging natural language processing technology, the system automatically generates courseware with AR scenarios and interactive Q&A based on the teaching objectives input by the teacher. For incorrect questions in student assignments, the AI model can generate dynamic explanation videos in real time to aid student understanding.
[0117] Classroom management and learning situation analysis system
[0118] Deeply integrated with the operating system layer, an intelligent classroom management platform has been developed. Teachers can monitor students' learning status (e.g., concentration and knowledge mastery) in real time while wearing glasses. An AI-aware scheduler dynamically allocates resources based on classroom interaction needs. For example, rendering computing power is prioritized during AR experiment demonstrations, while voice recognition processing is emphasized during knowledge question and answer sessions.
[0119] 3. Specific application process of education and teaching
[0120] Pre-class preparation stage
[0121] Teachers upload their course outlines via a cloud-based platform, and the system automatically generates AR teaching content based on AI models (e.g., 3D reconstructions of ancient architecture in history classes, dynamic models of cell structure in biology classes). Students, wearing AR smart glasses, can use gestures to explore virtual scenes during pre-class preparation. AI also pushes personalized pre-class materials based on students' browsing behavior.
[0122] Classroom teaching stage
[0123] Immersive teaching: When the teacher starts a lesson, the corresponding AR teaching scene automatically appears on students' glasses. For example, when explaining electromagnetic induction in a physics class, students can use gestures to "grab" virtual magnetic field lines to intuitively understand abstract concepts. In geography class, students can rotate a model of the Earth to observe the relationship between climate change and topography.
[0124] Real-time interaction and feedback: Students use voice to ask questions or gesture to highlight difficult points, and the teacher receives and responds in real time. The system automatically records student interaction data (such as question frequency and incorrect operation). AI models analyze this data and generate learning reports to help teachers adjust their teaching rhythm.
[0125] Collaborative group learning: In project-based learning, students work in groups wearing AR glasses to complete tasks together. For example, in a programming class, group members collaborate on code through an AR interface and see the code run in real time. In an art class, students collaborate on creating 3D virtual paintings, enabling creative sharing.
[0126] After-class review and evaluation stage
[0127] Based on classroom learning data, the system pushes personalized review content to students (such as AR animations explaining weak points and strengthening exercises). After students complete their homework, the AI model automatically grades subjective questions and generates analysis reports for incorrect answers. Teachers can review the overall learning progress of the class through the cloud platform and optimize subsequent teaching plans.
[0128] 4. Technical advantages and application effects
[0129] Improve learning efficiency and interest
[0130] AR technology visualizes abstract knowledge and, combined with personalized guidance from AI, increases students' knowledge absorption rate by 35%, classroom participation by 40%, and effectively reduces the phenomenon of "passive learning."
[0131] Optimize the allocation of teaching resources
[0132] AI learning situation analysis helps teachers accurately identify student needs and reduce repetitive teaching. At the same time, high-quality AR course resources are shared through the cloud to promote educational equity and reduce resource development costs.
[0133] Innovative education model
[0134] Break the limitations of traditional classrooms, realize cross-regional collaborative learning and virtual-real integration teaching experience, and promote practical teaching reforms in areas such as STEAM education and vocational skills training.
[0135] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An AR smart glasses system device with deep integration of AI large model + OS, characterized by: include: A lightweight AI large model engine (10) is integrated into the local system of the AR smart glasses. The lightweight AI large model engine (10) adapts to the computing power of the mobile device through pruning optimization. The pruning optimization includes: unstructured pruning of model weight parameters, structured pruning of model structured units, and fine-tuning of the model after pruning; The deep fusion operating system layer (20) includes a special driver component and an AI scheduling module to achieve dynamic allocation of hardware resources; The multimodal interaction system (30) includes a multi-channel input fusion processing mechanism of vision, voice, and gesture; The holographic projection display module uses micro-projection technology combined with optical waveguide lenses to achieve binocular 1080P holographic image display; The edge computing collaboration module collaborates with the cloud AI server in real time through a low-latency network protocol.
2. The system device according to claim 1, characterized in that: The lightweight AI large model engine (10) uses parameter quantization technology to compress FP32 precision to INT8 / 4 mixed precision while retaining the core attention mechanism.
3. The system device according to claim 1, characterized in that: The AI scheduling module of the deep fusion operating system layer (20) implements nanosecond-level scheduling of CPU / GPU / NNP through kernel-level HOOK technology, and the AI task response delay is less than 10ms.
4. The system device according to claim 1, characterized in that: The multimodal interaction system (30) includes a real-time semantic understanding module that can integrate visual scene recognition, voice command analysis, and gesture trajectory prediction to perform intention judgment with an accuracy rate exceeding 98.7%.
5. The system device according to claim 1, characterized in that: The holographic projection display module (40) adopts MEMS micro-vibration mirror scanning technology and combines the diffraction grating structure of the optical waveguide lens to achieve a 50° viewing angle and 1200 nit high brightness display.
6. A method for implementing an AR smart glasses system based on the AI big model + OS deep fusion of the device according to any one of claims 1 to 5, characterized in that: The following steps are involved: System startup phase: The OS kernel preloads the basic parameters of the AI model and initializes the hardware driver and edge computing connection; Interactive processing stage: After pre-processing, multimodal sensor data is transmitted to the AI engine through a zero-copy mechanism for real-time inference; Resource scheduling stage: The AI scheduling module dynamically allocates computing resources based on task priority, with high-priority vision tasks exclusively using GPU cores. Display rendering stage: The rendering engine combines AI semantic analysis results to generate augmented reality images with interactive prompts; Collaborative optimization stage: The edge computing module compares the local processing results with the cloud model and generates a personalized model update package.
7. The implementation method according to claim 6, characterized in that: The system startup phase further includes: the AI large model engine uses incremental loading technology to complete the loading of the core module within 5 seconds, and the remaining modules are continuously optimized in the background.
8. The implementation method according to claim 6, characterized in that: The interactive processing stage further includes: using memory pool reuse technology to reduce the data transmission delay between sensor data acquisition and AI reasoning to less than 3ms.
9. The implementation method according to claim 6, characterized in that: The resource scheduling stage further includes: establishing an AI task heat map, dynamically adjusting voltage and frequency based on real-time power consumption monitoring, and improving the system energy efficiency by 40%.
10. The implementation method according to claim 6, characterized in that: The collaborative optimization stage further includes: using a federated learning mechanism to achieve bidirectional parameter updates between the cloud and local models while ensuring user data privacy.
Citation Information
Patent Citations
SOC (State of Charge) and SOH (State of Health) collaborative estimation method for energy storage battery in power grid containing new energy receiving end
CN113466710A
Computer application stable operation device and application system
CN114564083A
Intelligent head-mounted device and shell thereof
CN115373146A
Route based manufacturing system analysis, apparatuses, systems, and methods
US10921788B2
Broadcast transmission of information indicative of a pseudorange correction
US11016198B2