Multi-modal data acquisition and intelligent analysis application method based on AI glasses

By synchronously collecting multimodal data through AI glasses and establishing temporal and spatial correlation indexes, combined with local cloud processing, the problems of low efficiency and insufficient privacy protection in the integration of multi-source information by smart devices are solved. This enables real-time information integration and personalized output, improving users' information acquisition and decision-making capabilities.

CN121785465APending Publication Date: 2026-04-03HUPENG (CHENGDU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing smart devices struggle to efficiently process various types of data, such as text, images, and sound, when faced with the integration of multi-source information. They are unable to flexibly adapt to user needs in different scenarios, and there is insufficient protection of personal privacy during information processing. Data transmission and storage pose risks, and they cannot achieve real-time information integration and effective output across scenarios.

Method used

AI glasses are used to simultaneously collect multimodal data through cameras, binocular cameras, array microphones, and eye-tracking sensors. The data is then linked by timestamps and spatial coordinates to establish an index. A lightweight local AI model is used for basic analysis, and the data is uploaded to a large cloud model for collaborative processing. An edge-cloud collaboration module and a federated learning framework are used to ensure privacy protection by uploading only feature parameters.

Benefits of technology

It enables real-time processing and personalized assistance of multimodal data, enhances users' ability to acquire information and make decisions in complex scenarios, balances efficiency and privacy, and provides scenario-based content output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121785465A_ABST
    Figure CN121785465A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data acquisition and intelligent analysis application method based on AI glasses. The method comprises the steps that the AI glasses collect text data through a camera and an OCR assembly, collect image video data through a binocular camera and an infrared sensor, collect audio data through an array microphone, and collect auxiliary data through an I MU and an eye movement tracking sensor; according to the AI glasses, a local lightweight AI model is used for basic analysis, and complex tasks are uploaded to a cloud large model for cooperative processing, including information summarization, deep mining and knowledge extraction; the AI glasses realize real-time pushing, scene assistance and content generation and output through an information application module according to an analysis result; according to the AI glasses, an end-cloud collaboration module and a federated learning framework are adopted, a local cloud processing strategy is dynamically adjusted, only feature parameters are uploaded, and it is ensured that original data are locally reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a method for multimodal data acquisition and intelligent analysis based on AI glasses. Background Technology

[0002] In the field of smart wearable devices, researching how to enhance users' information processing and decision-making capabilities in complex scenarios through technological means is particularly crucial. This area directly relates to people's efficiency and experience in daily work, study, and business activities. Especially in the era of information overload, smart devices need to help users quickly access valuable content while ensuring the security of personal information. However, this is no easy task, and the development of related technologies faces numerous challenges that urgently require breakthroughs.

[0003] Currently, although some smart devices on the market can collect and process information, they often fall short when integrating multi-source information, especially when processing multiple types of data such as text, images, and audio simultaneously, making efficient collaboration difficult. These devices typically cannot flexibly adapt to user needs in different scenarios, such as recording key points in meetings or providing supplementary information during study, often resulting in a poor user experience. More importantly, existing methods do not adequately protect personal privacy during information processing, and there are significant risks in data transmission and storage.

[0004] A deeper challenge lies in establishing data correlations amidst the efficient collection and processing of data from multiple sources, a significant technical hurdle. Data correlation refers to integrating information from different sources through temporal and spatial correspondences to form a complete information picture. Without this, the information collected by the device remains fragmented and fails to generate meaningful output. For example, in a meeting, the device might record the speaker's voice, images of the meeting content, and the user's focus. However, if this information cannot be linked chronologically and spatially, a complete meeting record cannot be generated, nor can targeted suggestions be provided to the user. This lack of correlation directly diminishes the device's auxiliary functions in complex scenarios.

[0005] Therefore, how to establish precise temporal and spatial correlations based on the efficient collection and processing of various types of data, so as to achieve real-time information integration and effective output across scenarios, has become a key problem that this research urgently needs to solve. Summary of the Invention

[0006] This invention provides a method for multimodal data acquisition and intelligent analysis based on AI glasses, which mainly includes:

[0007] AI glasses collect text data through cameras and OCR components, image and video data through binocular cameras and infrared sensors, audio data through array microphones, and auxiliary data through IMU and eye-tracking sensors.

[0008] The AI ​​glasses preprocess the collected multimodal data, including text data cleaning and standardization, key frame extraction of images and videos, audio-to-text tagging features, and establish a timestamp and spatial coordinate association index to form a spatiotemporal content fusion dataset.

[0009] The AI ​​glasses utilize a local lightweight AI model for basic analysis and upload complex tasks to a large cloud model for collaborative processing, including information summarization, in-depth mining, and knowledge extraction.

[0010] The AI ​​glasses, based on the analysis results, use the information application module to achieve real-time push notifications, contextualized assistance, and content generation output.

[0011] The AI ​​glasses employ an edge-cloud collaborative module and a federated learning framework to dynamically adjust local cloud processing strategies, uploading only feature parameters to ensure that the original data is retained locally.

[0012] Furthermore, the AI ​​glasses collect text data through a camera and OCR component, image and video data through a binocular camera and infrared sensor, audio data through an array microphone, and auxiliary data through an IMU and eye-tracking sensor. This includes: the camera and OCR component extracting text information from the scene in real time, supporting multilingual recognition and tilt correction; and the binocular camera and infrared sensor collecting dynamic video streams and static images, recording object, person, and environmental features, and supporting night mode and motion blur correction.

[0013] The array microphones collect speech signals, distinguish human voices from ambient sounds, and support sound source localization, noise reduction, and multi-channel separation.

[0014] The IMU and eye-tracking sensor record the trajectory of head movement gaze points to determine the information the user is focusing on.

[0015] Furthermore, the AI ​​glasses preprocess the collected multimodal data, including text data cleaning and standardization, keyframe extraction from images and videos, audio-to-text feature tagging, and establishing a timestamp-spatial coordinate association index to form a spatiotemporal content fusion dataset, including:

[0016] Text data is converted into structured text by removing redundant symbols;

[0017] Image and video data compression resolution extraction of keyframes;

[0018] Audio data is converted into text to mark intonation and pause features;

[0019] Text, image, and audio data are linked by timestamps and spatial coordinates, and a spatiotemporal content fusion dataset is formed by using the text position corresponding to the user's gaze point.

[0020] Furthermore, the AI ​​glasses utilize a local lightweight AI model for basic analysis and upload complex tasks to a large cloud model for collaborative processing, including information summarization, in-depth mining, and knowledge extraction, including:

[0021] A local lightweight AI model enables text semantic extraction, image object detection, and audio emotion recognition;

[0022] The cloud-based large model receives and uploads complex tasks, summarizes the information, and generates lecture minutes.

[0023] Cloud-based large-scale model analysis identifies the relationship between text and images, and recognizes the logical relationships between text and charts in PPT presentations;

[0024] The cloud-based big data model analyzes the matching degree between audio and video, and locates video demonstration segments based on the content of the speech;

[0025] The cloud-based big data model extracts technical terms, formulas, and cases from multimodal data to build personalized knowledge graphs.

[0026] Furthermore, the AI ​​glasses, based on the analysis results, utilize the information application module to achieve real-time push notifications, contextualized assistance, and content generation output, including:

[0027] Real-time push notifications via voice feedback;

[0028] Record the speakers' opinions in a meeting setting;

[0029] Generate to-do items and mark the points of disagreement.

[0030] Furthermore, the AI ​​glasses, based on the analysis results, utilize the information application module to achieve real-time push notifications, contextualized assistance, and content generation output, including:

[0031] Contextualized assistance identifies difficult formulas in textbooks within learning scenarios and automatically associates them with video explanations and relevant literature.

[0032] Contextualized assistance analyzes customer facial expressions and tone of voice in business scenarios to generate communication strategy suggestions;

[0033] Content is generated and exported as documents, mind maps, and short videos, and shared to cloud local storage via AI glasses.

[0034] Furthermore, the AI ​​glasses employ an edge-cloud collaborative module and a federated learning framework to dynamically adjust local cloud processing strategies, uploading only feature parameters to ensure that raw data is retained locally, including:

[0035] Through the federated learning framework, feature parameters are uploaded to the local model, while the raw data is stored on the device.

[0036] Dynamically adjust simple tasks locally and complete complex tasks in the cloud.

[0037] Furthermore, the camera and OCR component extract text information from the scene in real time, supporting multilingual recognition and tilt correction, including:

[0038] The OCR component recognizes text content from paper documents, electronic screens, and handwritten whiteboards;

[0039] Corrected text data is obtained based on the recognition results of tilt correction processing.

[0040] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0041] This invention discloses an AI glasses based on multimodal data fusion. Addressing the complex needs of information collection, processing, and real-time assistance in meeting, learning, and business scenarios, it proposes a unique business scenario problem: how to achieve real-time information analysis and contextualized content output across scenarios while efficiently collecting multimodal data and protecting privacy. This invention simultaneously collects text, images, audio, and auxiliary data using a high-definition camera, binocular cameras, array microphones, and eye-tracking sensors. It establishes a spatiotemporal correlation index using timestamps and spatial coordinates to ensure the accuracy of data fusion. Employing an edge-cloud collaborative and federated learning framework, a lightweight local model handles basic tasks, while a large cloud model deeply mines information, uploading only feature parameters to protect privacy. Finally, the information application module outputs contextualized content such as meeting minutes, learning association suggestions, and business strategies. The technical advantages of this invention lie in achieving real-time processing and personalized assistance of multimodal data, balancing efficiency and privacy, and significantly improving users' information acquisition and decision-making capabilities in complex scenarios. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of a multimodal data acquisition and intelligent analysis application method based on AI glasses according to the present invention.

[0043] Figure 2 This is another schematic diagram of a multimodal data acquisition and intelligent analysis application method based on AI glasses according to the present invention.

[0044] Figure 3 This is another schematic diagram of a multimodal data acquisition and intelligent analysis application method based on AI glasses according to the present invention.

[0045] Figure 4 This is another schematic diagram of a multimodal data acquisition and intelligent analysis application method based on AI glasses according to the present invention. Detailed Implementation

[0046] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] like Figure 1-4 This embodiment provides a method for multimodal data acquisition and intelligent analysis based on AI glasses, which may specifically include the following steps:

[0048] AI glasses collect text data through a high-definition camera and OCR component, image and video data through a binocular camera and infrared sensor, audio data through an array microphone, and auxiliary data through an IMU (Inertial Measurement Unit) and eye-tracking sensor.

[0049] Using a high-definition camera and optical character recognition components, text content is acquired from the surrounding environment, converting raw text data from paper documents or electronic screens into editable digital text information. Simultaneously, the spatial location and timestamp of the text are recorded, forming a text record with contextual markers. Using a binocular camera and infrared sensors, image and video data are acquired from the spatial location corresponding to the text record, capturing object features and dynamic changes in the environment. Infrared information is combined to perform lighting correction on the images, generating a visual data stream containing environmental details. Using an array microphone, audio signals are acquired from the scene covered by the visual data stream, separating human voice from background noise. Sound source localization is performed on the human voice portion, and the audio content is associated with the corresponding spatial location, forming a sound record matching the visual data stream. Using an inertial measurement unit and eye-tracking sensors, the user's focus of attention is extracted from the sound record and visual data stream, recording head movement trajectories and gaze focus areas. This auxiliary information is spatiotemporally aligned with the text record, visual data stream, and sound record to construct a multi-dimensional data set for comprehensive multimodal data acquisition based on AI glasses.

[0050] Specifically, the generation steps are as follows:

[0051] In one embodiment, the process of a high-definition camera and an optical character recognition (OCR) component acquiring text content from the surrounding environment involves the camera capturing a high-resolution image, and then the OCR component converting the characters in the image into digitized text information through pixel analysis and pattern matching. This conversion ensures the accurate extraction of the original text data while recording spatial location and timestamps to form text records with contextual tags. This improves data correlation and is beneficial for subsequent multimodal fusion because contextual tags allow other data types to be aligned with them, providing a more comprehensive understanding of the environment.

[0052] In one embodiment, the mechanism of acquiring image and video data from the spatial location corresponding to the text record using a binocular camera and an infrared sensor is to capture object features and dynamic changes by calculating depth information through the binocular camera, while the infrared sensor detects thermal signals to correct the image in areas with insufficient light. This generates a visual data stream containing environmental details, which can improve the acquisition quality under low light conditions and is beneficial for real-time scene reconstruction. This is because the corrected data stream provides accurate spatial references for audio and auxiliary data, avoids data isolation, and improves the overall coherence of acquisition.

[0053] In one embodiment, the acquisition of audio signals from the scene covered by the visual data stream using an array microphone includes the microphone array using beamforming technology to separate human voices from background noise, then performing sound source localization for the human voices, and associating the audio content with spatial location by calculating the time difference of arrival to form a sound record that matches the visual data stream. This association is beneficial for creating synchronized multi-sensory recordings because the matched sound record can enhance the interpretation of dynamic events, such as identifying the speaker's location, thereby providing audio context support for extracting the user's focus.

[0054] In one embodiment, the process of extracting the user's focus of attention from the audio recordings and visual data streams using an inertial measurement unit (IMU) and an eye-tracking sensor involves the IMU monitoring acceleration and gyroscope data to record head movement trajectories, while the eye-tracking sensor tracks the pupil position to determine the area of ​​focus. This auxiliary information is spatiotemporally aligned with the text recordings, visual data streams, and audio recordings to construct a multi-dimensional dataset. This facilitates the comprehensive acquisition of multimodal data based on AI glasses because spatiotemporal alignment ensures that all data is consistent in time and space, forming a cohesive dataset that supports deep analysis applications.

[0055] The AI ​​glasses preprocess the collected multimodal data, including text data cleaning and standardization, key frame extraction of images and videos, audio-to-text tagging features, and establish a timestamp and spatial coordinate association index to form a spatiotemporal content fusion dataset.

[0056] The AI ​​glasses perform preliminary processing on the collected multimodal data. First, text data is cleaned to remove irrelevant symbols and formatting errors, resulting in standardized text content. Simultaneously, keyframes are extracted from image and video data, retaining representative image segments. Audio data is converted into text content, and intonation and pause features are marked, forming a pre-processed multimodal data set. For this pre-processed multimodal data set, the AI ​​glasses further construct timestamp associations, uniformly labeling the collection time of each data type to ensure that text content, keyframe images, and audio text are aligned in the time dimension, forming a time-consistent data record. Based on this time-consistent data record, the AI ​​glasses introduce spatial coordinate information. Through the correspondence between the user's gaze point and the scene location, text content, keyframe images, and audio text are bound to specific spatial locations, establishing a spatial dimension association index to form a spatiotemporal fusion dataset. For this spatiotemporal fusion dataset, the AI ​​glasses integrate various data types at the content level, using timestamps and spatial coordinates as indexing criteria to ensure that different modalities form interconnected structured information, ultimately generating a spatiotemporal content fusion dataset for subsequent processing.

[0057] In one possible implementation, the AI ​​glasses can begin the process of cleaning the collected multimodal data by cleaning the text data. For example, for text information extracted from a scene, such as the content on a paper document, redundant punctuation or inconsistent formatting can be removed to obtain clear and standardized text content. This helps improve the accuracy of subsequent analysis because standardized data is easier to integrate with other modalities, avoids misjudgments caused by messy symbols, and helps improve the efficiency of overall data processing.

[0058] For example, when processing image and video data simultaneously, keyframe extraction preserves representative scene segments, such as capturing key moments of a speaker's facial expressions in a meeting scenario. This extraction highlights important visual information, reduces the burden of redundant data storage, and allows the system to operate more smoothly in a low-power environment, bringing the convenience of real-time processing. Furthermore, converting audio data into text and marking intonation and pause features—for example, marking high-pitched voices or long pauses after converting speech recordings to text—not only preserves emotional nuance but also facilitates association with visual data, enhancing the depth of analysis. The beneficial effect is to achieve a more comprehensive contextual understanding and avoid the loss of semantic richness from simple text.

[0059] It should be noted that these preliminary processing steps create a multimodal dataset that lays the foundation for subsequent steps, ensuring consistent data quality. When constructing timestamp associations for the pre-processed multimodal dataset, for example, uniformly labeling the collection time of each type of data, such as aligning text content at a certain moment with keyframe images and audio text from the same period, this creates time-consistent data records. This is beneficial for tracking the evolution of dynamic events because time alignment prevents discontinuity between modalities and improves the temporal accuracy of the analysis.

[0060] In one possible implementation, the process of introducing spatial coordinate information can be achieved by associating the user's gaze point with the scene location. For example, the gazed text location can be bound to the corresponding keyframe image and audio text to establish a spatial association index. Such binding forms a spatiotemporal fusion dataset, which has the benefit of enabling location-sensitive queries, making it easier for users to trace back information in specific areas, and improving the interactive experience.

[0061] For example, when integrating spatiotemporal fusion datasets at the content level, timestamps and spatial coordinates are used as indexes to ensure that data from different modalities are correlated. For instance, text, images, and audio from a spatial location at a certain point in a meeting can be integrated into structured information to ultimately generate a spatiotemporal content fusion dataset. This not only optimizes data retrieval but also supports in-depth applications such as automatic report generation, which is beneficial for intelligent user decision-making because the fused dataset provides a comprehensive view and reduces the need for manual intervention.

[0062] It should be noted that the entire process progresses step by step from initial sorting to final fusion, with each step strengthening the connectivity between data. For example, the time labeling of the initial set directly supports spatial binding, while the spatial index serves content integration, thereby improving the coherence and practical value of multimodal processing as a whole. This design is beneficial for the efficient application of AI glasses in real-time scenarios, avoids the problem of data silos, and ensures the structured and contextualized nature of the output results.

[0063] The AI ​​glasses utilize a local lightweight AI model for real-time basic analysis, including text semantic extraction, image object detection, and audio emotion recognition. Complex tasks are uploaded to a large cloud model for collaborative processing, including information summarization, in-depth mining, and knowledge extraction.

[0064] The AI ​​glasses, through a built-in lightweight processing unit, perform preliminary semantic extraction on the collected text data, transforming the original text content into first structured text data. Simultaneously, it performs object detection on the image data, generating first object annotation results, and extracts sentiment features from the audio data, forming a first sentiment tag set. The AI ​​glasses then perform preliminary fusion of the first structured text data, the first object annotation results, and the first sentiment tag set through a local processing unit, aligning the three based on timestamps to form a first multimodal association dataset, providing a unified data foundation for subsequent processing. For the first multimodal association dataset, the AI ​​glasses upload the data to a cloud processing platform via a wireless transmission channel. The cloud processing unit performs deep content mining on the dataset, generating a second summary content set and extracting key knowledge points to form a second knowledge structure graph. The cloud processing unit further integrates and refines the content based on the second summary content set and the second knowledge structure graph, transmitting the processed results back to the AI ​​glasses for presentation to the user via voice or a display interface, ensuring real-time access to in-depth analysis results based on multimodal data.

[0065] Specifically, the generation steps are as follows:

[0066] For example, when AI glasses process the collected text data, the lightweight processing unit first scans the original text content, such as identifying keywords and sentence structures from document images captured by the camera, and organizes this content into first structured text data through word segmentation and entity recognition processes. This processing can quickly filter out irrelevant noise, thereby improving the accuracy of subsequent analysis.

[0067] In one possible implementation, for image data, the unit applies object detection methods such as bounding box detection to label object positions and generate the first object labeling result. This helps to capture scene details in real time and avoid data loss. For audio data, by extracting pitch and rhythm features to form the first sentiment tag set, it can effectively capture the speaker's emotional changes and thus provide an emotional dimension for overall understanding. The benefit of doing so is that local processing reduces latency and ensures that users get immediate feedback in dynamic environments.

[0068] In one possible implementation, for the initial fusion of the first structured text data, the first target annotation results, and the first sentiment tag set, the local processing unit uses timestamps as a synchronization mechanism, for example, aligning the time points in the text with the image annotations and sentiment tags to form the first multimodal association dataset. This alignment process involves matching data fragments at the same time to build a unified spatiotemporal framework. The beneficial effect is to enhance the coherence of the data, making subsequent cloud analysis more efficient.

[0069] For example, in a meeting scenario, if text data records the content of speeches, images label the locations of participants, and emotion tags capture changes in tone, a complete event chain can be generated through fusion. This not only improves the usability of the data but also lays the foundation for in-depth data mining, avoiding information fragmentation caused by isolated processing.

[0070] For example, when uploading the first multimodal association dataset, the cloud processing unit receives the data and performs deep content mining. For instance, it generates a second summary content set by parsing the association elements in the dataset layer by layer. This involves summarizing key points such as merging text summaries and image descriptions, while extracting knowledge points such as the relationship between professional terms to form a second knowledge structure graph.

[0071] In one possible implementation, this mining process first decomposes the dataset into sub-modules and then reassembles them, revealing hidden patterns and thus enriching the depth of analysis. The beneficial effect is that cloud resources utilize powerful computing capabilities to handle complex tasks, reducing the burden on glasses devices and ensuring more comprehensive results. For example, in educational scenarios, mining the correlation between lecture highlights and visual aids from the dataset can help users quickly review knowledge.

[0072] In one possible implementation, for the integration of the second summary content set and the second knowledge structure graph, the cloud processing unit refines them into a concise output such as a comprehensive report, which is then sent back to the AI ​​glasses for presentation via voice broadcast or display.

[0073] For example, the integration process includes cross-validation of summaries and knowledge graphs to eliminate redundancy and generate the final result. This can provide scenario-based applications such as real-time decision support. The beneficial effect is that users obtain structured information rather than raw data, improve the interactive experience, and support the goal of real-time analysis through edge-cloud collaboration.

[0074] Based on the analysis results, the AI ​​glasses use the information application module to provide real-time contextualized assistance and content generation output, including automatically generating meeting minutes, linking learning difficulties with business strategy suggestions, and exporting documents, mind maps, and short videos.

[0075] The system acquires audio and image information from the meeting scenario using a multimodal data acquisition module. It separates the speaker's voice signal and records emotional fluctuations using an array microphone, while simultaneously capturing participants' facial expressions and body movements using a binocular camera, forming raw audio segments and raw image frames. These raw audio segments and image frames are then fed into a data preprocessing module to remove background noise from the audio and extract emotional tone features. Simultaneously, it performs lighting correction and facial feature point localization on the image frames, generating standardized audio and image feature sets. These standardized audio and image feature sets are then input into an information application module. Based on a pre-established mapping relationship between voice emotion and image expression, the module determines the speaker's viewpoint inclination and points of disagreement, automatically compiling meeting minutes and highlighting learning difficulties. Based on the meeting minutes content and learning difficulties, and combined with business scenario needs, it generates relevant communication suggestions. The results are then converted into document format, mind map style, or short video format and output to users via real-time push functionality, ensuring that content generation and scenario-based assistance closely align with the goals of automatic meeting minutes generation and business strategy recommendations.

[0076] In one possible implementation, audio and image information from the meeting scene is acquired from a multimodal data acquisition module. The speaker's voice signal is separated and emotional fluctuations are recorded using an array microphone. At the same time, a binocular camera captures the participants' facial expressions and body movements, forming raw audio segments and raw image frames. This ensures the comprehensiveness of data acquisition because the array microphone can accurately locate the sound source and avoid interference from mixed sounds, thereby improving the accuracy of subsequent analysis. For example, in a business meeting, when multiple people speak at the same time, the microphone separates each person's voice signal and records emotional fluctuations such as anger or excitement. Meanwhile, the binocular camera captures the corresponding frowning or nodding movements. These raw segments and frames serve as basic data, resulting in more reliable emotion recognition and helping to generate accurate meeting minutes.

[0077] For example, the original speech segments and original image frames are fed into the data preprocessing module to remove background noise from the speech and extract emotional tone features. At the same time, the image frames are subjected to lighting correction and facial feature point localization to generate standardized speech feature sets and standardized image feature sets. This processing can improve data quality because the speech signal is clearer after noise removal, and the extracted tone features, such as pitch changes, can reflect the speaker's emotions. Light correction ensures that the image is consistent under different lighting conditions, and facial feature point localization, such as the position of the eyes and mouth, helps to identify micro-expressions. These standardized sets, as a unified format, can bring more efficient fusion analysis results. For example, in a dimly lit conference room, the preprocessed image frames clearly show the participants' smiles, which, combined with the high-pitched tone of the speech, supports the reliability of emotion judgment.

[0078] In one possible implementation, standardized speech feature sets and standardized image feature sets are input into the information application module. Through a pre-established mapping relationship between speech emotion and image expression, the module determines the speaker's viewpoint inclination and points of disagreement, automatically compiling meeting minutes and marking learning difficulties. This mapping relationship is based on a correspondence table trained on a large number of samples, which can associate hesitant tone in speech with evasive eye contact in images, identifying points of disagreement such as opposing viewpoints. This allows for the automatic generation of minutes and the marking of difficult points such as technical terms. These judgments and compilations can provide decision-making assistance. For example, in a product discussion meeting, the module identifies a speaker's low voice accompanied by a frowning expression as a point of disagreement and marks it as a learning difficulty such as a disagreement on market strategy, helping users to quickly review the information.

[0079] For example, based on the content of meeting minutes and learning difficulties, relevant communication suggestions are generated in conjunction with business scenario needs. The results are then converted into document formats, mind maps, or short videos and delivered to users via real-time push notifications. This ensures that content generation and scenario-based assistance closely align with the goals of automatic meeting minutes generation and business strategy suggestions. For instance, suggestions such as mediation statements addressing points of disagreement can optimize communication efficiency. Conversion into mind maps visualizing the relationships between difficulties, or short video summaries, delivered to users' devices in real-time via push notifications, provides immediate application. For example, in negotiations, push notifications of mind maps displaying strategic suggestions such as adjusting pricing support immediate user responses, thereby improving the overall effectiveness of business interactions.

[0080] The AI ​​glasses employ an edge-cloud collaborative module and a federated learning framework to dynamically adjust local cloud processing strategies, uploading only feature parameters to ensure local retention of raw data and privacy protection.

[0081] During the data processing phase, the device first extracts key feature values ​​from locally collected multimodal information. Using pre-established feature extraction rules, it transforms the original image and audio content into structured feature parameters, completing initial localization processing and ensuring that subsequent uploaded content does not contain the original records. For the extracted feature parameters, the device uploads them to the cloud for further in-depth processing via an edge-cloud collaboration mechanism, while retaining all unprocessed original records locally to prevent the leakage of sensitive information. During cloud processing, the device and cloud nodes update parameters using a federated learning approach, jointly optimizing only the feature parameters, and then sending the updated parameters back to the local device to improve local processing capabilities. For the returned parameters, the device dynamically adjusts its local processing logic, applying the adjusted logic to subsequent multimodal information collection and feature extraction, ensuring continuous optimization of feature parameter extraction accuracy while protecting privacy, and meeting the business needs of the edge-cloud collaboration module.

[0082] Specifically, the generation steps are as follows:

[0083] For example, in the data processing stage, the process of the device extracting key feature values ​​from locally acquired multimodal information can be achieved through pre-established feature extraction rules. These rules include edge detection of image content and spectral analysis of audio content, transforming the original image into structured feature parameters containing edge contours and color distribution, and transforming the original audio into structured feature parameters containing pitch and rhythm. The benefit of doing so is to ensure that the content uploaded after local processing is only abstract parameters, thereby reducing the risk of data leakage.

[0084] In one possible implementation, for the image part, the main object contours in the scene are first identified as the first feature set, and then the distribution density of these contours is calculated as the second feature set. For the audio part, the human voice frequency band is first separated as the third feature set, and then its emotional intensity is quantified as the fourth feature set. The extraction of these feature sets makes subsequent processing more efficient and helps to improve the overall system response speed.

[0085] The extracted feature parameters are then uploaded to the cloud for further in-depth processing via an edge-cloud collaboration mechanism.

[0086] Understandably, this mechanism involves the design of the communication protocol between the local device and the cloud server. The specific way to retain all unprocessed raw records locally is to isolate the raw data through the device's built-in storage unit to prevent the leakage of sensitive information. The beneficial effect of doing so is to enhance user privacy protection, while allowing the cloud to use more powerful computing resources to optimize feature parameters.

[0087] In one possible implementation, the first and second feature sets are transmitted via an encrypted channel during upload. The cloud then applies aggregation calculations to fuse these parameters to generate enhanced model parameters, which are then sent back to the device. This allows the local device to obtain more accurate analysis capabilities without exposing the original image. For example, in a conference scenario, after the feature parameters are uploaded, the cloud can infer the speaker's emotions, while the original audio retained locally ensures the integrity of the recording and prevents external access.

[0088] When processing in the cloud, the device and cloud nodes update parameters based on federated learning.

[0089] Specifically, federated learning is a distributed machine learning method in which multiple devices jointly train a model without sharing the original data. They only jointly optimize the feature parameters. For example, after cloud nodes collect the feature parameters of multiple devices, they update the global model parameters by calculating the average gradient, and then send the updated parameters back to the local device to improve local processing capabilities. The benefit of doing this is that it enables continuous improvement of the model without transmitting sensitive data.

[0090] In one possible implementation, for joint optimization, the uploaded third and fourth feature sets are first anonymously aggregated with similar parameters from other devices, and the average optimized value is calculated as the update parameter. After being sent back, the local device applies these parameters to adjust its extraction rules, which is beneficial to improve the generalization ability of feature extraction in a multi-user environment. For example, in scientific research scenarios, after the audio feature parameters of different users are jointly optimized, the pronunciation patterns of professional terms can be better identified without leaking any personal voice records.

[0091] The process by which the device dynamically adjusts its local processing logic in response to the returned parameters.

[0092] Understandably, this adjustment involves integrating updated parameters into local rules and applying them to subsequent multimodal information collection and feature extraction. For example, the returned parameters are first compared with existing rules to generate new extraction thresholds, and then these thresholds are used to filter image and audio content during the next collection, ensuring continuous optimization of feature parameter extraction accuracy while protecting privacy. The beneficial effect of doing so is to achieve adaptive learning of the system and meet the business needs of the end-to-cloud collaboration module.

[0093] In one possible implementation, the adjusted logic, when applied to image acquisition, can more accurately capture dynamically changing object features, while when applied to audio, it can better separate background noise. This is beneficial for generating more reliable communication strategy suggestions in business scenarios. For example, when analyzing customer expressions, the optimized parameters ensure that feature extraction ignores irrelevant background, thereby providing accurate tone analysis, while all raw data remains locally, enhancing privacy protection.

[0094] In the multimodal data acquisition, text data acquisition supports multilingual recognition and tilt correction, image and video data acquisition supports motion blur correction in night mode, audio data acquisition supports sound source localization, noise reduction, and multi-channel separation, and auxiliary data recording of head movement gaze point trajectory.

[0095] In the multimodal data acquisition process, for text data, original text images in the scene are acquired using a high-definition camera. These images are then processed using optical character recognition (OCR) tools to correct tilted text images and classify and recognize multilingual text, generating a standardized text content data stream. For image and video data acquisition, dynamic video streams and static image data in the scene are simultaneously acquired from the device that acquired the original text images. Infrared sensors are used to enhance image clarity in low-light environments, and image processing tools correct motion blur, resulting in a clear set of image and video data. For audio data acquisition, scene information acquired from the image and video data set is combined with speech signals captured by an array microphone. Sound source localization tools determine the direction of the sound source, and environmental noise is separated and reduced to generate a clean audio data segment containing multi-channel separation. For auxiliary data acquisition, scene context information extracted from the clean audio data segment is combined with inertial measurement unit (IMU) recording of the user's head movement trajectory. Simultaneously, eye-tracking sensors capture the user's gaze path, forming user focus data for subsequent comprehensive processing and analysis of multimodal data, focusing on the collaborative acquisition of text, images, video, and audio.

[0096] Specifically, the generation steps are as follows:

[0097] For example, in the process of multimodal data acquisition, for text data, the original text image in the scene is acquired through a high-definition camera. This high-definition camera usually uses a high-resolution sensor to capture image data with rich details, thereby ensuring the basic quality of subsequent processing. The image is then processed in conjunction with an optical character recognition tool, which is an image analysis method based on pattern matching. It first converts the image into a grayscale image, then extracts the character outlines and compares them with a pre-stored character library. The tilted text image is then angle-corrected. This correction process involves detecting straight line features in the image and applying affine transformations to rotate and adjust the angle. Furthermore, multilingual text is classified and recognized. A trained classifier distinguishes language types such as English or Chinese and applies corresponding decoding rules to generate a standardized text content data stream. This improves data accuracy because correction and recognition reduce the risk of misinterpretation and are beneficial to the reliability of subsequent analysis.

[0098] In one possible implementation, for image and video data acquisition, dynamic video streams and static image data of the scene are simultaneously acquired from the device that acquired the original text images. This synchronization mechanism uses the device's internal clock to align the data timestamps of different sensors, ensuring the spatiotemporal consistency of the text data stream and the image and video. Infrared sensors are used to enhance image clarity in low-light environments. The infrared sensors supplement the insufficient visible light by capturing the invisible spectrum, thereby producing a brighter image. At the same time, motion blur is corrected by image processing tools. This correction process includes estimating motion vectors and applying deconvolution filtering to sharpen blurred areas, forming a clear set of image and video data. This improves data quality in nighttime or fast-moving scenes, helps capture complete environmental details, and avoids information loss.

[0099] For example, in audio data acquisition, scene information obtained from image and video datasets is combined with speech signals captured by an array microphone, which consists of multiple microphone units. The array microphone enhances directional sound pickup through phase difference calculation, and the direction of the sound source is determined by a sound source localization tool. This tool uses a time difference arrival algorithm to analyze the signal delay between microphones to calculate the azimuth angle, and performs environmental noise separation and noise reduction processing. The separation process involves spectral analysis to distinguish and filter the human voice frequency band from the noise frequency band, generating a clean audio data segment containing multi-channel separation. This isolates useful signals, which is beneficial for clearly recording the content in a multi-speaker environment and improving the overall accuracy of multimodal fusion.

[0100] In one possible implementation, for auxiliary data acquisition, scene context information extracted from pure audio data segments is combined with the user's head movement trajectory recorded by an inertial measurement unit (IMU). The IMU integrates accelerometers and gyroscopes to calculate three-dimensional displacement and rotation data. Simultaneously, an eye-tracking sensor captures the user's gaze path. This sensor uses infrared light reflection to track the pupil position and map it to display coordinates, forming user focus data for subsequent multimodal data processing and analysis. This focuses on the collaborative acquisition of text, images, video, and audio, which can correlate user intent with the acquired data, facilitates prioritizing the processing of the area of ​​interest, and achieves more intelligent scene understanding.

[0101] In the data preprocessing, redundant symbols are removed from text data and converted into structured text; keyframes are extracted from image and video data by compressing resolution; intonation and pauses are marked in audio data; and a multimodal data association index is established to correspond to the text and image positions based on the user's gaze point.

[0102] Redundant symbols are removed from the collected text data, and the original text content is converted into a structured text format. The timestamp and spatial location information of each text segment are recorded for subsequent correlation processing. For image and video data, the original image stream is compressed to extract keyframes, and the timestamps and spatial regions corresponding to the keyframes are labeled to ensure correspondence with the time and location information of the structured text. The audio data is characterized by intonation and pause features, generating timestamped audio feature sequences. Using user gaze point data, the text and image locations corresponding to the gaze points are matched with the audio feature sequences to establish a correlation index between multimodal data. The established multimodal data correlation index is integrated, and a unified temporal and spatial mapping relationship is formed for the text and image locations corresponding to the user gaze points, combined with the audio feature sequences, for subsequent data processing and content presentation.

[0103] In one possible implementation, the process of removing redundant symbols from the collected text data involves identifying and deleting parts such as repeated or irrelevant punctuation characters, converting the original text content into a structured text format. This makes the data easier to process later, while recording the timestamp and spatial location information of each text fragment helps improve the accuracy of data association.

[0104] For example, in a meeting setting, the collected handwritten whiteboard text may contain redundant underlines. By removing these symbols and converting it into structured text, key points can be clearly extracted, and timestamps such as the 5th minute after the meeting started and spatial locations such as the left side of the eyes can be added. This provides a basis for multimodal fusion, and the beneficial effects are reduced data noise and improved analysis efficiency.

[0105] For example, for image and video data, resolution compression of the raw image stream is performed by reducing the pixel density to reduce the file size, while extracting key frames such as frames with significant scene changes and labeling the timestamps and spatial regions corresponding to the key frames to ensure that they correspond to the time and location information of the structured text. This optimizes storage and transmission resources.

[0106] For example, in a lecture environment, video streams capture the speaker's gestures, extract keyframes containing formula demonstrations by compressing the resolution, and match the timestamps of the text. This helps to achieve real-time synchronization, avoids delays caused by data redundancy, and enhances the system's response speed, supporting the application of low-power devices such as AI glasses.

[0107] In one possible implementation, the process of generating a timestamped audio feature sequence by characterizing the intonation and pauses of audio data includes analyzing pitch and pause duration to label emotions or emphasis points, and matching the text and image positions corresponding to the gaze points with the audio feature sequence through user gaze point data to establish an association index between multimodal data. This can integrate information flows from different modalities.

[0108] For example, in a discussion, audio captures the speaker's excited tone and pauses. By marking and matching the text positions with the gaze points, it can reveal the user's focus. The beneficial effect is to improve the depth of information interpretation and avoid the limitations of a single modality.

[0109] For example, the process of integrating the established multimodal data association index targets the text and image locations corresponding to the user's gaze point, combines them with audio feature sequences, and forms a unified temporal and spatial mapping relationship for subsequent data processing and content presentation. This creates a comprehensive data view.

[0110] For example, in scientific research scenarios, the integrated index maps the literature text, related images, and audio explanations at the point of focus into a unified relationship, which is beneficial for generating a coherent knowledge graph. The beneficial effect is to realize scenario-based applications such as automatic summarization and enhance user decision support.

[0111] In one possible implementation, these processes begin with text preprocessing and gradually expand to multimodal association, ensuring that the output of each stage, such as structured text, directly supports the input of the next stage, such as keyframe matching, thus forming a tight logical chain that benefits the overall system's robustness.

[0112] For example, the location information of structured text is used to guide the selection of image compression, which is then fused with audio tags. The final mapping relationship closely follows the data preprocessing goal, supports the efficient establishment of multimodal indexes, and has the beneficial effect of optimizing the real-time analysis capabilities of AI glasses.

[0113] In the AI ​​large-scale model analysis and processing, a lightweight model is deployed locally to achieve basic analysis, while a large model in the cloud collaboratively realizes the generation of conference lecture summaries, logical association of PPT text and charts, audio and video matching and positioning, and construction of professional terminology knowledge graphs.

[0114] A lightweight processing component is deployed on a local device to perform preliminary semantic extraction on the collected text content, basic object detection on the image data, and sentiment feature annotation on the audio signal. This preliminary processed text semantic information, image object information, and audio sentiment information are integrated into a first dataset. The first dataset is uploaded to a cloud processing platform, where large-scale computing resources are used to deeply refine the text semantic information, generating key summaries of meetings or lectures. Simultaneously, the chart content in the image object information is logically associated with the text key points, forming a structured content mapping relationship. The audio sentiment information in the first dataset is matched and located with video clips. Timestamps are used to align the audio sentiment information with the video content, constructing an audio-video association index to generate a second dataset. The second dataset is compared with a pre-established professional terminology database to extract key terms from the text key summaries, constructing a knowledge graph framework. This knowledge graph framework is combined with the meeting / lecture summary content to form the final structured analysis result, supporting the deep association between meeting / lecture summary generation and professional content.

[0115] In one embodiment, a lightweight processing component is deployed on a local device to perform preliminary semantic extraction on the collected text content, such as identifying core keywords like "project progress" and "budget allocation" from meeting notes, thereby quickly capturing the basic meaning, which helps reduce subsequent cloud load and improve real-time response speed.

[0116] Specifically, this preliminary semantic extraction is achieved through word frequency statistics and simple contextual association, such as associating "progress" with "delay" to form a preliminary semantic chain, which is beneficial for providing basic input for deep processing in the cloud and avoiding delays caused by analyzing from scratch.

[0117] In one embodiment, basic object detection is performed on the image data, such as identifying chart elements like bar charts and lines in a PowerPoint presentation, and labeling their positions and types, which is beneficial to the accuracy of subsequent logical associations.

[0118] Specifically, basic object detection uses edge detection methods to scan image pixels to define object boundaries, such as separating the chart area from the background to form image object information, which is beneficial for maintaining data consistency when integrating multimodal data.

[0119] In one embodiment, the audio signal is labeled with emotional features, such as extracting tone tags of "excitement" or "emphasis" from lecture speeches, which is helpful to enrich the analysis dimensions by judging by tone changes and volume peaks.

[0120] Specifically, the sentiment annotation process involves spectral analysis to identify positive emotions represented by high-frequency bands, forming audio sentiment information, which is beneficial for collaboration with text and images and improves the contextual accuracy of the overall summary.

[0121] In one embodiment, this pre-processed information is integrated into a first dataset, for example, by aligning textual semantics, image targets, and audio sentiments into a unified structure by timestamps, which is beneficial for data integrity during upload.

[0122] Specifically, the integration process is achieved through a time synchronization mechanism. For example, text keywords at the same moment are bound to image frames and audio clips to form the first dataset, which is beneficial for efficient cloud processing. The first dataset is uploaded to the cloud processing platform, where large-scale computing resources are used to deeply extract the semantic information of the text. For example, a summary of key points of the meeting, such as "project schedule delays require budget adjustment," is generated from the initial keywords, which helps to produce concise and insightful output.

[0123] Specifically, in-depth refinement involves semantic expansion, such as expanding "budget allocation" to related influencing factors to form a summary of key points, which helps users quickly grasp the essence.

[0124] In one embodiment, the chart content and textual key points in the image target information are logically associated at the same time. For example, the bar chart data is matched with the text "budget allocation" to create a mapping such as the chart values ​​corresponding to the text description, which is helpful in revealing hidden relationships.

[0125] Specifically, logical connections are achieved through content similarity comparisons, such as comparing the matching degree between chart labels and text keywords to form a structured content mapping relationship, which helps enhance the visual support of the summary. Matching and locating audio sentiment information and video clips in the first dataset, for example, aligning the "emphasis" sentiment label with the speaker's gestures in the video, helps to accurately locate key moments.

[0126] Specifically, aligning audio emotional information with video content through timestamps, such as synchronizing audio peaks with video frame changes, and constructing an association index between audio and video, is beneficial for generating more dynamic analysis results and forming a second dataset.

[0127] In one embodiment, comparing a second dataset with a pre-established terminology database, such as extracting the term "AI model" from the abstract and matching it with the definition in the database, is beneficial for constructing a professional knowledge graph.

[0128] Specifically, the comparison process involves string matching and semantic querying, such as finding relevant nodes for "AI model" like "neural network," and constructing a knowledge graph framework, which is beneficial for deepening content understanding. Combining the knowledge graph framework with conference lecture summaries, such as integrating graph nodes into the summaries, forms structured results with links, which helps support the deep association between conference lecture summary generation and professional content, improving user decision-making efficiency.

[0129] The information application includes real-time information push, voice feedback, translation and annotation, and scenario-based assistance, including meeting minutes, to-do list annotation, disagreement learning, related literature, business facial expression and tone analysis strategies. Content generation supports direct sharing from glasses to cloud and local storage.

[0130] Real-time information streams are acquired from multimodal data acquisition devices. These streams include text content extracted from high-definition cameras, images and video clips captured by binocular cameras, and audio signals recorded by array microphones. These different types of data streams are initially classified and timestamped to ensure that subsequent processing can handle multiple information sources within the same scenario, generating a unified format of raw data. The raw data set is preprocessed: redundant symbols are removed from the text content and converted into structured text data; image and video clips undergo sharpness enhancement and motion blur correction; and audio signals undergo noise reduction and sound source localization, forming a cleaned and standardized data set to provide highly consistent basic information for subsequent scenario-based applications. This standardized data set is input into the information application process. In a meeting scenario, speaker content is extracted from the audio signal to generate a to-do list, and discrepancies are marked by comparing text and audio. In a business scenario, communication suggestions are generated by combining facial expressions in images and tone changes in audio, forming a scenario-based auxiliary information set. The contextualized auxiliary information set is pushed to the user via voice feedback. At the same time, the generated content can be directly shared to the cloud or local storage as a document. For real-time translation needs, foreign language content is converted into the user's language and annotated with voice, ensuring the immediacy and convenience of information transmission and completing the entire process from data acquisition to application output.

[0131] In one possible implementation, the process of acquiring real-time information streams from multimodal data acquisition devices can be accomplished through hardware components integrated into AI glasses. For example, a high-definition camera is responsible for capturing text in the scene, such as the content on a paper document, while a binocular camera records video clips of people's actions, and an array microphone captures the audio of conversations in a meeting. This ensures the comprehensiveness of the information stream because the initial classification processes text, images, and audio separately, and the timestamp synchronization makes them correspond to events at the same moment, thereby avoiding analytical biases caused by data chaos and facilitating the subsequent generation of accurate scene auxiliary information.

[0132] Specifically, the unified format raw data set generated after preliminary classification and timestamp synchronization serves as a foundation for preprocessing. For example, in a meeting scenario, if the raw data set includes the speaker's audio and corresponding video expressions, synchronization can accurately match tone and facial changes, which helps improve the reliability of data processing, avoids errors in labeling divergence points caused by asynchronous processing, and makes the overall application more real-time.

[0133] In one possible implementation, the preprocessing step of the raw dataset involves removing redundant symbols such as extraneous punctuation from text and converting it into structured text data. This makes the text easier for machines to read. For example, standardizing the content of handwritten whiteboards facilitates subsequent translation of foreign language documents. Meanwhile, enhancing the clarity of images and videos involves adjusting pixel contrast to correct motion blur. In business scenarios, the standardized dataset processed in this way can better identify customer facial expressions and is beneficial for generating accurate communication suggestions because the cleaning process reduces noise interference and improves data consistency.

[0134] Specifically, when standardized data sets are input into the information application process, the process of extracting speech content from audio to generate to-do list records for meeting scenarios includes parsing speech to text and identifying keywords such as "action items." At the same time, it compares the text record with the audio emotion to mark points of disagreement. For example, if the audio shows an excited tone but the text content is calm, it is marked as a potential conflict. This is consistent with the handling of business scenarios because facial features such as furrowed brows combined with changes in tone can generate suggestions such as "de-escalate the topic," forming a set of scenario-based auxiliary information that is beneficial for users to make quick decisions.

[0135] In one possible implementation, pushing contextualized auxiliary information sets to users via voice feedback involves converting analysis results into natural language broadcasts, such as real-time translation of foreign texts and voice annotation of object uses. It also supports direct sharing of content generation results, such as mind map documents, to the cloud. This ensures the immediacy of information delivery because the glasses device can seamlessly connect to the storage system, avoiding delays caused by manual operation, and is beneficial for the immediate application of analysis results in dynamic environments such as business negotiations.

[0136] Specifically, the complete process of ensuring the timeliness and convenience of information transmission, from data acquisition to output, enables the association of literature citations in practical applications such as scientific research conferences. This is because the pre-processed data supports in-depth analysis and forms a knowledge graph for push notifications. This is linked to privacy protection. By uploading only feature parameters rather than raw data through edge-cloud collaboration, user trust is enhanced, which is beneficial for promoting the use of AI glasses in multiple scenarios.

[0137] In the aforementioned end-to-cloud collaboration, simple tasks are completed locally while complex tasks are processed in the cloud. Federated learning uploads feature parameter raw data, which is encrypted and stored on the device, and strategies are dynamically adjusted to balance real-time computing power consumption.

[0138] In edge-cloud collaborative processing, tasks are initially categorized based on their complexity. Less computationally intensive tasks are assigned to local devices, while more complex tasks are uploaded to the cloud for processing. This categorization is based on a pre-established task complexity evaluation standard, determined by the amount of input data and the number of computational steps. After categorization, local devices only process simple tasks and generate preliminary results. For the preliminary results generated by local devices, federated learning is used to extract feature parameters. These feature parameters are obtained by dimensionality reduction of the preliminary results data. The original data is stored encrypted on the local device, while the feature parameters are uploaded to the cloud for further optimization and adjustment, ensuring data security while completing collaborative computation. Upon receiving the feature parameters, the cloud dynamically adjusts the allocation ratio of computing resources according to task processing needs. This allocation ratio is updated in real-time based on current network latency and the computing power status of local devices. If network latency exceeds a preset threshold, the proportion of local computing tasks is increased first; conversely, the proportion of cloud computing tasks is increased if network latency is low, resulting in an adjusted resource allocation scheme. Based on the adjusted resource allocation scheme, the processing flow of subsequent tasks is optimized. The optimization is achieved by updating the task classification standards and local device processing rules to ensure a balance between real-time performance and computing power consumption. The task allocation for end-to-cloud collaboration is continuously and dynamically adjusted to adapt to different scenario requirements.

[0139] Specifically, the generation steps are as follows:

[0140] In edge-cloud collaborative processing, the process of initially classifying tasks based on their complexity is called "cloud-edge".

[0141] Understandably, this categorization helps improve overall efficiency. For example, in real-time conferencing scenarios, simple tasks such as basic speech recognition can be assigned to local devices to avoid unnecessary network transmissions and thus reduce latency.

[0142] Specifically, the pre-established task complexity evaluation criteria are determined by comparing the amount of input data and the number of computation steps. For example, if the amount of input data is small, it is directly identified as a simple task and preliminary result data is generated. This can bring the beneficial effect of fast response because local processing does not need to wait for cloud feedback, thus supporting low latency requirements.

[0143] An implementation method for extracting feature parameters from preliminary result data generated by local devices using federated learning.

[0144] In one possible implementation, the preliminary result data is first subjected to dimensionality reduction processing to obtain feature parameters. For example, when analyzing image data of customer facial expressions, only the key facial expression vectors are extracted as feature parameters and uploaded, while the original image data is stored on the device in an encrypted manner. This not only protects privacy but also allows the cloud to optimize and adjust based on these parameters. The beneficial effect is to achieve secure collaborative computing. From multiple perspectives, this method can extract professional terms to construct knowledge graphs in scientific research scenarios because the cloud can further associate literature citation relationships after the feature parameters are uploaded, supporting personalized applications.

[0145] The specific process of dynamically adjusting the allocation ratio of computing resources based on task processing requirements after receiving feature parameters in the cloud.

[0146] Understandably, the allocation ratio is updated in real time based on the current network latency and the computing power status of local devices. For example, if the network latency exceeds a preset threshold, the proportion of local computing tasks is prioritized to increase. In emergency information pushes, for instance, simple annotations are processed locally, resulting in an adjusted resource allocation scheme. This approach achieves a beneficial balance between real-time performance and computing power consumption. From another perspective, in a meeting scenario, this helps to automatically record speakers' opinions and generate to-do lists, because after the adjustment, the cloud only processes complex emotional comparisons, reducing overall resource consumption. The adjusted resource allocation scheme is then used to optimize the processing flow of subsequent tasks.

[0147] One possible implementation is to ensure balance by updating task classification criteria and local device processing rules. For example, in business scenarios, optimization can generate communication strategy suggestions in real time and share them to the cloud via AI glasses. This approach has the beneficial effect of adapting to different scenario needs.

[0148] Understandably, this continuous dynamic adjustment supports end-to-cloud collaboration from multiple directions, such as prioritizing local tasks in low-latency scenarios to improve real-time performance, while relying on the cloud in complex tasks to optimize computing power, thus forming an efficient processing mechanism.

[0149] The method for synchronous acquisition of text, image, video, and audio data by sensors and spatiotemporal correlation indexing includes binding PPT text to corresponding video frames and the lecturer's voice, as well as using spatial coordinates to locate the user's gaze point to assist in judging the key information.

[0150] Raw information is acquired from multimodal data acquisition devices. This information includes text content captured by a high-definition camera, video frames recorded by a binocular camera, lecturer speech collected by an array microphone, and user gaze point positions determined by an eye-tracking sensor. This information is bound to timestamps to form a preliminary time-related data set. Content matching processing is performed on this time-related data set, comparing text content with video frames at corresponding timestamps and synchronously associating lecturer speech segments to form a three-dimensional content mapping relationship encompassing text, video, and speech, ensuring consistency of data from different sources in the temporal dimension. Based on this three-dimensional content mapping relationship, user gaze point data is combined with spatial coordinate positioning to determine the text or video areas that the user focuses on within a specific time period, generating attention highlight markers to help determine the importance of information. These attention highlight markers are integrated with the three-dimensional content mapping relationship to construct a complete spatiotemporal correlation index structure, ensuring accurate correspondence between text, video frames, and lecturer speech in the temporal and spatial dimensions, providing a foundation for subsequent content retrieval and analysis based on this index.

[0151] Specifically, the generation steps are as follows:

[0152] For example, the process of acquiring raw information from multimodal data acquisition devices can be understood as follows: text content captured by a high-definition camera, such as keywords on a document; video frames recorded by a binocular camera, such as dynamic presentation images; lecturer voice collected by an array microphone, such as narration segments; and the user's gaze point position determined by an eye-tracking sensor, such as focus coordinates. This information is then bound to timestamps to form a preliminary time-related data set. This ensures that the data is time-synchronized during the acquisition phase, which is beneficial for subsequent processing and avoids information gaps. In one possible implementation, the high-definition camera first scans the text content in the scene, such as the title on a PowerPoint presentation; the binocular camera simultaneously captures the image changes displayed in the video frames; the array microphone records the corresponding lecturer's voice explanation; and the eye-tracking sensor records the time when the user's gaze point falls on the text area. These are then bound to a unified timestamp, such as one marked in milliseconds, to form a time-related data set. This helps improve the accuracy of data integration because data from different modalities, once aligned in time, can better reflect the dynamics of the real scene, thus providing a reliable foundation for analysis.

[0153] For example, the process of content matching for time-related data sets involves comparing text content with video frames at corresponding timestamps, such as matching text descriptions in a PPT with charts appearing in a video, and simultaneously associating lecturer audio clips, such as key terms mentioned in the audio, to form a three-dimensional content mapping relationship that includes text, video, and audio. This ensures consistency of data from different sources in the time dimension, which enhances the semantic coherence of the data and helps avoid misunderstandings caused by isolated information. In one possible implementation, text content such as the keyword "product testing" is first extracted from the time-related data set and compared with video frames at the same timestamp, such as footage showing the testing process. After confirming a high degree of matching, the lecturer's audio clips, such as the recording of "the test needs to be completed this week," are associated to construct a three-dimensional content mapping relationship. This is beneficial for maintaining information integrity when dealing with complex scenarios because consistency in the time dimension allows multimodal data to mutually verify each other, improving the reliability of the overall analysis.

[0154] For example, by combining spatial coordinates with the three-dimensional content mapping relationship to locate user gaze points, it's possible to determine which text or video areas a user focuses on within a specific time period, such as when the gaze point lingers on a specific paragraph in a PowerPoint presentation. This generates attention markers to help determine the importance of information, highlighting user interests and facilitating personalized information extraction while avoiding irrelevant data interference. In one possible implementation, video frames and audio from the three-dimensional content mapping relationship are overlaid with spatial coordinates, such as the user's gaze point's xy position on the screen. If the attention duration exceeds a threshold, a marker is generated, such as highlighting the text area. This is helpful in emphasizing key points in educational scenarios because combining spatial data can accurately capture user behavior, making the determination of information importance more aligned with actual needs.

[0155] For example, the process of integrating attention markers with three-dimensional content mapping relationships to construct a complete spatiotemporal associative index structure ensures accurate correspondence between text, video frames, and speaker audio in both time and space dimensions. For instance, the index links text marked with gaze points to corresponding audio. This provides a foundation for subsequent content retrieval and analysis based on this index, enabling efficient queries and facilitating rapid location of key information while reducing processing latency. In one possible implementation, attention markers, such as key text labels, are first integrated into the three-dimensional content mapping relationship, and then expanded into a spatiotemporal associative index structure. Timestamps are bound to spatial coordinates to ensure accurate correspondence. This facilitates direct retrieval of gaze-related video frames and audio during meeting reviews, as the integration of spatiotemporal dimensions provides a comprehensive view, thus supporting in-depth analysis such as content summaries.

[0156] Based on the embodiments of the present invention described above, and through the above description, those skilled in the art can make various changes and modifications without departing from the technical concept of the present invention. The technical scope of the present invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A method for multimodal data acquisition and intelligent analysis based on AI glasses, characterized in that, include: AI glasses collect text data through cameras and OCR components, image and video data through binocular cameras and infrared sensors, audio data through array microphones, and auxiliary data through IMU and eye-tracking sensors. The AI ​​glasses preprocess the collected multimodal data, including text data cleaning and standardization, key frame extraction of images and videos, audio-to-text tagging features, and establish a timestamp and spatial coordinate association index to form a spatiotemporal content fusion dataset. The AI ​​glasses utilize a local lightweight AI model for basic analysis and upload complex tasks to a large cloud model for collaborative processing, including information summarization, in-depth mining, and knowledge extraction. The AI ​​glasses, based on the analysis results, use the information application module to achieve real-time push notifications, contextualized assistance, and content generation output. The AI ​​glasses employ an edge-cloud collaborative module and a federated learning framework to dynamically adjust local cloud processing strategies, uploading only feature parameters to ensure that the original data is retained locally.

2. The method as described in claim 1, characterized in that, The AI ​​glasses collect text data through a camera and OCR component, image and video data through a binocular camera and infrared sensor, audio data through an array microphone, and auxiliary data through an IMU and eye-tracking sensor, including: The camera and OCR component extract text information from the scene in real time, supporting multi-language recognition and tilt correction; The binocular camera and infrared sensor acquire dynamic video streams and static images, record the characteristics of objects, people and environment, and support night mode and motion blur correction. The array microphones collect speech signals, distinguish human voices from ambient sounds, and support sound source localization, noise reduction, and multi-channel separation. The IMU and eye-tracking sensor record the trajectory of head movement gaze points to determine the information the user is focusing on.

3. The method as described in claim 1, characterized in that, The AI ​​glasses preprocess the collected multimodal data, including text data cleaning and standardization, keyframe extraction from images and videos, audio-to-text tagging features, and establishing a timestamp-spatial coordinate association index to form a spatiotemporal content fusion dataset, including: Text data is converted into structured text by removing redundant symbols; Image and video data compression resolution extraction of keyframes; Audio data is converted into text to mark intonation and pause features; Text, image, and audio data are linked by timestamps and spatial coordinates, and a spatiotemporal content fusion dataset is formed by using the text position corresponding to the user's gaze point.

4. The method as described in claim 1, characterized in that, The AI ​​glasses utilize a local lightweight AI model for basic analysis and upload complex tasks to a large cloud model for collaborative processing, including information summarization, in-depth mining, and knowledge extraction, including: A local lightweight AI model enables text semantic extraction, image object detection, and audio emotion recognition; The cloud-based large model receives and uploads complex tasks, summarizes the information, and generates lecture minutes. Cloud-based large-scale model analysis identifies the relationship between text and images, and recognizes the logical relationships between text and charts in PPT presentations; The cloud-based big data model analyzes the matching degree between audio and video, and locates video demonstration segments based on the content of the speech; The cloud-based big data model extracts technical terms, formulas, and cases from multimodal data to build personalized knowledge graphs.

5. The method as described in claim 1, characterized in that, The AI ​​glasses, based on the analysis results, utilize an information application module to provide real-time push notifications, contextualized assistance, and content generation output, including: Real-time push notifications via voice feedback; Record the speakers' opinions in a meeting setting; Generate to-do items and mark the points of disagreement.

6. The method as described in claim 1, characterized in that, The AI ​​glasses, based on the analysis results, utilize an information application module to provide real-time push notifications, contextualized assistance, and content generation output, including: Contextualized assistance identifies difficult formulas in textbooks within learning scenarios and automatically associates them with video explanations and relevant literature. Contextualized assistance analyzes customer facial expressions and tone of voice in business scenarios to generate communication strategy suggestions; Content is generated and exported as documents, mind maps, and short videos, and shared to cloud local storage via AI glasses.

7. The method as described in claim 1, characterized in that, The AI ​​glasses employ an edge-cloud collaborative module and a federated learning framework to dynamically adjust local cloud processing strategies, uploading only feature parameters to ensure that raw data is retained locally, including: Through the federated learning framework, feature parameters are uploaded to the local model, while the raw data is stored on the device. Simple tasks can be dynamically adjusted to be completed locally, while complex tasks can be processed in the cloud.

8. The method as described in claim 2, characterized in that, The camera and OCR component extract text information from the scene in real time, supporting multilingual recognition and tilt correction, including: The OCR component recognizes text content from paper documents, electronic screens, and handwritten whiteboards; Based on the recognition results of the tilt correction process, the corrected text data is obtained.

Citation Information

Cited By

  • Intelligent interaction method, device and system

    CN122086223A

  • Intelligent audio glasses adaptive interaction method and device, equipment and medium

    CN122086249A