All-sensory data processing method and device based on AR glasses and electronic equipment

By integrating multimodal sensors and dynamic fusion networks on AR glasses, combining the time alignment model to identify synchronization events and generate feedback strategies, the problem of multi-sensory data information separation in complex scenarios is solved, and a more comprehensive and accurate scenario understanding and user experience is achieved.

CN120143986AInactive Publication Date: 2025-06-13GUANGZHOU GUDONG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510289791.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the prior art deals with complex scenarios, the lack of a unified feature extraction and fusion mechanism between multi-sensory data, resulting in information separation and the inability to form a complete scene cognition. Especially in complex environments where multimodal information requires joint judgment, it is difficult for traditional methods to efficiently correlate multiple modes.

Method used

By integrating multimodal sensors on AR glasses, they acquire all-sensory data such as vision, hearing, smell and touch, and use a multimodal dynamic fusion network to extract the data feature, combine the time alignment model to identify synchronous events, and generate multimodal feedback strategies based on the events to control AR glasses to execute feedback strategies.

Benefits of technology

It realizes a comprehensive perception of complex scenarios, improves the depth and breadth of information acquisition, solves the information fragmentation problem caused by modal independent processing in traditional systems, enhances the real-time and accuracy of the system, and significantly improves users' perception of key information and decision-making efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120143986A_ABST
    Figure CN120143986A_ABST
Patent Text Reader

Abstract

The invention provides a total sensory data processing method and device based on AR glasses and electronic equipment, and relates to the field of data processing. In the method, total sensory data for a target environment sent by a multi-modal sensor is obtained, and the multi-modal sensor is located on AR glasses; performing feature extraction on the total sensory data by adopting a multi-modal dynamic fusion network to obtain a visual feature, an auditory feature, an olfactory feature and a tactile feature; determining a synchronization event from the visual feature, the auditory feature, the olfactory feature and the tactile feature through a time alignment model; and according to the synchronization event, generating a multi-modal feedback strategy, and controlling the AR glasses to execute the multi-modal feedback strategy. By implementing the technical scheme provided by the invention, the use experience of a wearer of the AR glasses can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a full-sensory data processing method, device, and electronic device based on AR glasses. Background Art

[0002] With the rapid development of augmented reality technology, as a new type of human-computer interaction device, AR glasses are gradually being applied to multiple fields such as industrial manufacturing, medical assistance, and intelligent driving. In modern complex scenarios, users' demand for obtaining environmental information is no longer limited to a single sensory modality, but rather hopes to obtain a more comprehensive and accurate understanding of the scenario through the comprehensive perception of multi-sensory data such as vision, hearing, smell, and touch.

[0003] Currently, existing technologies usually rely on independent perception modules to process visual, auditory, olfactory, and tactile data separately. Although this independent processing method can complete the data analysis of a single modality, there is a lack of a unified feature extraction and fusion mechanism between different modalities, resulting in fragmented information and an inability to form a complete scene perception. Especially in complex environments where multi-modal information needs to be jointly judged, such as dangerous signal detection or task collaboration scenarios, traditional methods are difficult to efficiently associate multiple modalities, thus leading to a poor user experience for AR glasses wearers.

[0004] Therefore, there is an urgent need for a full-sensory data processing method, device, and electronic device based on AR glasses. Summary of the Invention

[0005] This application provides a full-sensory data processing method, device, and electronic device based on AR glasses, which is convenient for improving the user experience of AR glasses wearers.

[0006] In the first aspect of this application, a full-sensory data processing method based on AR glasses is provided. The method includes: obtaining full-sensory data for a target environment sent by a multi-modal sensor, where the multi-modal sensor is located on the AR glasses; using a multi-modal dynamic fusion network to extract features from the full-sensory data to obtain visual features, auditory features, olfactory features, and tactile features; determining synchronous events from the visual features, the auditory features, the olfactory features, and the tactile features through a time alignment model; generating a multi-modal feedback strategy according to the synchronous events, and controlling the AR glasses to execute the multi-modal feedback strategy.

[0007] By adopting the above technical solution, all-sensory data such as vision, hearing, smell, and touch are obtained through multi-modal sensors, realizing a comprehensive perception of the target environment. Compared with single-modal information processing, this method can capture key details in complex scenes more comprehensively, improving the depth and breadth of information acquisition. The multi-modal dynamic fusion network is used to extract features from all-sensory data, realizing unified modeling and correlation analysis of different sensory information. This solves the problem of information fragmentation caused by independent modal processing in traditional systems and lays a foundation for subsequent synchronous event recognition and feedback strategy generation. Through the time alignment model, the problem of time asynchrony caused by different acquisition frequencies and response speeds of multi-modal information is solved, and multi-modal synchronous events can be accurately identified. In this way, the critical moments in multi-modal information can be captured, enhancing the real-time performance and accuracy of the system. According to synchronous events, a multi-modal feedback strategy is dynamically generated to achieve personalized and multi-modal linkage feedback to the user. The system can select appropriate feedback modalities according to the type and priority of events, improving the interaction effect and user experience. By controlling the AR glasses to execute the multi-modal feedback strategy, integrating visual display, audio prompt, tactile feedback, and even odor release, etc., the information transmission is made more efficient, intuitive, and immersive. Especially in complex or high-risk scenarios, it can significantly improve the user's perception ability of key information and decision-making efficiency. Therefore, it is convenient to improve the usage experience of AR glasses wearers.

[0008] Optionally, the obtaining of the all-sensory data of the target environment sent by the multi-modal sensor specifically includes: receiving the visual modal data sent by the RGB + depth camera; receiving the auditory modal data sent by the microphone array; receiving the olfactory modal data sent by the odor sensor; receiving the tactile modal data sent by the pressure sensor; performing data synchronization, data cleaning, and standardization processing on the visual modal data, the auditory modal data, the olfactory modal data, and the tactile modal data to obtain the all-sensory data.

[0009] By adopting the above technical solutions, visual, auditory, olfactory, and tactile data are collected respectively through multiple sensors, achieving multi-sensory coverage of the target environment. This design ensures the comprehensiveness of environmental perception and provides rich basic information for subsequent data processing and analysis. During the multi-modal data processing, the time deviation problem caused by the differences in the acquisition frequencies and response speeds of different sensors is solved through synchronization operations. Ensuring the consistency of each modal data in the time dimension lays the foundation for accurately identifying multi-modal events. The data cleaning step is introduced to effectively remove the possible noise and redundant information during sensor acquisition and retain high-quality data. This helps to enhance the reliability of feature extraction and reduce the interference of error information on subsequent analysis. Through standardization processing, the scales and formats of different modal data are unified, solving the problem of different data dimensions and numerical ranges. This preprocessing operation facilitates the fusion of multi-modal data and improves the efficiency and accuracy of feature extraction and analysis. Data synchronization, cleaning, and standardization processing can effectively cope with complex and changing environments and improve the applicability and robustness of the system in dynamic scenarios. This ensures that reliable full-sensory data can still be obtained in an environment with high uncertainty.

[0010] Optionally, the multi-modal dynamic fusion network is used to extract features from the full-sensory data to obtain visual features, auditory features, olfactory features, and tactile features, which specifically include: according to the multi-modal dynamic fusion network, using ResNet+Transformer to extract the expression and scene feature vectors from the full-sensory data; based on the expression and scene feature vectors, determining the visual features; using the short-time Fourier transform to generate an audio time-frequency map from the full-sensory data; according to the multi-modal dynamic fusion network, extracting the spectral feature vectors from the audio time-frequency map through CNN; based on the spectral feature vectors, determining the auditory features.

[0011] By adopting the above technical solution, through the multi-modal dynamic fusion network, specialized feature extraction techniques are used for different modal data, making full use of the characteristics of each modal data, and achieving accurate and efficient feature extraction. Combining the powerful image feature extraction ability of ResNet and the spatio-temporal modeling ability of Transformer, it is able to accurately extract delicate expression changes and complex scene information in the visual modality. Through the deep fusion of expression and scene features, the integrity and accuracy of visual information understanding are improved. The audio signal is converted into a time-frequency diagram, realizing the effective combination of time-domain and frequency-domain information, making the sound feature extraction more comprehensive. Key features are extracted from the spectrogram through CNN to ensure that auditory information can be accurately modeled, providing high-quality feature vectors for subsequent modal fusion. The multi-modal dynamic fusion network comprehensively extracts and processes visual and auditory and other modal features within the same framework through joint modeling, solving the problem of information fragmentation between modalities in traditional multi-modal systems. In this way, not only can single-modal characteristics be captured, but also deep relationships between modalities can be explored, providing higher accuracy for the recognition of synchronous events.

[0012] Optionally, the multi-modal dynamic fusion network is used to extract features from the full-sensory data to obtain visual features, auditory features, olfactory features, and tactile features. Specifically, it further includes: according to the multi-modal dynamic fusion network, sparse coding is used to analyze and obtain an odor signal pattern vector from the full-sensory data; based on the odor signal pattern vector, the olfactory features are determined; according to the multi-modal dynamic fusion network, multi-scale convolution is used to identify a dynamic change feature vector from the full-sensory data; based on the dynamic change feature vector, the tactile features are determined.

[0013] By adopting the above technical solutions, sparse coding is used to analyze odor signals, and key pattern vectors are extracted from high-dimensional and complex olfactory data to accurately capture the characteristic patterns of odors. This method can effectively reduce redundant information and improve the efficiency and accuracy of olfactory feature extraction. Through the analysis of the odor signal pattern vectors, the system can achieve accurate modeling and classification of complex odor environments and enhance the environmental perception ability. Using multi-scale convolution to identify dynamic change features can capture the dynamic changes from subtle to large-scale in tactile data. This method is applicable to various tactile scenarios and enhances the system's perception ability of complex tactile events. Extracting tactile features based on the dynamic change feature vectors can achieve accurate parsing of the tactile modality and provide high-quality tactile data for multi-modal fusion. Combining the feature extraction mechanisms of the four modalities of vision, audition, olfaction, and touch, the method comprehensively covers the diversity of multi-sensory data; the special optimization for the olfactory and tactile modalities further improves the refinement degree of feature extraction. Through the feature extraction of olfactory and tactile data by the multi-modal dynamic fusion network, which maintains the same deep modeling framework as the vision and audition modalities, the problem of difficult collaborative processing caused by modality heterogeneity in traditional systems is solved. The correlation between multi-modal information is enhanced, providing important support for the accurate identification of subsequent synchronous events.

[0014] Optionally, determining synchronous events from the visual features, the auditory features, the olfactory features, and the tactile features through the time alignment model specifically includes: calculating the correlation between a first feature and a second feature in the same time window to obtain a correlation coefficient, where the first feature is any one of the visual features, the auditory features, the olfactory features, and the tactile features, and the second feature is any one of the visual features, the auditory features, the olfactory features, and the tactile features other than the first feature; determining the magnitude relationship between the correlation coefficient and a preset threshold; if the correlation coefficient is greater than or equal to the preset threshold, determining that the first feature and the second feature belong to the synchronous event.

[0015] By adopting the above technical solutions, the time correlation between modalities is clarified by calculating the correlation of different modality features within the same time window, thereby effectively solving the problem of asynchronous time of multimodal data and improving the accuracy of event recognition. The method allows for correlation analysis between any two modality features, adapts to possible combinations of modality features in different scenarios, and enhances the applicability and generality of the system. According to the correlation between different modalities, important feature modalities can be flexibly determined to avoid information fragmentation between modalities or over-reliance on a single modality. Using the correlation coefficient as an indicator provides a quantitative way to judge the synchronization between modality features, making the recognition of synchronous events more reliable and controllable. By setting a correlation coefficient threshold, noise and irrelevant feature matches are effectively filtered, ensuring that only highly correlated modality pairs are recognized as synchronous events, and improving the robustness of the system.

[0016] Optionally, the correlation between the first feature and the second feature within the same time window is calculated to obtain a correlation coefficient, and the following specific calculation formula is adopted: ; where C ij is the correlation coefficient, i is the first feature, j is the second feature, F i is the first feature vector, F j is the second feature vector, Cov(F i , F j ) is the covariance of the first feature vector and the second feature vector, Var(F i ) is the variance of the first feature vector, and Var(F j ) is the variance of the second feature vector.

[0017] By adopting the above technical solutions, through the standardization of covariance and variance, the relationship between different modality features is quantified using the correlation coefficient, ensuring the consistency and scientific nature of the correlation calculation. The formula is based on eigenvector operations and does not depend on the form of specific modality data, having wide adaptability. The correlation calculation can effectively capture the linear relationship between different modality features, providing a basis for the recognition of synchronous events. By describing the joint change trend of two modality features through covariance, the correlation between modalities can be better reflected, rather than relying solely on the characteristics of individual modalities. The standardization process of variance excludes the influence of the magnitude difference of single modality data, ensuring the fairness of the comparison of feature relationships.

[0018] Optionally, generating a multimodal feedback strategy according to the synchronization event and controlling the AR glasses to execute the multimodal feedback strategy specifically includes: determining a target category corresponding to the synchronization event according to a preset category, where the preset category includes a dangerous event, a notification event, and an interaction event; generating the multimodal feedback strategy based on the target category, and by controlling the AR glasses to execute the multimodal feedback strategy to prompt the wearer of the AR glasses.

[0019] By adopting the above technical solution, by classifying the synchronization event as a dangerous event, a notification event, or an interaction event, the system can select the most appropriate feedback strategy according to the category, ensuring the pertinence and effectiveness of the feedback. The classification mechanism covers a variety of common event scenarios, can adapt to complex and changeable application environments, and improves the flexibility of the system. The synchronization event is directly associated with the target category, ensuring that the generated feedback strategy can accurately match the current event requirements and avoid redundant or irrelevant feedback prompts. The feedback content is more in line with the event characteristics, providing clear and definite prompts for the wearer, reducing unnecessary interference, and improving the interaction experience. The generated feedback strategy can flexibly call multimodal feedback, enhancing the richness and sensory coverage of the feedback information. Based on the target category, the system can reasonably design the combination method of multimodal feedback, making different modal information complement each other and improving the feedback efficiency and effect. Automatically matching the target category based on the synchronization event and generating a feedback strategy without manual intervention realizes intelligent feedback control of the entire process. By quickly controlling the AR glasses to execute the feedback strategy, the system can prompt the wearer at the first time when the event occurs, ensuring timely response.

[0020] In a second aspect of the present application, a full-sensory data processing device based on AR glasses is provided. The full-sensory data processing device includes an acquisition module and a processing module. Among them, the acquisition module is used to acquire full-sensory data of a target environment sent by a multimodal sensor, and the multimodal sensor is located on the AR glasses; the processing module is used to extract features from the full-sensory data by using a multimodal dynamic fusion network to obtain visual features, auditory features, olfactory features, and tactile features; the processing module is further used to determine a synchronization event from the visual features, the auditory features, the olfactory features, and the tactile features through a time alignment model; the processing module is further used to generate a multimodal feedback strategy according to the synchronization event and control the AR glasses to execute the multimodal feedback strategy.

[0021] In a third aspect of the present application, an electronic device is provided. The electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, and both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method described above.

[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions, and when the instructions are executed, the method described above is executed.

[0023] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: Through multimodal sensors, full sensory data such as vision, hearing, smell, and touch are obtained to achieve comprehensive perception of the target environment. Compared with single-modal information processing, this method can more comprehensively capture key details in complex scenes and improve the depth and breadth of information acquisition. The multimodal dynamic fusion network is used to extract features from all sensory data, realizing unified modeling and correlation analysis of different sensory information. This solves the problem of information fragmentation caused by independent modal processing in traditional systems, laying the foundation for subsequent synchronous event identification and feedback strategy generation. Through the time alignment model, the time asynchrony problem caused by different multimodal information acquisition frequencies and response speeds is solved, and multimodal synchronous events can be accurately identified. In this way, key moments in multimodal information can be captured, enhancing the real-time and accuracy of the system. According to the synchronous events, multimodal feedback strategies are dynamically generated to achieve personalized and multimodal linkage feedback for users. The system can select the appropriate feedback modality according to the type and priority of the event to improve the interactive effect and user experience. By controlling AR glasses to execute multimodal feedback strategies, integrating visual display, audio prompts, tactile feedback and even odor release, information transmission is made more efficient, intuitive and immersive. Especially in complex or high-risk scenarios, it can significantly improve users' perception of key information and decision-making efficiency, thus facilitating the improvement of the user experience of AR glasses wearers. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A schematic diagram of a flow chart of a method for processing full sensory data based on AR glasses provided in an embodiment of the present application; Figure 2 Another schematic diagram of a flow chart of a full sensory data processing method based on AR glasses provided in an embodiment of the present application; Figure 3 A schematic diagram of a module of a full sensory data processing device based on AR glasses provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0025] Explanation of the reference numerals: 31, acquisition module; 32, processing module; 41, processor; 42, communication bus; 43, user interface; 44, network interface; 45, memory. DETAILED DESCRIPTION

[0026] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments.

[0027] In the description of the embodiments of this application, words such as "for example" or "for illustration" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for illustration" is intended to present the relevant concepts in a specific manner.

[0028] In the description of the embodiments of this application, the term "plurality" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0029] With the rapid development of augmented reality technology, AR glasses, as a new type of human-computer interaction device, are being widely used in many fields such as industrial manufacturing, medical assistance, and intelligent driving. In modern complex scenarios, users' demand for obtaining environmental information has shifted from a single sensory modality to multi-sensory comprehensive perception, hoping to obtain a more comprehensive and accurate understanding of the scene through the fusion of multi-modal data such as vision, hearing, smell, and touch.

[0030] However, existing technologies usually rely on independent perception modules to process data of different modalities separately. Although this separate processing method can complete the analysis of single-modal information, it lacks cross-modal feature extraction and fusion mechanisms, resulting in the difficulty of forming effective linkage of multi-modal information and incomplete scene cognition. Especially in complex scenarios that require multi-modal collaborative judgment, such as dangerous signal detection or task collaboration, the modality fragmentation of traditional methods limits the efficiency of information association, making it difficult to meet the requirements of real-time and accuracy, thus reducing the usage experience and operation effect of AR glasses wearers.

[0031] To solve the above technical problems, this application provides a full-sensory data processing method based on AR glasses, referring to Figure 1 , Figure 1Schematic flowchart of a full-sensory data processing method based on an AR glasses provided by an embodiment of this application. This processing method is applied to a server and specifically includes steps S110 to S140. The above steps are as follows: S110. Obtain full-sensory data for a target environment sent by a multi-modal sensor, where the multi-modal sensor is located on the AR glasses.

[0032] Specifically, as the core node for data processing, the server receives environmental data collected by the multi-modal sensors on the AR glasses. This means that the AR glasses are not only responsible for collecting data but also uploading this data to the server for further processing. The multi-modal sensor is a hardware module on the AR glasses used to capture different types of information in the environment. The data collection is based on specific targets or scenarios in the environment, such as dangerous signal detection, operation task assistance, etc. The multi-modal sensors work together to generate full-sensory data, providing a basis for subsequent intelligent analysis.

[0033] For example, in a chemical plant environment, the RGB camera on the AR glasses captures real-time images of the pipeline appearance to detect whether there are cracks or leakage marks. The microphone array collects the operating sounds of the equipment to identify abnormal noises, such as the high-frequency vibration sound of the pump. The odor sensor detects whether there are leaking toxic gases in the air, such as methane or ammonia. The pressure sensor senses whether there is vibration or external impact acting on the AR glasses. The server receives these multi-modal data from the AR glasses, analyzes and finds that the gas concentration has increased abnormally, accompanied by specific noises and cracks on the pipeline surface, and finally generates an alarm to prompt chemical plant employees that there may be a leakage hazard.

[0034] In a possible implementation manner, obtaining full-sensory data for a target environment sent by a multi-modal sensor specifically includes: receiving visual modal data sent by an RGB + depth camera; receiving auditory modal data sent by a microphone array; receiving olfactory modal data sent by an odor sensor; receiving tactile modal data sent by a pressure sensor; performing data synchronization, data cleaning, and standardization processing on the visual modal data, auditory modal data, olfactory modal data, and tactile modal data to obtain full-sensory data.

[0035] Specifically, visual modality data is acquired by an RGB + depth camera. Such a camera can capture color images of the environment and depth information related to the distance from the camera. For example, in a room, RGB data can be used to identify the color and shape of furniture, while depth data can help measure the distance between the furniture and the camera. Auditory modality data is acquired by a microphone array. The microphone array can capture sounds and determine the direction of the sound source. For example, in a street scene, the honking of a car can be captured and the location of the sound source can be determined. Olfactory modality data is acquired by odor sensors. Such sensors detect the gas components in the environment through chemical analysis. For example, in a kitchen scene, the odor sensors can detect the smell of coffee or toasted bread. Tactile modality data is acquired by pressure sensors. Such sensors can sense touch or pressure. For example, when touching a virtual object, the pressure sensors can sense the applied force and feedback the tactile information. Among them, the above sensors are all integrated on the AR glasses.

[0036] Since the sampling frequencies of different sensors may vary, such as the camera capturing 30 frames per second and the microphone capturing thousands of audio samples per second, it is necessary to perform time synchronization on the multi-modal data to ensure that they correspond to events at the same time point. Data cleaning is used to remove noise or outliers in the data. For example, blurry images in visual data are removed, and background noise in auditory data is filtered out. Normalization processing transforms the multi-modal data into a unified representation form so that it can be processed by subsequent models. For example, RGB images are transformed into tensors of a fixed size, and audio data is transformed into time-frequency diagrams.

[0037] S120. Use a multi-modal dynamic fusion network to extract features from the full-sensory data to obtain visual features, auditory features, olfactory features, and tactile features.

[0038] Specifically, the multimodal dynamic fusion network is a deep learning network structure that can process multiple modal data at the same time, such as vision, hearing, smell and touch. It can adaptively adjust the fusion strategy of each modal feature according to the characteristics of different modalities. By extracting and fusing features from multiple senses, a unified representation is generated to facilitate subsequent decision-making or feedback. Visual features are key features of scenes, objects or expressions extracted from RGB+depth images. For example, the texture of a beach and the dynamics of waves are identified. Auditory features are the spectrum, timing information or sound source location of sounds extracted from audio signals. For example, the sound of waves and wind on the beach are identified. Olfactory features are chemical pattern features extracted from odor signals. For example, the smell of seawater in the salty and humid air in the beach environment is analyzed. Tactile features are mechanical properties extracted from pressure or touch data. For example, the soft touch of the user's feet on the beach is detected. The server centrally processes data from multiple modal sensors and uses the multimodal dynamic fusion network to complete unified feature extraction. By extracting and fusing features of each modality, key information in complex scenes is identified, and the association understanding of multimodal data is achieved.

[0039] In one possible implementation, a multimodal dynamic fusion network is used to perform feature extraction on all sensory data to obtain visual features, auditory features, olfactory features, and tactile features, specifically including: according to the multimodal dynamic fusion network, ResNet+Transformer is used to extract expression and scene feature vectors from all sensory data; based on the expression and scene feature vectors, visual features are determined; short-time Fourier transform is used to generate an audio time-frequency diagram from all sensory data; according to the multimodal dynamic fusion network, CNN is used to extract spectral feature vectors from the audio time-frequency diagram; based on the spectral feature vectors, auditory features are determined.

[0040] Specifically, ResNet+Transformer is used for visual feature extraction. ResNet is a powerful convolutional neural network architecture that focuses on extracting visual features from images, such as edges, textures, object shapes, etc. Transformer is a network that is good at processing temporal relationships and is used to capture contextual information of scenes or expressions. First, local detail features are extracted from RGB images using ResNet. Then, high-level features such as the overall structure of the scene and the dynamic changes in the user's expression are extracted by using Transformer in combination with contextual information, thereby generating feature vectors representing the scene and expression.

[0041] Auditory feature extraction is performed through short-time Fourier transform, that is, converting the time-domain signal (such as audio) into a frequency-domain signal to generate a time-frequency graph, which is used to represent the distribution of the audio signal in time and frequency. CNN is used to extract key spectral features from the time-frequency graph. First, the audio signal is converted into a time-frequency graph through short-time Fourier transform. Then, spectral features such as the energy change and rhythm pattern of specific frequencies are extracted from the time-frequency graph using CNN, thereby generating a spectral feature vector for auditory analysis. Visual features are extracted through expression and scene feature vectors, representing the key information of the visual modality. Auditory features are extracted through spectral feature vectors, representing the sound information of the auditory modality.

[0042] For example, assume the scenario is experiencing multimodal feedback on the beach. The RGB+depth camera captures the beach scene, including the blue sea, yellow sand, and the dynamic changes in the sky. ResNet extracts the edge features of the waves, the texture features of the sand, and the facial expressions of people (such as smiling). The Transformer analyzes the time context and extracts the dynamic changes in the scene (such as the continuous movement of the waves). Among them, the expression feature vector represents the feature of the user's smile. The scene feature vector represents the overall layout of the waves, sand, and coconut trees. The sound signals collected by the microphone array include the sound of the waves, the wind, and the cries of seagulls. A time-frequency graph is generated through short-time Fourier transform, showing the change in the frequency distribution of the sound signal over time. CNN extracts key features from the time-frequency graph, such as the low-frequency noise of the waves, the high-frequency calls of seagulls, and the continuous distribution of the wind sound. The spectral feature vector represents the characteristics and rhythm of different sounds in the environment. The multimodal dynamic fusion network combines the visual features (scene and expression) and the auditory features (sound spectrum) to generate a multimodal feature vector representing the current beach environment. The server generates an immersive feedback strategy (such as enhancing the wave dynamics in AR glasses and playing the real sound of the waves) that matches the beach scene based on these features.

[0043] In a possible implementation, a multimodal dynamic fusion network is used to extract features from the full-sensory data, obtaining visual features, auditory features, olfactory features, and tactile features. Specifically, it further includes: analyzing the odor signal pattern vector from the full-sensory data using sparse coding according to the multimodal dynamic fusion network; determining the olfactory features based on the odor signal pattern vector; identifying the dynamic change feature vector from the full-sensory data using multi-scale convolution according to the multimodal dynamic fusion network; and determining the tactile features based on the dynamic change feature vector.

[0044] Specifically, sparse coding is used for olfactory feature extraction. Sparse coding is a signal representation method that aims to decompose complex signals into a set of sparse basis vectors (i.e., vectors with only a few non-zero elements). In odor analysis, it can be used to extract the features of key odor components. From the data of odor sensors, key odor patterns are extracted using the sparse coding algorithm, such as the combination patterns of different odor molecules. These patterns are represented as odor signal pattern vectors, thereby generating olfactory feature vectors that can represent specific odor characteristics. Multi-scale convolution is a deep learning method for signal processing that combines convolution kernels of different scales and can capture dynamic tactile features. In tactile analysis, it is used to identify dynamic changes such as pressure intensity and vibration frequency.

[0045] From the data of pressure sensors, dynamic change features (such as vibration patterns or force distributions) in tactile signals are extracted using multi-scale convolution, and these dynamic changes are represented as feature vectors of tactile signals, thereby generating tactile feature vectors that can characterize dynamic tactile perception. Olfactory features are odor signal pattern vectors generated by sparse coding. Tactile features are dynamic change feature vectors extracted by multi-scale convolution.

[0046] For example, the scenario is multi-modal feature extraction in a garden experience. Odor sensors capture various odor components in the garden, such as the fragrance of flowers, the smell of plants, and the odor of moist soil. Using sparse coding, the odor signals are decomposed into specific pattern vectors. The main components are extracted, such as the characteristic components of the fragrance of flowers (e.g., jasmine, rose) and the component of soil moisture. The olfactory feature vector represents the combined characteristics of different odor components in the garden, such as {jasmine fragrance intensity = 0.8, soil odor = 0.5}. Tactile data captured by pressure sensors, such as the force when the user's hand touches a flower and the subtle vibration when the wind blows on the skin. The tactile signals are analyzed by multi-scale convolution to extract dynamic changes at different forces (such as the low-frequency signal of a gentle touch and the high-frequency signal of the vibration of the wind). The key patterns in the tactile sense are identified (such as the length of the touch time, the intensity of the vibration). The tactile feature vector represents the tactile characteristics of the dynamic changes, such as {touch pressure = 0.3, vibration frequency = 0.7}. The multi-modal dynamic fusion network combines the olfactory features (odor vectors) and tactile features (dynamic change vectors) to generate a comprehensive feature representation describing the garden experience. Based on these features, the AR glasses can provide immersive feedback of the real scene: olfactory simulation uses an odor diffusion module to release the fragrance of jasmine. Tactile simulation reproduces the slight touch of the wind on the back of the hand through a tactile feedback device.

[0047] S130. Determine synchronous events from visual features, auditory features, olfactory features, and tactile features through a time alignment model.

[0048] Specifically, the time alignment model is an algorithm for analyzing the temporal relationships of multimodal data, used to detect the synchronization characteristics between different modalities. By identifying whether the time points at which different modality features such as visual, auditory, olfactory, and tactile occur are highly correlated, it is determined whether they belong to the same event. Calculate the correlation (such as covariance or correlation coefficient) of each modality feature within the same time window. Judge whether the correlation exceeds a set threshold (i.e., high correlation means these features belong to the same event). A synchronous event refers to an event in which the perceptual signals of different modalities occur simultaneously or have a high temporal correlation in multimodal features. Synchronous events can help the system determine whether the sources of multimodal data belong to the same scene or interaction behavior, thereby enhancing the system's ability to understand the context of the data. The server calculates the correlation between visual features, auditory features, olfactory features, and tactile features through the time alignment model. If the correlation exceeds the threshold, it is considered that these features belong to a synchronous event.

[0049] In a possible implementation, the time alignment model is used to determine synchronous events from visual features, auditory features, olfactory features, and tactile features, specifically including: calculating the correlation between a first feature and a second feature within the same time window to obtain a correlation coefficient, where the first feature is any one of the visual feature, auditory feature, olfactory feature, and tactile feature, and the second feature is any one of the visual feature, auditory feature, olfactory feature, and tactile feature other than the first feature; judging the magnitude relationship between the correlation coefficient and a preset threshold; if the correlation coefficient is greater than or equal to the preset threshold, it is determined that the first feature and the second feature belong to a synchronous event.

[0050] Specifically, the time window sets a short time period (such as a few milliseconds or seconds) to compare whether multimodal features are highly correlated within this time period. Any two features of the first feature and the second feature are paired (for example, visual and auditory, auditory and olfactory, olfactory and tactile). If the correlation coefficient exceeds the set threshold (such as 0.8), it is considered that these two features are highly synchronized in time. The preset threshold is set customarily in advance. When the correlation coefficients of two or more features all exceed the threshold, it indicates that they may originate from the same event, that is, they belong to a synchronous event.

[0051] In a possible implementation, the correlation between the first feature and the second feature within the same time window is calculated to obtain a correlation coefficient, and the following calculation formula is specifically adopted: ; Specifically, C ij is the correlation coefficient, with a range of (-1, 1). A positive value indicates a positive correlation, a negative value indicates a negative correlation, and the larger the absolute value, the stronger the correlation. i is the first feature, j is the second feature, F i is the first feature vector, F jis the second eigenvector, Cov(F i , F j ) is the covariance of the first eigenvector and the second eigenvector, representing the strength of the linear relationship between the two-modal features. Var(F i ) is the variance of the first eigenvector, and Var(F j ) is the variance of the second eigenvector, which is used to represent the distribution range of the features.

[0052] For example, assume the scenario is that a user is experiencing a fountain in a park. The user wears AR glasses and stands beside the fountain in the park. The fountain is gushing, and the user's visual, auditory, olfactory, and tactile sensors respectively collect the following information: Vision: See the fountain water column rising and water droplets splashing. Auditory: Hear the impact sound of the fountain water flow and the sound of water falling. Olfactory: Smell the smell of plants and water vapor mixed in the moist air. Tactile: Feel the coolness of the water droplets splashing from the fountain gently falling on the skin. The visual feature extracts the dynamic changes of the water column from the camera image to form an eigenvector. The auditory feature captures the impact sound of the water flow from the microphone array to form a spectral eigenvector. The olfactory feature extracts the specific odor signal in the moist air from the odor sensor to form an odor eigenvector. The tactile feature detects the vibration of the water droplets contacting the skin from the pressure sensor to form a tactile eigenvector. The server selects a time window (e.g., 1 second) to align the timestamps of the visual, auditory, olfactory, and tactile features.

[0053] For example, the correlation coefficient between vision and audition (the water flow and sound of the fountain) is 0.9. The correlation coefficient between olfaction and touch (the moisture of the water droplets and the sense of contact) is 0.85. The correlation coefficient between audition and olfaction (the sound of water and the smell of moist air) is 0.88. If all the calculated correlation coefficients are higher than the set threshold (e.g., 0.8), it is determined that these features belong to a synchronous event. Among them, the scene features of the fountain are recognized as a synchronous event by the server. By determining the strong correlation between all features in the fountain scene, it shows that they originate from the same event, which helps the server understand the user's current scene and enhance the context awareness ability. Finally, the server can, according to the type of synchronous event (such as the experience of natural scenery), control the AR glasses to enhance the dynamic effect of water flow visually; play a clearer fountain sound effect auditorily; provide olfactory feedback by releasing the smell of moist air through an odor simulation device; simulate the feeling of water droplet contact on the tactile feedback device. Through the time alignment model, the server can accurately identify synchronous events of different modalities, solving the problems of multi-modal information fragmentation and difficult time sequence alignment. This not only improves the relevance and consistency of multi-modal data but also provides more intelligent and immersive feedback capabilities for AR devices.

[0054] S140. Generate a multi-modal feedback strategy according to the synchronous event, and control the AR glasses to execute the multi-modal feedback strategy.

[0055] Specifically, according to the type of synchronization event (such as a danger event, a notification event, or an interaction event), an appropriate multi-modal feedback strategy is formulated. The strategy is based on the real-time requirements of the user's current scenario and a preset response mechanism (such as prompting, guiding, enhancing, etc.). The visual feedback is enhanced visual information displayed by AR. The auditory feedback is playing an audio prompt or background sound. The olfactory feedback is releasing an odor related to the scenario. The tactile feedback is providing vibration or pressure simulation. The strategy generated by the server is transmitted to the AR glasses through the network, and the AR glasses perform specific operations (such as display, sound effect playback, odor release, tactile simulation, etc.). The feedback form is adjusted in real time according to the user's current needs to ensure accurate feedback.

[0056] In a possible implementation manner, with reference to Figure 2 , Figure 2 FIG. is another schematic flowchart of a full-sensory data processing method based on AR glasses provided by an embodiment of the present application, specifically including steps S210 to S220. The above steps are as follows: S210, determining a target category corresponding to the synchronization event according to a preset category, where the preset category includes a danger event, a notification event, and an interaction event; S220, generating a multi-modal feedback strategy based on the target category, and prompting the wearer of the AR glasses by controlling the AR glasses to execute the multi-modal feedback strategy.

[0057] Specifically, the synchronization event is classified into three preset categories according to its nature. A danger event may pose a threat to the user's safety (such as approaching an obstacle, a dangerous animal). A notification event is important information that needs to be transmitted (such as receiving a message reminder, environmental change). An interaction event is a scenario where the user needs to actively participate (such as navigation guidance, virtual object interaction). Based on the specific requirements of the target category, a feedback strategy is formulated, including visual, auditory, olfactory, and tactile feedback. The strategy will be dynamically adjusted according to the content of the synchronization event to ensure the real-time and effectiveness of the feedback. The AR glasses, as the output terminal of the feedback, perform corresponding operations according to the strategy generated by the server to prompt the wearer on how to respond to the current scenario.

[0058] For example, when a user wears AR glasses and walks on the street, an unnoticed obstacle (such as a trash can or a pothole on the ground) suddenly appears in front, which may cause the user to trip. The visual feature is that the camera captures the obstacle in front. The auditory feature is that the environmental sound analysis shows that the footsteps are approaching the obstacle. The olfactory and tactile features have no significant impact, but the event can be judged by the combination of vision and hearing. The server classifies this event as a "dangerous event". The visual feedback is to highlight the position of the obstacle on the display interface of the AR glasses and display a safe path. The auditory feedback is to play a warning sound (such as "There is an obstacle ahead, please pay attention"). No olfactory feedback is required. The tactile feedback is to remind the user to slow down or turn by vibration. The AR glasses immediately display a path correction suggestion when the user approaches the obstacle, and at the same time issue a vibration warning and a voice prompt.

[0059] The present application also provides a full-sensory data processing device based on AR glasses. Referring to Figure 3 , Figure 3 is a schematic diagram of the modules of a full-sensory data processing device based on AR glasses provided by an embodiment of the present application. The full-sensory data processing device is a server, and the server includes an acquisition module 31 and a processing module 32. Among them, the acquisition module 31 acquires full-sensory data for a target environment sent by a multi-modal sensor, and the multi-modal sensor is located on the AR glasses; the processing module 32 uses a multi-modal dynamic fusion network to extract features from the full-sensory data to obtain visual features, auditory features, olfactory features, and tactile features; the processing module 32 determines synchronous events from the visual features, auditory features, olfactory features, and tactile features through a time alignment model; the processing module 32 generates a multi-modal feedback strategy according to the synchronous events and controls the AR glasses to execute the multi-modal feedback strategy.

[0060] In a possible implementation manner, the acquisition module 31 acquires full-sensory data for a target environment sent by a multi-modal sensor, specifically including: the acquisition module 31 receives visual modal data sent by an RGB + depth camera; the acquisition module 31 receives auditory modal data sent by a microphone array; the acquisition module 31 receives olfactory modal data sent by an odor sensor; the acquisition module 31 receives tactile modal data sent by a pressure sensor; the processing module 32 performs data synchronization, data cleaning, and normalization processing on the visual modal data, auditory modal data, olfactory modal data, and tactile modal data to obtain full-sensory data.

[0061] In a possible implementation, the processing module 32 uses a multi-modal dynamic fusion network to extract features from the full-sensory data, obtaining visual features, auditory features, olfactory features, and tactile features. Specifically, it includes: The processing module 32 extracts the expression and scene feature vectors from the full-sensory data using ResNet+Transformer according to the multi-modal dynamic fusion network; The processing module 32 determines the visual features based on the expression and scene feature vectors; The processing module 32 generates an audio time-frequency map from the full-sensory data using the short-time Fourier transform; The processing module 32 extracts the spectral feature vectors from the audio time-frequency map through CNN according to the multi-modal dynamic fusion network; The processing module 32 determines the auditory features based on the spectral feature vectors.

[0062] In a possible implementation, the processing module 32 uses a multi-modal dynamic fusion network to extract features from the full-sensory data, obtaining visual features, auditory features, olfactory features, and tactile features. Specifically, it further includes: The processing module 32 analyzes and obtains the odor signal pattern vectors from the full-sensory data using sparse coding according to the multi-modal dynamic fusion network; The processing module 32 determines the olfactory features based on the odor signal pattern vectors; The processing module 32 identifies the dynamic change feature vectors from the full-sensory data using multi-scale convolution according to the multi-modal dynamic fusion network; The processing module 32 determines the tactile features based on the dynamic change feature vectors.

[0063] In a possible implementation, the processing module 32 determines synchronous events from the visual features, auditory features, olfactory features, and tactile features through a time alignment model. Specifically, it includes: The processing module 32 calculates the correlation between the first feature and the second feature in the same time window to obtain the correlation coefficient. The first feature is any one of the visual features, auditory features, olfactory features, and tactile features, and the second feature is any one of the visual features, auditory features, olfactory features, and tactile features other than the first feature; The processing module 32 determines the magnitude relationship between the correlation coefficient and the preset threshold; If the correlation coefficient is greater than or equal to the preset threshold, the processing module 32 determines that the first feature and the second feature belong to synchronous events.

[0064] In a possible implementation, the processing module 32 calculates the correlation between the first feature and the second feature in the same time window to obtain the correlation coefficient. The specific calculation formula is as follows: ; where C ij is the correlation coefficient, i is the first feature, j is the second feature, F i is the first feature vector, F j is the second feature vector, Cov(F i , F j) is the covariance of the first eigenvector and the second eigenvector, Var(F i ) is the variance of the first eigenvector, Var(F j ) is the variance of the second eigenvector.

[0065] In a possible implementation, the processing module 32 generates a multimodal feedback policy according to the synchronization event and controls the AR glasses to execute the multimodal feedback policy, specifically including: the processing module 32 determines the target category corresponding to the synchronization event according to the preset categories, and the preset categories include dangerous events, notification events, and interaction events; the processing module 32 generates a multimodal feedback policy based on the target category and controls the AR glasses to execute the multimodal feedback policy to prompt the wearer of the AR glasses.

[0066] It should be noted that: when the device provided in the above embodiment realizes its functions, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.

[0067] This application also provides an electronic device. Refer to Figure 4 , Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of this application. The electronic device may include: at least one processor 41, at least one network interface 44, a user interface 43, a memory 45, and at least one communication bus 42.

[0068] Among them, the communication bus 42 is used to realize the connection and communication between these components.

[0069] Among them, the user interface 43 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 43 may further include a standard wired interface and a wireless interface.

[0070] Among them, the network interface 44 may optionally include a standard wired interface and a wireless interface (such as Wi-Fi, interface).

[0071] Among them, the processor 41 may include one or more processing cores. The processor 41 connects various parts within the entire server through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 45, and by invoking the data stored in the memory 45, it performs various functions of the server and processes data. Optionally, the processor 41 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 41 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 41 and may be implemented separately by a single chip.

[0072] Among them, the memory 45 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 45 includes a non-transitory computer-readable storage medium. The memory 45 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 45 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory 45 may also be at least one storage device located far from the aforementioned processor 41. As Figure 4 shown, the memory 45, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a full-sensory data processing method based on AR glasses.

[0073] In Figure 4In the electronic device shown, the user interface 43 is mainly used to provide an interface for the user to input and obtain the data input by the user; while the processor 41 can be used to call an application program stored in the memory 45, which is a method for processing all-sensory data based on AR glasses. When executed by one or more processors, the electronic device executes one or more of the methods in the above embodiments.

[0074] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0075] The present application also provides a computer-readable storage medium storing instructions. When executed by one or more processors, the electronic device executes one or more of the methods described in the above embodiments.

[0076] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0077] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of the device or unit can be in electrical or other forms.

[0078] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0079] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0080] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.

[0081] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and the disclosure of the practical truth, those skilled in the art will easily think of other implementation manners of the present disclosure. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A full sensory data processing method based on AR glasses, characterized in that: The method comprises: Acquire full sensory data for the target environment sent by a multimodal sensor, wherein the multimodal sensor is located on the AR glasses; Using a multimodal dynamic fusion network to extract features from the full sensory data to obtain visual features, auditory features, olfactory features, and tactile features; Determining synchronization events from the visual features, the auditory features, the olfactory features, and the tactile features through a time alignment model; A multimodal feedback strategy is generated according to the synchronization event, and the AR glasses are controlled to execute the multimodal feedback strategy.

2. The full sensory data processing method based on AR glasses according to claim 1 is characterized in that: The obtaining of full sensory data for the target environment sent by the multimodal sensor specifically includes: Receive visual modality data sent by RGB+depth camera; receiving auditory modality data sent by a microphone array; Receiving olfactory modality data sent by the odor sensor; receiving tactile modality data sent by the pressure sensor; The visual modality data, the auditory modality data, the olfactory modality data, and the tactile modality data are synchronized, cleaned, and standardized to obtain the full sensory data.

3. The full sensory data processing method based on AR glasses according to claim 1 is characterized in that: The multimodal dynamic fusion network is used to extract features from the full sensory data to obtain visual features, auditory features, olfactory features and tactile features, specifically including: According to the multimodal dynamic fusion network, ResNet+Transformer is used to extract expression and scene feature vectors from the full sensory data; Determining the visual feature based on the expression and scene feature vector; generating an audio frequency spectrum from the full sensory data using a short-time Fourier transform; According to the multimodal dynamic fusion network, extracting a spectrum feature vector from the audio time-frequency graph through CNN; The auditory feature is determined based on the frequency spectrum feature vector.

4. The full sensory data processing method based on AR glasses according to claim 3 is characterized in that: The multimodal dynamic fusion network is used to extract features from the full sensory data to obtain visual features, auditory features, olfactory features and tactile features, and specifically includes: According to the multimodal dynamic fusion network, sparse coding is used to analyze the odor signal pattern vector from the full sensory data; Determining the olfactory feature based on the odor signal pattern vector; According to the multimodal dynamic fusion network, a dynamic change feature vector is obtained by using multi-scale convolution to identify from the full sensory data; The tactile feature is determined based on the dynamically changing feature vector.

5. The full sensory data processing method based on AR glasses according to claim 1 is characterized in that: The determining of synchronization events from the visual features, the auditory features, the olfactory features, and the tactile features by using a time alignment model specifically includes: Calculating the correlation between a first feature and a second feature in the same time window to obtain a correlation coefficient, wherein the first feature is any one of the visual feature, the auditory feature, the olfactory feature, and the tactile feature, and the second feature is any one of the visual feature, the auditory feature, the olfactory feature, and the tactile feature except the first feature; Determine the magnitude relationship between the correlation coefficient and a preset threshold; If the correlation coefficient is greater than or equal to the preset threshold, it is determined that the first feature and the second feature belong to the synchronous event.

6. The full sensory data processing method based on AR glasses according to claim 5 is characterized in that: The correlation between the first feature and the second feature in the same time window is calculated to obtain a correlation coefficient, specifically using the following calculation formula: ; Among them, C ij is the correlation coefficient, i is the first feature, j is the second feature, F i is the first eigenvector, F j is the second eigenvector, Cov(F i ,F j ) is the covariance of the first eigenvector and the second eigenvector, Var(F i ) is the variance of the first eigenvector, Var(F j ) is the variance of the second eigenvector.

7. The full sensory data processing method based on AR glasses according to claim 1 is characterized in that: The generating a multimodal feedback strategy according to the synchronization event, and controlling the AR glasses to execute the multimodal feedback strategy specifically includes: Determine the target category corresponding to the synchronization event according to the preset category, wherein the preset category includes a dangerous event, a notification event, and an interactive event; Based on the target category, the multimodal feedback strategy is generated, and the multimodal feedback strategy is executed by controlling the AR glasses to prompt the wearer of the AR glasses.

8. A full sensory data processing device based on AR glasses, characterized in that: The full sensory data processing device comprises an acquisition module (31) and a processing module (32), wherein: The acquisition module (31) is used to acquire full sensory data for a target environment sent by a multimodal sensor, wherein the multimodal sensor is located on the AR glasses; The processing module (32) is used to extract features from the full sensory data using a multimodal dynamic fusion network to obtain visual features, auditory features, olfactory features and tactile features; The processing module (32) is further used to determine synchronization events from the visual features, the auditory features, the olfactory features and the tactile features through a time alignment model; The processing module (32) is further used to generate a multimodal feedback strategy according to the synchronization event, and control the AR glasses to execute the multimodal feedback strategy.

9. An electronic device, characterized in that: The electronic device comprises a processor (41), a memory (45), a user interface (43) and a network interface (44), wherein the memory (45) is used to store instructions, the user interface (43) and the network interface (44) are both used to communicate with other devices, and the processor (41) is used to execute the instructions stored in the memory (45) so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.

Citation Information

Cited By

  • Intelligent glasses gesture interaction method, device and equipment

    CN120743119A

  • Multi-sensory virtual reality equipment control method and multi-sensory virtual reality mask

    CN121541786A

  • Multi-sensory virtual reality device control method and multi-sensory virtual reality mask

    CN121541786B