Multi-modal data processing method and device based on AR glasses and electronic equipment

By using lightweight convolutional neural networks, deep convolutional neural networks and long and short-term memory networks on AR glasses, and determining processing priorities through scoring mechanisms, the problem of excessive consumption of computing resources for multimodal data processing of AR glasses is solved, and efficient and real-time data processing is achieved.

CN120217052AInactive Publication Date: 2025-06-27GUANGZHOU GUDONG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510333619.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When performing multimodal data processing on AR glasses in the prior art, computing resources consume too much, resulting in response delay and low data processing efficiency.

Method used

Lightweight convolutional neural networks are used to extract visual modal data features, deep convolutional neural networks are used to extract audio modal data features, long and short-term memory networks are used to extract IMU modal data features, and the processing priority of each modal data is determined through the scoring mechanism.

Benefits of technology

It reduces the complexity of model calculation, improves the real-time and efficiency of data processing, avoids information redundancy caused by full data processing, and ensures efficient response of AR glasses in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217052A_ABST
    Figure CN120217052A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data processing method and device based on AR glasses and electronic equipment, and relates to the field of data processing. Obtaining visual modal data, audio modal data and IMU modal data for the target equipment; performing feature extraction on the visual modal data by adopting a lightweight convolutional neural network corresponding to a visual modal encoder to obtain dynamic visual features; performing feature extraction on the audio modal data by adopting a deep convolutional neural network corresponding to an audio modal encoder to obtain audio key features; performing feature extraction on the IMU modal data by adopting a long short-term memory network corresponding to an IMU modal encoder to obtain time sequence features; scoring the dynamic visual features, the audio key features and the time sequence features to obtain a scoring result; and according to a scoring result, determining processing priorities corresponding to the visual modal data, the audio modal data and the IMU modal data by the AR glasses. By implementing the scheme, the data processing efficiency of the AR glasses can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly relates to a multi-modal data processing method, device, and electronic device based on an AR glasses. Background Art

[0002] In the field of industrial manufacturing, AR glasses, as intelligent devices, are widely used in the auxiliary inspection tasks of factories. By integrating multiple sensors such as cameras, microphones, and IMUs, AR glasses can collect multi-modal data such as vision, audio, and motion in real time, providing support for the status monitoring, fault diagnosis, and operation guidance of factory equipment.

[0003] Currently, related multi-modal data processing technologies usually adopt a full-volume processing strategy, that is, uniformly analyze or fuse the collected data of each modality. However, full-volume processing consumes a large amount of computing resources. Especially on AR glasses, this computational overhead may lead to response delays and reduce the overall data processing efficiency.

[0004] Therefore, there is an urgent need for a multi-modal data processing method, device, and electronic device based on AR glasses. Summary of the Invention

[0005] This application provides a multi-modal data processing method, device, and electronic device based on AR glasses, which is convenient for improving the data processing efficiency of AR glasses.

[0006] In the first aspect of this application, a multi-modal data processing method based on AR glasses is provided. The method includes: obtaining a multi-modal data set for a target device, where the multi-modal data set includes visual modality data, audio modality data, and IMU modality data; using a lightweight convolutional neural network corresponding to a visual modality encoder to extract features from the visual modality data to obtain dynamic visual features; using a deep convolutional neural network corresponding to an audio modality encoder to extract features from the audio modality data to obtain audio key features; using a long short-term memory network corresponding to an IMU modality encoder to extract features from the IMU modality data to obtain time series features; scoring the dynamic visual features, the audio key features, and the time series features to obtain a scoring result; and determining the processing priorities corresponding to the visual modality data, the audio modality data, and the IMU modality data by the AR glasses according to the scoring result.

[0007] By adopting the above technical solutions, visual modality, audio modality, and IMU modality data are obtained simultaneously, providing multi-angle information sources for the status monitoring of the target device. The fusion of different modality data can complement each other, improving the comprehensiveness and robustness of detection. The lightweight convolutional neural network is adopted, which is efficient and energy-saving when running on the AR glasses. The deep convolutional network is good at processing spectrogram data and accurately extracting abnormal sound patterns. The long short-term memory network is used to analyze time series features and adapt to the dynamically changing motion data. This design makes full use of the characteristics of each modality data, reduces the computational complexity of the model, and ensures real-time performance. The dynamic visual features, audio key features, and time series features are extracted respectively, focusing on the data patterns related to the target task and avoiding information redundancy generated when processing all data. The feature extraction stage takes the modality characteristics as the core, improving the distinguishability and effectiveness of the features and providing high-quality data input for subsequent scoring and priority assignment. Score according to the feature importance, assign dynamic priorities to each modality, and ensure that the most relevant modality is processed first in the current task. Adjust the modality processing priorities according to the scoring results, avoid full-scale analysis of all modalities, reduce redundant data processing, and improve real-time performance and computational efficiency. The system can flexibly switch the dominant modality and dynamically adjust the priority assignment according to environmental changes to ensure that important data will not be delayed due to insufficient computing resources. Therefore, it is convenient to improve the data processing efficiency of the AR glasses.

[0008] Optionally, the lightweight convolutional neural network corresponding to the visual modality encoder is adopted to extract features from the visual modality data to obtain dynamic visual features, which specifically includes: determining the visual image corresponding to the visual modality data according to the visual modality data; adopting the lightweight convolutional neural network to perform image analysis on the visual image to obtain the changing area in the visual image, so as to obtain the dynamic visual features.

[0009] By adopting the above technical solution, the visual image is analyzed by a lightweight convolutional neural network, and only the change regions related to the task are extracted, avoiding processing the static or irrelevant regions in the entire visual image. This improves the task adaptability of visual features, reduces the interference of irrelevant information on subsequent processing, and enhances the accuracy and efficiency of detection. By using the lightweight convolutional neural network, while maintaining a high detection accuracy, the computational overhead is significantly reduced, making it suitable for AR glasses to operate. Avoiding full-scale processing of the entire image and focusing the limited computational resources on the dynamic regions improves the real-time performance and response speed of the system. Through the extraction of dynamic visual features, the system can timely sense the state changes of the target device. In complex scenarios, focusing on the changing regions can more effectively identify key information and enhance the robustness of the system. Through the detection of the changing regions of the visual image, the feature focus of the visual modality is dynamically adjusted, enabling the system to flexibly process different types of visual information according to the scene and task requirements, providing high-quality visual feature inputs for subsequent modality priority scoring and multimodal fusion.

[0010] Optionally, a deep convolutional neural network corresponding to the audio modality encoder is used to extract features from the audio modality data to obtain audio key features, which specifically includes: determining a spectrogram corresponding to the audio modality data according to the audio modality data; using the deep convolutional neural network to determine high-frequency noise and / or abnormal audio in the spectrogram to obtain the audio key features.

[0011] By adopting the above technical solution, by analyzing the high-frequency noise and abnormal audio in the spectrogram, key information related to the device operation state or faults can be effectively captured. The interference of background noise or irrelevant audio is reduced, ensuring that the extracted audio features are highly relevant to the task and improving the detection accuracy. The deep convolutional neural network demonstrates excellent feature extraction capabilities when processing spectrograms, capable of mining deep patterns and abnormal characteristics in audio signals. Automatically analyzing the spectrogram without the need for manually designing complex audio feature extraction rules, it has strong adaptability and is easy to deploy. By filtering out irrelevant low-frequency signals or background noise and only focusing on the high-frequency noise and abnormal audio in the spectrogram, the redundant information in the audio data is significantly reduced. This reduces the computational burden on the system and simultaneously accelerates the efficiency of subsequent processing steps. In complex environments such as factories, device operation is usually accompanied by various background noises. Through spectrogram analysis and high-frequency feature extraction, the system can effectively distinguish abnormal device audio from environmental noise. This enhances the task adaptation ability and scene robustness of the audio modality data.

[0012] Optionally, use the long short-term memory network corresponding to the IMU modality encoder to extract features from the IMU modality data to obtain time series features, which specifically includes: calculating the movement trajectory of the user's head direction, determining the user's attention area of the target device, where the user wears the AR glasses; through the long short-term memory network, when the user faces the user's attention area, detecting fine movement mutations to obtain the time series features.

[0013] By adopting the above technical solution, calculate the user's head direction and movement trajectory through IMU data, accurately locate the attention area when the user wears AR glasses, and ensure that the system can focus on processing target device-related information according to the user's intention. Effectively improve the interaction experience between the AR glasses and the user, and avoid wasting resources on the analysis of non-attention areas. The LSTM network is good at processing time series data. By analyzing the fine mutations in the head movement trajectory, it can capture key actions or changes related to the task. Enhance the system's perception ability of dynamic behaviors, and is applicable to scenarios where fine changes may be key information. Combining IMU modality data and time series feature extraction, the system can adjust the processing priority of the target device in real time when the user moves or turns dynamically, realizing seamless docking with the user's behavior. Improve the adaptability and user experience of the AR glasses, especially in scenarios such as quick inspections or multi-device inspections, and avoid information omission caused by lagging processing. Using the head movement trajectory and attention area information, the system can preferentially process the data related to the target device that the user is currently interested in, and avoid full-scale analysis of the data in the entire area. Reduce the computational burden of the AR glasses, and at the same time accelerate the overall efficiency of multi-modal data processing. The time series features extracted by IMU modality can work in coordination with visual modality and audio modality, providing supplements in the spatio-temporal dimension for multi-modal data fusion and system decision-making, and improving the overall detection accuracy and task adaptation ability.

[0014] Optionally, score the dynamic visual features, the audio key features, and the time series features to obtain a scoring result, which specifically includes: if it is determined that the dynamic visual features indicate that the red light of the target device is flashing, generate a first score; if it is determined that the audio key features indicate that the target device emits a sharp noise, generate a second score; if it is determined that the time series features indicate that the user turns to the target device, generate a third score, where the first score is higher than the third score, the second score is higher than the third score, and the score difference between the first score and the second score is within a preset range; based on the first score, the second score, and the third score, determine the scoring result.

[0015] By adopting the above technical solutions, the dynamic visual features, audio key features, and time series features are combined. The importance of each modality is quantified through a scoring mechanism to form a comprehensive evaluation result, ensuring the scientificity and accuracy of decision-making. The strong indicative features such as red light flashing and sharp noise are comprehensively considered in combination with the time series features of user behavior to avoid misjudgment that may be caused by single-modal data. Through scoring settings, the strong indicative features are highlighted and given a higher scoring priority to ensure that the system can quickly focus on potential anomalies or emergency situations. Dynamic adjustment of priorities is achieved in multi-modal data processing to optimize the resource utilization efficiency. In the scoring mechanism, the priorities of red light flashing and sharp noise are set to be close and the scoring difference is within a preset range to ensure balanced analysis of different modal features and adaptation to diverse scenario requirements. Deviations caused by too high a score for a single modality are avoided, improving the system's adaptability and decision-making reliability in complex environments. The time series features of the user are incorporated into the scoring system to enable the system to dynamically respond to changes in the user's focus. The intelligent level of the interaction between the AR glasses and the user is enhanced to ensure that the target device information most relevant to the user is preferentially processed in a multi-target environment. Through the weighted combination of the scoring mechanism, abnormal devices or task targets to be preferentially processed can be quickly locked, improving the system's response speed and accuracy in multi-target scenarios. The scoring mechanism processes each modal feature through differential weighting, avoiding redundant processing of unimportant or secondary features, concentrating resources on analyzing key data, and improving the efficiency of multi-modal data processing.

[0016] Optionally, determining the corresponding processing priorities of the AR glasses for the visual modal data, the audio modal data, and the IMU modal data according to the scoring result specifically includes: determining a first processing priority according to the first score; determining a second processing priority according to the second score; determining a third processing priority according to the third score, the first processing priority is higher than the third processing priority, the second processing priority is higher than the third processing priority, and the first processing priority is the same as the second processing priority.

[0017] By adopting the above technical solutions, the processing priority is assigned based on the scoring results, ensuring that the most critical modal data is processed first to quickly respond to potential emergencies or abnormal conditions. This avoids unnecessary over-analysis of low-priority data by the system and improves the task processing efficiency. By setting the first processing priority to be the same as the second processing priority, the equivalent importance of visual and audio modalities in anomaly detection is highlighted, meeting the diverse scene requirements. The processing order of each modality is dynamically adjusted to ensure that the system can flexibly respond to changes in various signal combinations in the scene. The lowest priority is assigned to the low-priority time series features to avoid wasting resources on data processing that has a weak correlation with the current task. The computing resources are concentrated to process high-priority modal data, reducing the computing and power consumption burdens of the AR glasses and improving the overall system efficiency. The processing order of different modal data is adjusted based on the priority, providing a structured processing logic for multi-modal fusion. High-priority modal data can participate in decision-making faster, complementing low-priority data to achieve more accurate comprehensive judgment. Through the setting of dynamic priorities, the system can prioritize the most important needs or the most urgent abnormal conditions of the user during real-time processing, ensuring an efficient human-computer interaction experience.

[0018] Optionally, the method further includes: receiving the personalized processing priority sent by the user device for the multi-modal data set; processing the visual modal data, the audio modal data, and the IMU modal data according to the personalized processing priority.

[0019] By adopting the above technical solutions, by adjusting the processing order of visual, audio, and IMU modal data according to the user's personalized processing priority, it is ensured that the system can better adapt to the specific needs of each user. It meets the different needs of different users for information processing priorities in different scenarios, enhancing the personalization and flexibility of human-computer interaction. The data processing priority is dynamically adjusted according to the user's real-time needs, enabling the system to quickly respond to the most urgent or critical data and improving the system's response speed to emergency tasks or important information. It is particularly suitable for rapidly changing industrial environments or scenarios with multi-task operations, ensuring that users can quickly obtain the most relevant feedback. The introduction of personalized priorities enables the system to more intelligently allocate computing resources, avoiding excessive resource occupation by low-priority tasks and reducing the system burden. It realizes efficient data processing, saves battery life and computing resources, and meets the real-time processing requirements of AR glasses. The system can flexibly adjust the data processing strategy according to the user's specific preferences or scene requirements, enhancing the system's adaptability to changing environments. Users can adjust the priorities of each modal data according to different task requirements, enabling the system to operate efficiently in different application scenarios.

[0020] In a second aspect of the present application, a multimodal data processing device based on an AR glasses is provided. The multimodal data processing device includes an acquisition module and a processing module. Among them, the acquisition module is used to acquire a multimodal data set for a target device, and the multimodal data set includes visual modal data, audio modal data, and IMU modal data; the processing module is used to extract features from the visual modal data by using a lightweight convolutional neural network corresponding to a visual modal encoder to obtain dynamic visual features; the processing module is further used to extract features from the audio modal data by using a deep convolutional neural network corresponding to an audio modal encoder to obtain audio key features; the processing module is further used to extract features from the IMU modal data by using a long short-term memory network corresponding to an IMU modal encoder to obtain time series features; the processing module is further used to score the dynamic visual features, the audio key features, and the time series features to obtain a scoring result; the processing module is further used to determine the processing priorities of the AR glasses for the visual modal data, the audio modal data, and the IMU modal data respectively according to the scoring result.

[0021] In a third aspect of the present application, an electronic device is provided. The electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method described above.

[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described above is executed.

[0023] In summary, one or more technical solutions provided in the present application have at least the following technical effects or advantages: By simultaneously acquiring visual modality, audio modality, and IMU modality data, it provides multi - perspective information sources for the status monitoring of the target device. The fusion of different modality data can complement each other, improving the comprehensiveness and robustness of detection. An efficient and energy - saving lightweight convolutional neural network is adopted to run on the AR glasses. The deep convolutional network is good at processing spectrogram data and accurately extracting abnormal sound patterns. The long short - term memory network is used to analyze time - series features and adapt to dynamically changing motion data. This design makes full use of the characteristics of each modality data, reduces the computational complexity of the model, and ensures real - time performance. Dynamically visual features, audio key features, and time - series features are extracted respectively, focusing on the data patterns related to the target task and avoiding information redundancy generated when processing all data. The feature extraction stage takes the modality characteristics as the core, improving the distinguishability and effectiveness of the features, and providing high - quality data input for subsequent scoring and priority assignment. Score according to the feature importance, assign dynamic priorities to each modality, and ensure that the most relevant modality is preferentially processed in the current task. Adjust the modality processing priority according to the scoring results, avoid full - scale analysis of all modalities, reduce redundant data processing, and improve real - time performance and computational efficiency. The system can flexibly switch the dominant modality and dynamically adjust the priority assignment according to environmental changes, ensuring that important data will not be delayed due to insufficient computing resources. Therefore, it is convenient to improve the data processing efficiency of the AR glasses. Description of the Drawings

[0024] Figure 1 It is a schematic flowchart of a multi - modality data processing method based on an AR glasses provided by an embodiment of the present application; Figure 2 It is another schematic flowchart of a multi - modality data processing method based on an AR glasses provided by an embodiment of the present application; Figure 3 It is a schematic diagram of modules of a multi - modality data processing device based on an AR glasses provided by an embodiment of the present application; Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application.

[0025] Description of the reference numerals: 31, acquisition module; 32, processing module; 41, processor; 42, communication bus; 43, user interface; 44, network interface; 45, memory. Detailed Embodiments

[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0027] In the description of the embodiments of the present application, words such as "for example" or "for illustration" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for illustration" is intended to present the relevant concepts in a specific manner.

[0028] In the description of the embodiments of the present application, the term "a plurality of" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0029] In the field of industrial manufacturing, as an intelligent device, AR glasses have been widely used in the auxiliary inspection work of factories. By integrating sensors such as cameras, microphones, and IMUs, AR glasses can collect multi-modal data such as vision, audio, and motion in real time, thereby providing support for equipment status monitoring, fault diagnosis, and operation guidance.

[0030] Currently, multi-modal data processing technologies usually adopt a full-volume processing method, that is, unified analysis or fusion of the collected data of each modality. However, this full-volume processing method requires a large amount of computing resources. Especially on devices such as AR glasses, an excessive computing burden may lead to response delays, thereby reducing the overall data processing efficiency.

[0031] To solve the above technical problems, the present application provides a multi-modal data processing method based on AR glasses. Refer to Figure 1 , Figure 1 which is a schematic flow chart of a multi-modal data processing method based on AR glasses provided by the embodiments of the present application. This multi-modal data processing method is applied to the controller of AR glasses and includes steps S110 to S160. The above steps are as follows: S110. Obtain a multi-modal data set for a target device, where the multi-modal data set includes visual modal data, audio modal data, and IMU modal data.

[0032] Specifically, the controller of the AR glasses is the core control unit in the AR glasses, responsible for processing and coordinating data from different sensors. These sensors include visual sensors such as cameras, audio sensors such as microphones, and motion sensors such as IMUs. The controller collects environmental information through these sensors and converts it into a multi-modal dataset. A multi-modal dataset refers to a collection of data from different types of sensors. Visual modal data is the image or video data captured by the camera, providing visual information about the environment, target objects, or device status. Audio modal data is the audio data captured by the microphone, such as the working sound of a machine, an alarm sound, or abnormal noise of a device failure. IMU modal data is the motion data captured by the IMU, usually containing the outputs of the accelerometer and gyroscope, reflecting the movement, position change, and rotation of the device or user, etc.

[0033] For example, assume that on a factory production line, AR glasses are used to assist workers in checking the operating status of equipment. The controller will obtain data from multiple sensors. The camera of the AR glasses captures a certain piece of equipment on the production line, and displays the indicator lights, fault warning signs, or the operating status of the machine in real time. For example, when a device fails, the camera recognizes the red flashing light on the device. The microphone of the AR glasses may capture abnormal noise from the device. For example, if the device emits abnormal high-frequency sharp noise, it may indicate a device fault. The IMU sensor of the AR glasses monitors the movement of the wearer's head or the glasses. If the user turns their head to view a certain piece of equipment, the IMU data can record the head rotation trajectory, thus helping to determine the equipment that the user is concerned about. These different modal data will be aggregated into a multi-modal dataset and processed by the controller of the AR glasses, comprehensively analyzing this data to make decisions or provide real-time feedback. In this way, the AR glasses can not only provide visual information in real time, but also help users more comprehensively understand the device status through audio and motion data, thereby providing more accurate fault diagnosis and operation guidance.

[0034] S120. Use a lightweight convolutional neural network corresponding to the visual modal encoder to extract features from the visual modal data to obtain dynamic visual features.

[0035] Specifically, the visual modality encoder refers to the system component in the AR glasses responsible for processing and analyzing visual data. In the embodiments of this application, the visual modality data comes from the camera of the AR glasses, and the captured image or video information is input into the visual modality encoder for analysis. The convolutional neural network is a deep learning model widely used for processing image and video data. In the convolutional neural network, the convolutional layer can automatically extract features from the image, such as edges, shapes, textures, etc. The lightweight convolutional neural network is an optimization aimed at reducing computational complexity and storage requirements and is suitable for AR glasses. These lightweight models improve the running efficiency by reducing the number of parameters and the number of network layers while maintaining good performance. Feature extraction is the process of extracting representative information or features from the original data. In the scenario of image processing, features may include information such as the outline of an object, color changes, and motion. These features can help the controller understand the image or video content for further analysis or decision-making. Dynamic visual features refer to the changing parts in an image or video. For example, in industrial inspection, dynamic visual features may include changes in equipment status, movement of machine parts, or flashing indicator lights. These features can reflect the real-time status and behavior of the target device and help identify potential faults or abnormal conditions.

[0036] In a possible implementation manner, a lightweight convolutional neural network corresponding to the visual modality encoder is used to extract features from the visual modality data to obtain dynamic visual features, specifically including: determining the visual image corresponding to the visual modality data according to the visual modality data; using the lightweight convolutional neural network to perform image analysis on the visual image to obtain the changing area in the visual image, so as to obtain dynamic visual features.

[0037] Specifically, the visual image is a specific image or video frame extracted from the visual modality data. It may be a static image or a frame in a video, representing the visual scene of a certain moment of the device, environment, or machine. Image analysis is to process the input visual image to identify the patterns, features, or key areas therein. In the embodiments of this application, image analysis refers to using the lightweight convolutional neural network to identify the changing parts in the image, and these changes represent the dynamic visual features related to the target task. The changing area refers to the part of the image that has changed. These changes can be the movement of an object, color changes, flashing of device indicator lights, etc. In industrial inspection, the changing area may indicate changes in equipment status, such as equipment failure or changes in working status. Dynamic visual features refer to the changing parts in an image, and these changes may indicate changes or abnormalities in the working status of the device. For example, the flashing of a red indicator light, the movement of machine parts, etc. can all be used as dynamic visual features to judge the running status of the device.

[0038] S130. Use the deep convolutional neural network corresponding to the audio modality encoder to extract features from the audio modality data to obtain the audio key features.

[0039] Specifically, the audio modality encoder is a system component built into the AR glasses, dedicated to processing and analyzing audio data. In the application scenario of AR glasses, the audio modality data is usually collected through the built-in microphone. This data may include the machine sounds in the factory environment, the operator's speech, the alarm sounds of the equipment, etc. Although the convolutional neural network was initially mainly used for image analysis, it can also be applied to audio analysis. Especially after converting the audio signal into a spectrogram, the convolutional neural network can automatically extract important features from the audio data. The audio modality data is the sound data collected by the microphone. This data is usually the original audio waveform, but for deep learning processing, it may be first converted into a spectrogram, such as Mel-frequency cepstral coefficients, for further analysis by the model. The audio key features are the valuable information or patterns extracted from the audio data, and these features may reveal important changes or anomalies in the audio signal. For example, the sharp noise emitted by the equipment, the difference between the normal operation sound and the fault sound of the machine, or the change in background noise, etc.

[0040] In a possible implementation, using the deep convolutional neural network corresponding to the audio modality encoder to extract features from the audio modality data to obtain the audio key features specifically includes: determining the spectrogram corresponding to the audio modality data according to the audio modality data; using the deep convolutional neural network to determine the high-frequency noise and / or abnormal audio in the spectrogram to obtain the audio key features.

[0041] Specifically, audio modal data usually exists in the form of a time series. For the convenience of analysis and processing, audio data is often converted into a spectrogram. A spectrogram is a graph obtained by performing a Fourier transform on an audio signal and can represent the energy distribution of the audio signal at different frequencies. Common methods for representing spectrograms include Mel-frequency cepstral coefficients and short-time Fourier transform, etc. A deep convolutional neural network is a neural network architecture used for processing images and audio. It extracts local features from the input data through convolutional layers and reduces the complexity of the data through pooling layers. In audio analysis, it can automatically extract useful features from the spectrogram without the need to manually design a feature extraction algorithm. High-frequency noise generally refers to noise with a relatively high frequency in an audio signal, which may be generated by equipment failures, mechanical vibrations, or other abnormal conditions. Such noise may affect the normal operation of the equipment and needs to be detected in a timely manner. Abnormal audio refers to sounds that are different from the audio characteristics under normal operating conditions, such as sharp noises, metal friction sounds, alarm sounds, etc. emitted by the equipment. Audio key features are important information extracted from the spectrogram, and these features can reveal significant changes or abnormalities in the audio signal. For example, an abnormal increase in the energy of the high-frequency part of the spectrogram may indicate that the equipment has emitted a fault warning sound or generated abnormal noise.

[0042] S140. Use the long short-term memory network corresponding to the IMU modal encoder to extract features from the IMU modal data to obtain time series features.

[0043] Specifically, an IMU is a device used to measure the acceleration, angular velocity, and magnetic field of an object and consists of an accelerometer, a gyroscope, and a magnetometer. It can capture multi-dimensional data of motion, such as information about the acceleration, rotation angle, and direction of an object. In the application of AR glasses, the IMU modal data is used to capture the motion of the wearer, such as head rotation, eye gaze direction, etc. The IMU modal encoder is a component in AR glasses specifically used to process IMU data. It preprocesses the original data of the IMU and passes it to the subsequent neural network model. In this process, the role of the encoder is to convert the original motion data into features suitable for further analysis and learning. The long short-term memory network is a special recurrent neural network that can maintain a long-term memory when processing time series data and can overcome the common problem of gradient disappearance in long sequences.

[0044] In a possible implementation, use the long short-term memory network corresponding to the IMU modal encoder to extract features from the IMU modal data to obtain time series features, specifically including: calculating the motion trajectory of the user's head direction, determining the user's attention area of the target device, where the user wears AR glasses; through the long short-term memory network, when the user faces the user's attention area, detecting fine motion mutations to obtain time series features.

[0045] Specifically, through IMU data, LSTM can calculate the movement trajectory of the user's head. That is, the change in the rotation direction and angle of the user's head over time. Through the head movement trajectory, the server can infer the area that the user is looking at or paying attention to, such as the target device. The movement trajectory of the user's head not only helps to understand the user's head direction, but also can speculate the area of interest of the user through the head orientation. For example, if the user's head is facing a certain device, it can be considered that the user is paying attention to that device. By analyzing time series data, LSTM can detect subtle changes in the user's head movement. For example, if the user suddenly turns to a certain device, or makes a small head adjustment within the area of interest, LSTM can identify these subtle changes and extract them as time series features. This is very important for real-time monitoring of user behavior. Finally, the time series features extracted by LSTM can be used to reflect the head movement pattern and attention state of the user. These features will help the server understand the user's current behavior and thus make corresponding responses.

[0046] For example, the IMU sensor of the AR glasses collects data on the wearer's head movement, which includes the acceleration of the head in the X, Y, and Z axes, as well as the angular velocity, etc. For example, when the wearer turns their head, the gyroscope records the angle and speed of the wearer's head rotation. After receiving this data, the LSTM calculates the movement trajectory of the wearer's head by analyzing the time series. For example, when the wearer starts to turn their head from one area to another, the LSTM can capture this change and infer the movement path of the wearer's head in space. Suppose the wearer starts to look at a certain device, the LSTM will determine the wearer's area of interest by analyzing the head movement trajectory. For example, if the wearer turns their head towards device A, the server can determine that device A is the current area of interest. Suppose the wearer is continuously looking at device A and suddenly the wearer fine-tunes the angle of their head. Such a subtle movement change may indicate that the wearer is more focused on a certain detail of device A. The LSTM can detect these subtle mutations and generate corresponding time series features. For example, for the action of the wearer slightly adjusting the head angle, the LSTM can recognize this change and then infer whether the wearer's attention is concentrated. Finally, the LSTM extracts time series features based on the movement trajectory and subtle changes of the user's head. The server knows that the wearer is interested in device A rather than device B. For example, the wearer may stay in front of device A for a long time, and the server can recognize this pattern. The wearer's head may have a slight adjustment in a certain direction, and the server can incorporate these details into the analysis to better understand the wearer's degree of attention. Once the LSTM extracts these time series features, the server can make a feedback according to the wearer's attention state. For example, the server can automatically display the status or operation guide of the device when the wearer is looking at a certain device, or provide more detailed help information when the wearer fine-tunes the head angle.

[0047] S150. Score the dynamic visual features, audio key features, and time series features to obtain a scoring result.

[0048] Specifically, after obtaining these features, the controller will score these features. The purpose of scoring is to rank the importance of different modality data, so as to determine which data needs to be processed first. By scoring the visual, audio, and time series features, the controller of the AR glasses can determine the priority of data processing according to these scoring results. The advantage of doing this is to improve the processing efficiency, avoid comprehensive processing of all data, thereby reducing the consumption of computing resources, and being able to flexibly adjust the response strategy according to the real-time situation. This method enables the AR glasses to make more accurate and timely responses in complex environments, improving the efficiency of device monitoring and auxiliary inspection.

[0049] In a possible implementation, the dynamic visual features, audio key features, and time series features are scored to obtain a scoring result, which specifically includes: if it is determined that the dynamic visual features indicate that the red light of the target device is flashing, a first score is generated; if it is determined that the audio key features indicate that the target device emits a sharp noise, a second score is generated; if it is determined that the time series features indicate that the user turns to the target device, a third score is generated, the first score is higher than the third score, the second score is higher than the third score, and the score difference between the first score and the second score is within a preset range; based on the first score, the second score, and the third score, the scoring result is determined.

[0050] Specifically, when the red light of the target device flashes in the visual modality data, this change is regarded as a possible device failure or warning. This situation generates a first score. Since the red light flashing usually indicates an emergency, this score may be relatively high, indicating a high level of urgency for this visual feature and the need for priority processing. If a sharp noise is detected in the audio modality data, for example, a mechanical device suddenly emits an alarm sound or a warning tone, this may be a signal of a device failure and generates a score. This situation generates a second score. The sharp noise also means a problem that requires immediate attention, so this score may also be relatively high, indicating a high level of urgency for this audio feature. If the user turns to the target device while wearing AR glasses, the time series features in the IMU data can indicate the change in the wearer's attention area or direction. This usually means that the user is paying attention to a certain device or operation. This situation generates a third score. Compared with visual or audio data, the user's turning action is usually not an emergency failure signal but more like the user's attention or operation behavior towards the device. Therefore, this score may be relatively low, indicating a low level of urgency. The first score and the second score represent the features of the visual and audio modalities respectively, and they both reflect possible failures or emergencies, so their scores will be relatively high. The third score represents the user's behavior, indicating that the wearer is concentrating, although it is important, it usually does not involve emergency issues. Therefore, its score is relatively low. The difference between the first score and the second score must be within a preset range, indicating that the first score and the second score have an equal possibility. This range is used to ensure that the level of urgency between the two is comparable. If the difference in the level of urgency between the red light flashing and the sharp noise is too large, it may be necessary to readjust the scoring weights. The controller determines the overall scoring result based on all the scoring results. For example, if the first score and the second score are both relatively high while the third score is relatively low, the controller will give priority to processing the visual and audio modality data to ensure that device failures or warnings are responded to in a timely manner. The time series features can be processed later because they do not involve immediate emergencies.

[0051] For example, assume that the user is performing equipment monitoring in a factory through AR glasses and has collected the following data: The red light of Device A is flashing. The AR glasses extract this change through a visual encoder, and the dynamic visual features are scored. Assume that the server assigns a relatively high score to this emergency change, which is called the first score. At the same time, Device B emits a sharp noise, and the AR glasses recognize this abnormal audio through an audio encoder. Since a sharp noise usually indicates that the device may be malfunctioning, the key audio features are given a relatively high score, which is called the second score. At this time, the wearer turns their line of sight to Device A, and the IMU modal data indicates that the wearer's head has turned towards Device A. Although this indicates that the wearer's attention is focused on this device, compared with a device failure, its urgency is lower. Therefore, the time series feature score corresponding to this behavior is relatively low, which is called the third score. The first score is obtained from the flashing red light, which indicates that the device may be malfunctioning and the score is relatively high. The second score is obtained from the sharp noise, which also indicates a device problem and the score is relatively high. The third score is obtained from the wearer's head turning, and the score is relatively low. Assume that the score gap between the first score and the second score is within a preset reasonable range. Then, the controller of the AR glasses will preferentially process the visual and audio data with higher first and second scores to handle the emergency problems of the device. The time series data with a lower third score will be processed later. Through this scoring method, the server can distinguish the urgency and importance of different modal data, ensuring that the AR glasses can efficiently respond to potential problems in a complex environment. This priority processing based on scoring can avoid unnecessary waste of computing resources, while ensuring that the most important data can be preferentially processed at critical moments, improving the overall processing efficiency and user experience.

[0052] S160. Determine the processing priorities corresponding to the visual modal data, audio modal data, and IMU modal data of the AR glasses according to the scoring results.

[0053] Specifically, the previous steps have scored the visual, audio, and IMU data. Each modal data has received different scores according to the situation it reflects, such as a fault indication, abnormal noise, or user behavior. The higher the score, the more urgently this data modal needs to be processed. Based on the scoring results, the controller assigns a processing priority to the data of each modal. The priority determines which data should be processed first and which can be processed later. The purpose of this process is to ensure that important and urgent data is preferentially processed, rather than all data being treated equally.

[0054] In a possible implementation manner, according to the scoring results, the processing priorities corresponding to the visual modality data, audio modality data, and IMU modality data of the AR glasses are determined, specifically including: determining a first processing priority according to the first score; determining a second processing priority according to the second score; determining a third processing priority according to the third score, where the first processing priority is higher than the third processing priority, the second processing priority is higher than the third processing priority, and the first processing priority is the same as the second processing priority.

[0055] Specifically, since visual data, such as a red light flashing, may indicate a device failure or the need for immediate attention, its score is the highest and the processing priority is the first. Audio data, such as a sharp noise, may also indicate a device anomaly or danger. Although its score is lower than that of visual data, it is still urgent, so its priority is the second. IMU data is usually only used to analyze the user's behavior or intention, and its priority is lower in case of an emergency, so its priority is the third. The first processing priority is higher than the third processing priority, and the second processing priority is higher than the third processing priority. The scoring gap between the first and second priorities is within a preset range, so they can be regarded as priority processing tasks of the same level.

[0056] For example, assume that in a factory environment, the controller of the AR glasses is processing data from different modalities: The red light of device A is flashing, indicating a device failure, so the score of the visual modality data is 90, and it is assigned to the first processing priority. Device B emits a sharp alarm sound, indicating a possible safety issue. The score of the audio modality data is 80, so it is assigned to the second processing priority. The user turns their head and faces device A, and the IMU data indicates a change in the user's head direction. The score of the IMU data is 50, so it is assigned to the third processing priority. Therefore, the first priority: visual data, with a score of 90, is the most urgent and needs to be processed first; the second priority: audio data, with a score of 80, is the second most urgent and is processed second; the third priority: IMU data, with a score of 50, is relatively less urgent and is processed last. Through this scoring mechanism and priority arrangement, the AR glasses can effectively manage the processing order of different modality data. In case of an emergency, visual and audio data will be processed first to ensure that device failures and safety issues are responded to in a timely manner. The IMU data is relatively less urgent and can be processed later. The advantage of this is to avoid waste of computing resources that may be caused by processing all data simultaneously, and improve the system response speed and processing efficiency.

[0057] In a possible implementation manner, with reference to Figure 2 , Figure 2Another flowchart of a multimodal data processing method based on an AR glasses provided by an embodiment of the present application. It includes steps S210 to S220, and the above steps are as follows: S210, receiving the personalized processing priority for the multimodal data set sent by the user device; S220, processing the visual modal data, audio modal data, and IMU modal data according to the personalized processing priority.

[0058] Specifically, the personalized processing priority means that each user can set the processing priority order of different types of data according to their own needs or working environment. For example, in some tasks, visual data may be more urgent than audio data, or vice versa. The user device will send this priority information to the controller of the AR glasses to ensure that the data is processed according to the priority set by the user. Among them, the user device is a mobile device used by the user that can perform data interaction with the AR glasses, such as a mobile phone. According to the personalized priority sent by the user device, the AR glasses will process these data differently. For example, if the user sets the audio data as the highest priority, the AR glasses will first process the audio data to analyze whether there is a sharp noise or abnormal sound. Through the personalized priority, the user can control which data needs to be processed with higher priority and which data can be postponed in different task scenarios. For example, when performing equipment fault inspection, the visual signal of the equipment may be more important, while in security monitoring, the audio signal may require more attention.

[0059] The present application also provides a multimodal data processing device based on an AR glasses, referring to Figure 3 , Figure 3 It is a module schematic diagram of a multimodal data processing device based on an AR glasses provided by an embodiment of the present application. The multimodal data processing device is the controller of the AR glasses. The controller includes an acquisition module 31 and a processing module 32. Among them, the acquisition module 31 acquires the multimodal data set for the target device. The multimodal data set includes visual modal data, audio modal data, and IMU modal data; the processing module 32 uses a lightweight convolutional neural network corresponding to the visual modal encoder to extract features from the visual modal data to obtain dynamic visual features; the processing module 32 uses a deep convolutional neural network corresponding to the audio modal encoder to extract features from the audio modal data to obtain audio key features; the processing module 32 uses a long short-term memory network corresponding to the IMU modal encoder to extract features from the IMU modal data to obtain time series features; the processing module 32 scores the dynamic visual features, audio key features, and time series features to obtain a scoring result; the processing module 32 determines the processing priorities corresponding to the visual modal data, audio modal data, and IMU modal data of the AR glasses according to the scoring result.

[0060] In a possible implementation, the processing module 32 uses a lightweight convolutional neural network corresponding to the visual modality encoder to extract features from the visual modality data to obtain dynamic visual features, which specifically includes: the processing module 32 determines a visual image corresponding to the visual modality data according to the visual modality data; the processing module 32 uses the lightweight convolutional neural network to perform image analysis on the visual image to obtain a changed area in the visual image, so as to obtain dynamic visual features.

[0061] In a possible implementation, the processing module 32 uses a deep convolutional neural network corresponding to the audio modality encoder to extract features from the audio modality data to obtain audio key features, which specifically includes: the processing module 32 determines a spectrogram corresponding to the audio modality data according to the audio modality data; the processing module 32 uses the deep convolutional neural network to determine high-frequency noise and / or abnormal audio in the spectrogram, so as to obtain audio key features.

[0062] In a possible implementation, the processing module 32 uses a long short-term memory network corresponding to the IMU modality encoder to extract features from the IMU modality data to obtain time series features, which specifically includes: the processing module 32 calculates the movement trajectory of the user's head direction to determine the user's attention area of the target device, and the user wears AR glasses; the processing module 32 uses the long short-term memory network to detect a subtle movement mutation when the user faces the user's attention area, so as to obtain time series features.

[0063] In a possible implementation, the processing module 32 scores the dynamic visual features, audio key features, and time series features to obtain a scoring result, which specifically includes: if the processing module 32 determines that the dynamic visual features indicate that the red light of the target device is flashing, a first score is generated; if the processing module 32 determines that the audio key features indicate that the target device emits a sharp noise, a second score is generated; if the processing module 32 determines that the time series features indicate that the user turns to the target device, a third score is generated, the first score is higher than the third score, the second score is higher than the third score, and the score difference between the first score and the second score is within a preset range; the processing module 32 determines the scoring result based on the first score, the second score, and the third score.

[0064] In a possible implementation, the processing module 32 determines the processing priorities corresponding to the visual modality data, audio modality data, and IMU modality data according to the scoring result, which specifically includes: the processing module 32 determines a first processing priority according to the first score; the processing module 32 determines a second processing priority according to the second score; the processing module 32 determines a third processing priority according to the third score, the first processing priority is higher than the third processing priority, the second processing priority is higher than the third processing priority, and the first processing priority is the same as the second processing priority.

[0065] In a possible implementation, an acquisition module 31 receives the personalized processing priority for the multimodal data set sent by the user device; a processing module 32 processes the visual modal data, audio modal data, and IMU modal data according to the personalized processing priority.

[0066] It should be noted that when the device provided in the above embodiment realizes its functions, only the division of the above function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process can be found in the method embodiment, which will not be repeated here.

[0067] This application also provides an electronic device. Refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of this application. The electronic device may include: at least one processor 41, at least one network interface 44, a user interface 43, a memory 45, and at least one communication bus 42.

[0068] Among them, the communication bus 42 is used to realize the connection and communication between these components.

[0069] Among them, the user interface 43 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 43 may further include a standard wired interface and a wireless interface.

[0070] Among them, the network interface 44 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0071] Among them, the processor 41 may include one or more processing cores. The processor 41 connects various parts within the entire server through various interfaces and lines, and executes various functions of the server and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 45, and by calling the data stored in the memory 45. Optionally, the processor 41 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 41 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 41 and may be implemented separately by a single chip.

[0072] Among them, the memory 45 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 45 includes a non-transitory computer-readable storage medium. The memory 45 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 45 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 45 may further be at least one storage device located far from the aforementioned processor 41. As Figure 4 shown, the memory 45, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a multi-modal data processing method based on AR glasses.

[0073] In Figure 4In the electronic device shown, the user interface 43 is mainly used to provide an interface for the user to input and obtain the data input by the user; while the processor 41 can be used to call an application program stored in the memory 45 for a multi-modal data processing method based on AR glasses. When executed by one or more processors, the electronic device is caused to execute the method as described in one or more of the above embodiments.

[0074] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0075] The present application also provides a computer-readable storage medium storing instructions. When executed by one or more processors, the electronic device is caused to execute the method as described in one or more of the above embodiments.

[0076] In the above embodiments, the descriptions of the various embodiments have their respective focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0077] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0078] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0079] In addition, in each embodiment of the present application, the various functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0080] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.

[0081] The foregoing are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and the disclosure of the practical truth, those skilled in the art will readily think of other implementation manners of the present disclosure. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A multimodal data processing method based on AR glasses, characterized in that: The method comprises: Acquire a multimodal data set for a target device, wherein the multimodal data set includes visual modality data, audio modality data, and IMU modality data; Using a lightweight convolutional neural network corresponding to the visual modality encoder, extracting features from the visual modality data to obtain dynamic visual features; Using a deep convolutional neural network corresponding to the audio modality encoder to extract features from the audio modality data to obtain audio key features; Using a long short-term memory network corresponding to the IMU modal encoder, feature extraction is performed on the IMU modal data to obtain time series features; Scoring the dynamic visual features, the audio key features, and the time series features to obtain a scoring result; According to the scoring results, determine the processing priority of the AR glasses for the visual modality data, the audio modality data, and the IMU modality data.

2. The multimodal data processing method based on AR glasses according to claim 1, characterized in that: The lightweight convolutional neural network corresponding to the visual modality encoder is used to extract features from the visual modality data to obtain dynamic visual features, specifically including: Determining, according to the visual modality data, a visual image corresponding to the visual modality data; The lightweight convolutional neural network is used to perform image analysis on the visual image to obtain the changed area in the visual image to obtain the dynamic visual feature.

3. The multimodal data processing method based on AR glasses according to claim 1, characterized in that: The deep convolutional neural network corresponding to the audio modality encoder is used to extract features from the audio modality data to obtain audio key features, specifically including: Determine, according to the audio modal data, a frequency spectrogram corresponding to the audio modal data; The deep convolutional neural network is used to determine high-frequency noise and / or abnormal audio in the spectrogram to obtain the audio key features.

4. The multimodal data processing method based on AR glasses according to claim 1, characterized in that: The long short-term memory network corresponding to the IMU modal encoder is used to extract features from the IMU modal data to obtain time series features, specifically including: Calculating a movement trajectory of a user's head direction to determine a user focus area of ​​the target device, wherein the user wears the AR glasses; Through the long short-term memory network, when the user faces the user focus area, subtle motion mutations are detected to obtain the time series features.

5. The multimodal data processing method based on AR glasses according to claim 1, characterized in that: Scoring the dynamic visual features, the audio key features, and the time series features to obtain a scoring result specifically includes: If it is determined that the dynamic visual feature indicates that a red light of the target device is flashing, generating a first score; If it is determined that the audio key feature indicates that the target device emits a sharp noise, generating a second score; If it is determined that the time series feature indicates that the user turns to the target device, a third score is generated, the first score is higher than the third score, the second score is higher than the third score, and the score difference between the first score and the second score is within a preset range; The scoring result is determined based on the first score, the second score, and the third score.

6. The multimodal data processing method based on AR glasses according to claim 5, characterized in that: Determining, according to the scoring result, the processing priority of the AR glasses for the visual modality data, the audio modality data, and the IMU modality data, specifically includes: Determining a first processing priority according to the first score; determining a second processing priority according to the second score; A third processing priority is determined according to the third score, the first processing priority is higher than the third processing priority, the second processing priority is higher than the third processing priority, and the first processing priority is the same as the second processing priority.

7. The multimodal data processing method based on AR glasses according to claim 1, characterized in that: The method further comprises: receiving a personalized processing priority for the multimodal data set sent by a user device; The visual modality data, the audio modality data, and the IMU modality data are processed according to the personalized processing priority.

8. A multimodal data processing device based on AR glasses, characterized in that: The multimodal data processing device comprises an acquisition module (31) and a processing module (32), wherein: The acquisition module (31) is used to acquire a multimodal data set for a target device, wherein the multimodal data set includes visual modal data, audio modal data, and IMU modal data; The processing module (32) is used to use a lightweight convolutional neural network corresponding to the visual modality encoder to extract features from the visual modality data to obtain dynamic visual features; The processing module (32) is further used to use a deep convolutional neural network corresponding to the audio modality encoder to perform feature extraction on the audio modality data to obtain audio key features; The processing module (32) is further used to extract features from the IMU modal data using a long short-term memory network corresponding to the IMU modal encoder to obtain time series features; The processing module (32) is further used to score the dynamic visual features, the audio key features and the time series features to obtain a scoring result; The processing module (32) is further used to determine the processing priority of the AR glasses for the visual modality data, the audio modality data and the IMU modality data according to the scoring result.

9. An electronic device, characterized in that: The electronic device comprises a processor (41), a memory (45), a user interface (43) and a network interface (44), wherein the memory (45) is used to store instructions, the user interface (43) and the network interface (44) are both used to communicate with other devices, and the processor (41) is used to execute the instructions stored in the memory (45) so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.

Citation Information

Cited By

  • Mixed reality intelligent inspection method and system based on multi-modal fusion

    CN120579148A