Adaptive interaction methods, devices, equipment, and media for smart audio glasses

By acquiring and analyzing multi-source data from smart audio glasses, an environmental interaction context vector is generated, which solves the problem of poor interaction accuracy in existing technologies and achieves a more accurate and personalized interactive experience.

CN122086249APending Publication Date: 2026-05-26BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING SUPERHEXA CENTURY TECH CO LTD
Filing Date
2026-04-23
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing smart audio glasses interaction methods rely on single voice data or simple head movement data, resulting in poor interaction accuracy in complex environments and an inability to accurately understand user scenarios and intentions.

Method used

By acquiring multi-source data (head motion data and environmental audio data), performing time synchronization processing, extracting multimodal features, jointly analyzing head motion features and environmental audio features, generating an environmental interaction context vector, determining the interaction target, and generating a dynamic interaction strategy based on the vector.

Benefits of technology

It improves the accuracy of speech recognition, reduces interaction failures or misoperations, and provides a more accurate, intelligent, and adaptive personalized interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122086249A_ABST
    Figure CN122086249A_ABST
Patent Text Reader

Abstract

This application provides an adaptive interaction method, device, and medium for smart audio glasses, belonging to the fields of smart wearable devices and artificial intelligence technology. The method includes: performing time synchronization processing and multimodal feature extraction on multi-source data to obtain head motion features and environmental audio features; jointly analyzing the head motion features and environmental audio features to determine cross-modal correlations, and generating an environmental interaction context vector based on the current working mode of the smart audio glasses to determine the corresponding interaction target; thereby selecting a corresponding feature subset from the head motion features and environmental audio features; fusing the feature subset based on the interaction target; and performing dynamic strategy prediction based on the fused features to obtain a dynamic interaction strategy, thereby adjusting the audio output and / or interaction mode of the smart audio glasses. This application can achieve a more accurate, intelligent, and adaptive personalized interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of smart wearable devices and artificial intelligence technology, and more specifically, relates to a smart audio glasses adaptive interaction method, device, and medium. Background Technology

[0002] In today's era of rapid technological development, smart audio glasses, as an innovative wearable device, not only possess the visual functions of traditional glasses but also integrate advanced functions such as audio playback and voice interaction, providing users with a convenient way to enjoy audio and interact in mobile scenarios.

[0003] Currently, existing smart audio glasses interaction methods are mainly based on single voice data input or simple head sensor data-assisted interaction. For example, they rely solely on voice recognition technology to understand the user's voice commands, thereby realizing basic interactive functions such as audio playback control and information query; or they use head nodding or shaking movements to assist in confirming voice commands or performing simple operation switching.

[0004] However, this interaction method has several significant drawbacks. First, relying solely on single voice data leads to a sharp drop in speech recognition accuracy under complex audio interference environments, resulting in interaction failures or misoperations. Second, when combined with simple head motion sensor data, the lack of comprehensive analysis and deep fusion of multiple data sources makes it impossible to fully and accurately understand the user's behavioral intentions and changes in the scene's state. For example, head movement patterns may be similar in different scenarios such as resting or working, but the actual needs are drastically different, making it difficult to accurately determine the user's intentions based on subtle differences, and thus unable to provide personalized interaction strategies. Summary of the Invention

[0005] The purpose of this application is to provide an adaptive interaction method, device, or medium for smart audio glasses, in order to solve the technical problem that relying solely on single voice data or simple head movement data leads to poor interaction accuracy in complex environments and an inability to accurately understand user scenarios and intentions, thereby achieving a more accurate, intelligent, and adaptive personalized interactive experience.

[0006] A first aspect of this application provides an adaptive interaction method for smart audio glasses, comprising: Acquire multi-source data, perform time synchronization processing on the multi-source data, and obtain processed multi-source data; the multi-source data includes head motion data and environmental audio data. Multimodal feature extraction was performed on the processed multi-source data to obtain head motion features and environmental audio features, respectively. Joint analysis of head motion features and environmental audio features is performed to determine cross-modal correlations. Combined with the current working mode of the smart audio glasses, an environmental interaction context vector is generated. The environmental interaction context vector represents the current contextual semantics that integrates the user's action intention and the environmental state. The corresponding interaction target is determined based on the environmental interaction context vector; the interaction target represents the task category that needs to be performed in the current context semantics. Based on the interaction target, a subset of features corresponding to the interaction target is selected from head motion features and environmental audio features; Feature subsets are fused based on interactive objectives to generate fused features; Dynamic policy prediction is performed based on fusion features to obtain dynamic interaction policies, and the audio output and / or interaction mode of the smart audio glasses are adjusted according to the dynamic interaction policies.

[0007] A second aspect of this application provides an adaptive interaction device for smart audio glasses, comprising: The multi-source data processing module is used to acquire multi-source data, perform time synchronization processing on the multi-source data, and obtain processed multi-source data; the multi-source data includes head motion data and environmental audio data; The feature extraction module is used to extract multimodal features from the processed multi-source data, obtaining head motion features and environmental audio features respectively. The environmental interaction context vector determination module is used to jointly analyze head motion features and environmental audio features to determine cross-modal correlations. Combined with the current working mode of the smart audio glasses, it generates environmental interaction context vectors. The environmental interaction context vectors represent the current contextual semantics that integrate the user's action intentions and the environmental state. The interaction target determination module is used to determine the corresponding interaction target based on the environmental interaction context vector; the interaction target represents the task category that needs to be performed in the current context semantics. The feature subset determination module is used to select a feature subset corresponding to the interaction target from head motion features and environmental audio features based on the interaction target. The feature fusion acquisition module is used to fuse feature subsets based on the interaction target to generate fused features; The dynamic interaction strategy determination module is used to predict dynamic strategies based on fusion features, obtain dynamic interaction strategies, and adjust the audio output and / or interaction mode of the smart audio glasses according to the dynamic interaction strategies.

[0008] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described adaptive interaction method for smart audio glasses.

[0009] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described adaptive interaction method for smart audio glasses.

[0010] The beneficial effects of the smart audio glasses adaptive interaction method, device, equipment, and medium provided in this application embodiment are as follows: This application embodiment acquires multi-source data (head motion data and environmental audio data), performs time synchronization processing on the multi-source data to obtain processed multi-source data, and then performs multimodal feature extraction and other operations on the processed multi-source data. This application embodiment comprehensively utilizes multiple data methods, not relying solely on voice data. When facing complex environmental audio interference, it can use other information such as head motion data to assist in judgment, thereby improving the accuracy of speech recognition and reducing interaction failures or misoperations. This application embodiment also performs joint analysis of head motion features and environmental audio features to determine cross-modal correlations. Combined with the current working mode of the smart audio glasses, it generates an environmental interaction context vector to more comprehensively grasp the user's intent and scene state. Based on the environmental interaction context vector, it determines the interaction target, selects the corresponding feature subset, generates fused features, and performs dynamic strategy prediction based on the fused features to provide personalized interaction strategies. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating an embodiment of the adaptive interaction method for smart audio glasses provided in this application. Figure 2 This is a structural block diagram of an adaptive interaction device for smart audio glasses provided in an embodiment of this application; Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0014] It should be noted that the terms "first," "second," etc., used in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein.

[0015] It is understood that in the embodiments of this application, data such as user head movement data and environmental audio data are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0017] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the adaptive interaction method for smart audio glasses provided in this application. The adaptive interaction method for smart audio glasses provided in this application embodiment can be executed by smart audio glasses, and the method may include: S101: Acquire multi-source data, perform time synchronization processing on the multi-source data, and obtain the processed multi-source data.

[0018] In this embodiment, the multi-source data includes head motion data and ambient audio data. This embodiment uses an inertial measurement unit (IMU) integrated within the smart audio glasses to collect head motion data in real time. This head motion data typically contains information in multiple dimensions, such as three-axis acceleration, three-axis angular velocity, and head attitude angles (e.g., pitch, yaw, and roll angles) and their rates of change calculated using a sensor fusion algorithm. Ambient audio data is collected in real time using a microphone array integrated within the smart audio glasses. This environmental data is a time-domain audio signal and may include ambient background noise, specific acoustic events (e.g., car horns, human conversations, and music). The microphone array can also estimate the preliminary location of the sound source.

[0019] This embodiment performs time synchronization processing on multi-source data, enabling the synchronized multi-source data to accurately reflect the physical state at the same moment. This embodiment encapsulates aligned head motion data and environmental audio data to generate processed multi-source data. This processed multi-source data can be time-indexed data frames or data packets, where each time point or time period corresponds to a set of synchronized head motion parameters (e.g., acceleration, angular velocity, and displacement) and environmental audio parameters (e.g., time-domain signals or preliminary frequency-domain characteristics).

[0020] S102: Perform multimodal feature extraction on the processed multi-source data to obtain head motion features and environmental audio features respectively.

[0021] This embodiment extracts features from the processed multi-source data. First, the head motion data is low-pass filtered to remove high-frequency noise. Then, a Kalman filter is used to compensate for and eliminate the effects of device jitter or slow drift, resulting in stable and accurate information on head orientation and spatial position changes. Based on the Kalman filter method described above, this embodiment can determine the head's posture in three-dimensional space (which can be represented by Euler angles) to describe the head's orientation; simultaneously, it can obtain the head's motion parameters, which may include linear acceleration, angular velocity, and displacement.

[0022] In this embodiment, the temporal features of the preprocessed motion parameters are extracted within a set sliding time window. For example, the mean and variance of motion amplitude (reflecting the average intensity and fluctuation of the action), the peak and trough of motion (capturing sudden and violent movements), the zero-crossing rate (measuring the frequency of changes in the direction of the action), and the energy integral are extracted.

[0023] This embodiment uses Fast Fourier Transform (FFT) to perform time-frequency transformation on the aforementioned motion parameters, converting the time-domain motion signal (selecting a composite signal of one or more axes of linear acceleration or angular velocity as the motion signal) to the frequency domain to obtain the corresponding spectrum. Based on this spectrum, frequency domain features are extracted, including the main frequency components (the core rhythm or rate of periodic movements of the user's head), frequency band energy distribution (distinguishing between slow head rotations and rapid nodding movements), and spectral entropy (quantifying the regularity or complexity of the motion pattern). These frequency domain features can quantify the rhythm, energy distribution, and regularity of head movements.

[0024] This embodiment utilizes frequency domain features for higher-level semantic action recognition. Parameters such as head posture angle, angular velocity, and linear acceleration in a specific direction (gravity reference direction) under continuous timestamps are used to generate a motion sequence in chronological order. This motion sequence characterizes the continuous movement state and changing trend of the head over a certain time period. By analyzing this motion sequence, action patterns with specific semantics are identified. For example, by detecting specific combinations and thresholds of angular velocity and posture angle, discrete action units such as nodding, shaking, raising the head, and turning the head to locate sounds, along with their duration, amplitude, and frequency, can be identified.

[0025] This embodiment preprocesses environmental audio data, including pre-emphasis, framing, and windowing, to obtain a preprocessed audio signal that enhances high-frequency components and meets the requirements of short-time stationarity analysis. Pre-emphasis processing uses a first-order finite impulse response high-pass filter to enhance high-frequency components and compensate for high-frequency attenuation during transmission or acquisition. Framing divides the pre-emphasized continuous audio signal into a series of short frames. For example, a fixed-duration (20-40 ms) sliding window is used, with a certain overlap rate (30%-50%) between adjacent frames to smooth inter-frame variations and capture the short-time stationarity characteristics of the signal. Windowing multiplies each frame signal by a preset window function (e.g., a Hamming or Hanning window) to reduce spectral leakage at both ends of the frame caused by framing.

[0026] This embodiment extracts basic acoustic features from each frame of preprocessed audio signal, such as Mel frequency cepstral coefficients to characterize the short-time power spectrum characteristics of sound; spectral shape features (centroid, bandwidth, and roll-off point) to describe the spectral shape; and time-domain features (short-time energy and zero-crossing rate) to characterize the basic physical properties of the sound signal in the time dimension.

[0027] This embodiment performs higher-level semantic feature extraction based on basic acoustic features. Using preset rules, it identifies the acoustic scene category of the current environment (e.g., a quiet office, a noisy street, or a multi-person conversation area) and detects specific acoustic events (e.g., the occurrence of speech, vehicle horns, or music playback) and their start time, duration, and relative intensity.

[0028] In this embodiment, the extracted head motion features and environmental audio features are organized into structured feature vectors or feature sequences. These features are typically organized according to the same or corresponding time base (e.g., several frames per second) to ensure that each feature is aligned with the features in the processed multi-source data in the time dimension. Finally, the head motion feature sequence and the environmental audio feature sequence are output.

[0029] S103: Perform joint analysis of head motion features and environmental audio features to determine cross-modal correlations, and generate environmental interaction context vectors by combining the current working mode of the smart audio glasses.

[0030] In this embodiment, the environmental interaction context vector representation integrates the user's action intent and the current contextual semantics of the environmental state.

[0031] One specific application scenario of this embodiment is as follows: the smart audio glasses are in music playback mode. At this time, by analyzing the environmental audio features, two acoustic events with clear spatial attributes are detected: a car horn sound from the left and a bicycle bell sound from the right. Subsequently, through head movement feature analysis, it is identified that the user made a movement of first turning their head to the left and then turning their head to the right.

[0032] This embodiment can analyze the temporal and spatial correlations between head movement features and environmental audio features to determine cross-modal associations. For example, the action of "turning the head to the left" is closely related to the environmental audio of "a car horn on the left" in time and is consistent in spatial direction (left side), thus determining a cross-modal association: the user turning the head to the left may be in response to the horn sound on the left, and turning the head to the right may be in response to the bicycle bell sound on the right. The order of the head turning actions may imply that the user's current attention focus is more inclined to the horn sound on the left, and the action implies an intention to pay attention to and confirm the potential dangerous sound source.

[0033] This embodiment interprets the cross-modal association in a contextualized manner based on the current operating mode (music playback) of the smart audio glasses. For example, in music playback mode, the aforementioned cross-modal association (paying attention to sudden dangerous sounds on the left) is interpreted as: the user has a potential need for safe listening in the current environment.

[0034] This embodiment encodes the interpretation results, key features (e.g., the direction of the sound source of interest, the type of event), and the working mode together to generate a unified environmental interaction context vector. In this embodiment, this vector can represent the complex scenario of "a user turning their head to pay attention to a sudden horn sound on the left and a bicycle bell sound on the right while playing music."

[0035] In one embodiment of this application, head motion features and environmental audio features are jointly analyzed to determine cross-modal correlations, and an environmental interaction context vector is generated by combining the current operating mode of the smart audio glasses, including: Cross-modal correlation networks are used to jointly analyze head motion features and environmental audio features to obtain cross-modal correlation relationships; these relationships are then used to quantify the dynamic correlation between head motion and environmental audio. Based on cross-modal associations, determine potential user intent; Encode the current working mode to obtain the encoding vector of the current working mode; The potential user intent, head movement features, and environmental audio features are vectorized and concatenated, and then fused with the encoded vector to generate an environmental interaction context vector.

[0036] The cross-modal association network includes a first attention subnetwork and a second attention subnetwork. Based on cross-modal association networks, joint analysis of head motion features and environmental audio features is performed to obtain cross-modal association relationships, including: Temporally aligned encoding is performed on head motion features and ambient audio features to obtain temporally aligned head motion codes and ambient audio codes; Head motion encoding is used as the query vector, and ambient audio encoding is used as the key vector and value vector. The query vector, key vector, and value vector are input into the first attention sub-network to obtain the first cross-modal attention map. The first cross-modal attention map is used to characterize the semantic contribution of ambient audio information to head motion at each moment. The ambient audio encoding is used as the query vector, and the head motion encoding is used as the key vector and value vector. The query vector, key vector and value vector are input into the second attention sub-network to obtain the second cross-modal attention map. The second cross-modal attention map is used to characterize the semantic contribution of head motion information to the ambient audio at each time step. The first cross-modal attention map and the second cross-modal attention map are fused to generate a cross-modal attention weight matrix; The head motion encoding and environmental audio encoding are weighted and fused based on the cross-modal attention weight matrix to obtain the cross-modal joint feature vector; Cross-modal associations are obtained based on cross-modal joint feature vectors.

[0037] In this embodiment, the head motion feature sequence and the environmental audio feature sequence, which have been synchronized in time, are input into the first encoder and the second encoder, respectively. The first encoder performs deep feature extraction on the head motion feature sequence and outputs a head motion code that is aligned in the feature dimension and contains temporal information. The second encoder performs deep feature extraction on the environmental audio feature sequence and outputs an environmental audio code that is aligned in the feature dimension and contains temporal information.

[0038] In this embodiment, head motion encoding is used as the query vector, and ambient audio encoding is used as the key and value vectors, respectively, and input into the first attention sub-network. The query vector, key vector, and value vector are projected onto a learnable weight matrix through the linear transformation layer of the first attention sub-network, mapping the three vectors to the same feature dimension. These three vectors with the same feature dimension are then input into the attention layer of the first attention sub-network to calculate the dot product similarity between each time step of the head motion (query) and all time steps of the ambient audio (key), obtaining the original attention score matrix. This original attention score matrix is ​​then input into the scaling layer of the first attention sub-network to scale the original attention score matrix, obtaining a scaled score matrix. The scaled score matrix is ​​then input into the activation layer of the first attention sub-network, where a softmax function is applied row-wise to the scaled score matrix to generate an attention weight matrix. The attention weight matrix is ​​then input into the output layer of the first attention sub-network to output the first cross-modal attention map. The first cross-modal attention map is used to characterize the semantic contribution of ambient audio information to each time step of the head motion.

[0039] In this embodiment, the ambient audio encoding is used as the query vector, and the head motion encoding is used as the key vector and value vector, respectively, and input into the second attention sub-network. The internal layer structure of the second attention sub-network is the same as that of the first attention sub-network. Similarly, the second attention sub-network processes the input vectors with reversed roles in a symmetrical structure: using the ambient audio encoding as the query vector and the head motion encoding as the key vector and value vector, and sequentially passing through the same linear transformation layer, attention layer, scaling layer, activation layer, and output layer, finally outputting the second cross-modal attention map. The second cross-modal attention map is used to characterize the semantic contribution of head motion information to the ambient audio at each time step.

[0040] This embodiment fuses the first and second cross-modal attention maps (e.g., by matrix addition, averaging, or concatenation followed by fusion using a lightweight network) to generate a cross-modal attention weight matrix. This matrix integrates bidirectional attention information, characterizing the inter-modal correlation strength at all times.

[0041] This embodiment performs weighted fusion of head motion encoding and ambient audio encoding based on a cross-modal attention weight matrix. Specifically, this weight matrix can be used as a guide to interactively weight and sum or concatenate head motion encoding and ambient audio encoding to generate a cross-modal joint feature vector that integrates bimodal depth information.

[0042] In this embodiment, the cross-modal joint feature vector is used as the cross-modal association relationship.

[0043] This embodiment can also input the aforementioned head movement features and environmental audio features into a cross-modal association network. This cross-modal association network analyzes the degree of matching between movement and audio signals at each time point through its internal correlation calculation mechanism. In this embodiment, the cross-modal association network can determine that there is a strong positive correlation between the user's left-turning head movement time interval and the occurrence time and spatial location information of the car horn sound on the left; subsequently, there is also a strong positive correlation between the user's right-turning head movement time interval and the occurrence time and spatial location information of the bicycle bell sound on the right. Thus, the positive correlation is taken as the cross-modal association relationship.

[0044] This embodiment utilizes a first attention subnetwork and a second attention subnetwork to calculate cross-modal attention maps from different directions, comprehensively revealing the semantic contribution of ambient audio to head motion and head motion to ambient audio. The bidirectional analysis method provided in this embodiment can more meticulously and deeply explore the complex relationships between the two modalities, improving the comprehensiveness and accuracy of the association analysis.

[0045] In this embodiment, no matter how complex and varied the ambient audio or how diverse the user's head movement patterns are, a series of operations such as temporal alignment, attention calculation, and feature fusion can effectively extract the associated information, enabling the smart audio glasses to maintain good interactive performance in different scenarios and enhancing the system's adaptability and robustness to complex environments.

[0046] In this embodiment, the head movement can be a sequence of continuous head-turning actions, first quickly to the left and then quickly to the right. This sequence includes the direction, angular velocity, and precise timestamp of each action. The acoustic events identified by the environmental audio features can be car horns and bicycle bells. Spatial audio analysis determines that the two sound sources originate from the "left" and "right" sides, respectively, and includes their respective start times and intensity information.

[0047] This embodiment analyzes the cross-modal correlation relationship output by the cross-modal correlation network. This correlation relationship indicates that the user's head rotation is not random or aimless, but rather precisely and sequentially points towards the sources of two sudden ambient sounds. This embodiment calculates the time difference between the start time of the acoustic event and the start time of the corresponding head movement, and compares the spatial difference between the azimuth angle of the sound source and the pointing angle of the head movement. This embodiment pairs head movements with event features that satisfy the temporal following condition (e.g., time difference within 100-800 milliseconds) and the spatial alignment condition (e.g., spatial difference less than 20 degrees). In this embodiment, the action of turning the head to the left forms one pair with the horn sound on the left, and the action of turning the head to the right forms another pair with the bell sound on the right. Based on this strong temporal following and spatial alignment pattern between the "specific acoustic event" and the "directional head movement pointing towards the sound source," this embodiment infers that the user's potential behavioral purpose is to actively explore the sound source to assess the environmental situation. Based on this, the potential user's intention is determined to be "to conduct a safety exploration of sudden ambient sounds on both sides."

[0048] In this embodiment, the current operating mode of the device (smart audio glasses) is determined to be "music playback mode". This mode has a predefined unique encoding vector. This encoding vector contains contextual information indicating that in this mode, the device's primary task is to play music content, and it may need to maintain a certain level of awareness of environmental safety.

[0049] This embodiment transforms the identified potential user intent of "conducting a safety investigation into sudden ambient noise on both sides" into a corresponding intent vector representation. This intent vector, along with key head movement features at the current moment (e.g., completion status of head turning, current head orientation), and key environmental audio features (e.g., remaining ambient noise level, whether there is still sudden noise), are vectorized and concatenated to form a preliminary semantic fragment that integrates the states of "person" and "environment." This preliminary semantic fragment is then fused with an encoded vector representing "music playback mode" to obtain an environmental interaction context vector. In this embodiment, the fusion process ensures that the final vector (environmental interaction context vector) can represent both "the user is investigating the environment due to sudden noise" and "the device is currently in music playback mode."

[0050] Among these, determining potential user intent based on cross-modal associations includes: Cross-modal correlations are analyzed to extract co-modal features that characterize the temporal and intensity-related changes of head motion and ambient audio. These co-modal features include one or more combinations of temporal following tightness, intensity change coupling, and event co-occurrence frequency between head motion and ambient audio. Based on the characteristics of the collaborative pattern, multiple candidate basic intentions are identified; Based on the current working mode of the smart audio glasses and the candidate intent priority strategy, multiple candidate basic intents are prioritized, and the candidate basic intent with the highest priority is taken as the potential user intent.

[0051] In this embodiment, by parsing cross-modal correlations, cooperative mode features are extracted. These cooperative mode features include one or more combinations of the temporal following tightness of head movements and environmental audio, the coupling degree of intensity changes, and the event co-occurrence frequency. The specific method for extracting cooperative mode features is as follows: Specifically, for each detected acoustic event, its timestamp is located in the cross-modal attention weight matrix, and a peak point where attention significantly increases is searched along the time axis in the head motion modality; the difference between the event start time and the peak attention point is determined; and statistical analysis is performed on multiple differences within the time window to obtain a quantitative value of the temporal follow-up tightness. In this embodiment, the smaller the difference and the more concentrated its distribution, the tighter the temporal follow-up.

[0052] This embodiment analyzes the first time difference between the starting time of the user's "turning head to the left" action and the starting time of the "car horn sound on the left" event, and the second time difference between the "turning head to the right" action and the "bicycle bell sound on the right" event. Both the first and second time differences are less than a preset time difference threshold (the preset time difference threshold can be 200 milliseconds), indicating that the user's head-turning action closely follows the corresponding sound event, demonstrating a high degree of temporal correlation. The preset time difference threshold in this embodiment can be set based on the typical human reaction time range to sudden stimuli in ergonomics, used to capture the reasonable physiological response interval between the user's perception of sound and the making of a definite head-turning action.

[0053] In this embodiment, within the identified time interval of significantly increased attention, the intensity representation of head movement and the intensity representation of ambient audio are simultaneously acquired within that interval; the statistical correlation between the intensity sequences of the two is obtained using the Pearson correlation coefficient method. This correlation coefficient is the quantification value of the coupling degree of intensity change. In this embodiment, the higher the correlation, the more synchronous the intensity changes of the represented action and the sound. Both acoustic events in this embodiment are sudden events, and their intensities reach their peaks almost simultaneously, demonstrating strong coupling of intensity changes.

[0054] Within a set analysis time window, based on cross-modal attention weights and preset thresholds, the number of times a specific type of head movement event and a specific type of acoustic event are simultaneously identified as associated is recorded. This number of simultaneous associations is taken as the co-occurrence frequency. This co-occurrence frequency is normalized to obtain a quantified value of the event co-occurrence frequency. A higher co-occurrence frequency indicates that the combination pattern is more typical in the current context. For example, in this embodiment, within the short analysis window, the "turning head to the left" event co-occurs once with the "sudden high-volume event on the left (horn)" event; and the "turning head to the right" event co-occurs once with the "sudden high-volume event on the right (bell)" event.

[0055] This embodiment matches the extracted collaborative pattern features with a preset intent feature library to generate multiple candidate basic intents. The preset intent feature library in this embodiment is used to map the collaborative pattern features to candidate basic intents with clear behavioral semantics. The preset intent feature library can be obtained by experts annotating various multi-source data and analyzing the regular collaborative patterns between head movements and ambient audio under different intents.

[0056] This embodiment determines a preset candidate intent priority strategy based on the current music playback mode. The preset candidate intent priority strategy is as follows: when there are candidate basic intents directly related to environmental safety (detecting / locating sound sources, actively scanning the environment to assess safety), the priority of such candidate basic intents is set higher than that of ordinary, interfered candidate basic intents. Based on this candidate intent priority strategy, the generated candidate basic intents are prioritized.

[0057] In this embodiment, the candidate underlying intent of "actively scanning the environment to assess security" more comprehensively describes the user's behavior pattern of continuously and purposefully exploring potential risks on both sides, and is strongly correlated with security. Therefore, it is determined to be the highest priority in this embodiment. Based on this, this embodiment determines "actively scanning the environment to assess security" as the final potential user intent.

[0058] This embodiment extracts semantically meaningful behavioral patterns and, by combining the working mode of smart audio glasses with a preset candidate intent priority strategy, infers the user intent that best fits the current scenario, providing a clear direction for subsequent precise interaction.

[0059] S104: Determine the corresponding interaction target based on the environmental interaction context vector.

[0060] In this embodiment, the interaction target represents the type of task that needs to be performed in the current context semantics.

[0061] In this embodiment, based on the determined environmental interaction context vector, the composite contextual semantics represented by the vector are parsed into a clear and executable task instruction, i.e., the interaction goal is determined. The potential user intent (e.g., actively scanning the environment to assess safety), the current device operating mode (music playback mode), and key environmental states (bilateral sudden safety-related acoustic events) contained in the vector are parsed out. Next, the vector is compared with multiple candidate contextual vectors pre-stored in the music playback mode for similarity. Each candidate contextual vector is associated with a preset interaction goal, such as maintaining immersive playback, temporarily lowering the volume, pausing playback and activating ambient sound enhancement, or triggering a safety warning. In this embodiment, since the environmental interaction context vector strongly points to the user's behavior of interrupting music immersion due to safety concerns, the current interaction goal is determined to be "enhanced environmental perception." This interaction goal clarifies that the core task category to be executed in the next stage is to prioritize ensuring the user's clear perception of environmental sounds, providing a fundamental basis for subsequently formulating specific audio output and interaction control strategies.

[0062] In one embodiment of this application, before determining the corresponding interaction target based on the environmental interaction context vector, the method further includes: Obtain the scene vector set corresponding to the smart audio glasses in different working modes, where each scene vector in the scene vector set corresponds to a candidate interaction target; Among them, determining the corresponding interaction target based on the environmental interaction context vector includes: Determine the corresponding set of scenario vectors based on the current working mode of the smart audio glasses; Obtain the similarity between the environmental interaction context vector and each context vector in the context vector set; The maximum similarity among all similarities is selected. If the maximum similarity exceeds the preset scenario matching threshold, the candidate interaction target associated with the scenario vector corresponding to the maximum similarity is taken as the interaction target.

[0063] This embodiment pre-constructs a set of scenario vectors (pre-stored in multiple candidate scenario vectors under the music playback mode) for each working mode of the smart audio glasses. Each scenario vector can be manually defined or learned from historical data to describe a specific interaction scenario that may occur in that working mode. For example, for the "music playback mode," the scenario vector set may include scenario vector A, scenario vector B, and scenario vector C. Scenario vector A represents a scenario where the user is immersed in listening, the environment is quiet, and there is no intention for the user to actively interact; its associated candidate interaction goal could be "maintain immersive playback." Scenario vector B represents a scenario where a sudden safety-related acoustic event (horn blaring) occurs in the environment, and the user shows an intention to actively investigate; its associated candidate interaction goal could be "temporarily enhance environmental awareness." Scenario vector C represents a scenario where the user actively issues a voice command; its associated candidate interaction goal could be "activate the voice assistant."

[0064] In this embodiment, after obtaining the current environmental interaction context vector, the similarity of each scenario vector is calculated based on the scenario vector set corresponding to the current working mode; the maximum similarity value is compared with the preset similarity threshold to determine the interaction target.

[0065] In this embodiment, the current working mode of the smart audio glasses is read. Taking the aforementioned scenario as an example, the current mode is music playback mode, thereby calling the scene vector set corresponding to this working mode. The cosine similarity method is used to determine the similarity between the environmental interaction context vector and each scene vector, obtaining a first similarity, a second similarity, and a third similarity. This similarity is used to characterize the semantic proximity between the current real scene and each candidate scene. The maximum similarity value is selected from the first similarity, the second similarity, and the third similarity, and compared with a preset similarity threshold. This preset similarity threshold is used to ensure that the matching result has sufficient confidence and avoids false triggering.

[0066] If the maximum similarity value exceeds a preset similarity threshold, the current scenario is determined to have successfully matched the corresponding candidate scenario, and the candidate interaction target associated with the scenario vector corresponding to the maximum similarity value is taken as the interaction target. If the maximum similarity value does not exceed the preset similarity threshold, the current environment interaction context vector is determined to have not formed a valid match with any preset scenario vectors. In this case, the current interaction target will not be updated, and the existing interaction target and device behavior will remain unchanged; alternatively, a preset default interaction target corresponding to the current working mode can be enabled (for example, in music playback mode, the default interaction target is to maintain immersive playback) to ensure the continuity and stability of device interaction.

[0067] For example, the environmental interaction context vector generated in this embodiment strongly represents the semantics of "the user actively conducts a security investigation due to sudden safety sounds (horn, bell) on both sides during music playback". When this vector is compared with each scenario vector in the "music playback mode" scenario vector set, it best matches the semantics of scenario vector B (representing "sudden security event + user investigation"), that is, the similarity (second similarity) between the environmental interaction context vector and scenario vector B is the highest. If the second similarity exceeds a preset similarity threshold, it is determined that the current scenario successfully matches the corresponding preset scenario, and the candidate interaction target associated with scenario vector B corresponding to the second similarity is determined as the current final interaction target.

[0068] This embodiment obtains a set of scene vectors for smart audio glasses under different working modes, with each scene vector corresponding to a candidate interaction target. The interaction target is determined based on the similarity between the environmental interaction context vector and the scene vector. This embodiment comprehensively considers the scene information under multiple working modes and the current environmental context, which can more accurately match the appropriate interaction target, reduce misjudgments, and improve the accuracy and reliability of the interaction.

[0069] This embodiment combines environmental interaction context vectors to enable smart audio glasses to dynamically select interaction targets based on the current specific environment. This allows the smart audio glasses to provide interactive services that meet user needs and environmental characteristics in different scenarios, enhancing their adaptability to complex and changing environments. Furthermore, this embodiment selects candidate interaction targets associated with the context vectors corresponding to the highest similarity as the interaction target. When the highest similarity exceeds a preset context matching threshold, interaction is confirmed. This explicit judgment mechanism can quickly and effectively determine the interaction target, reducing user waiting time and making the interaction process smoother and more natural, thereby improving the user experience of the smart audio glasses.

[0070] S105: Based on the interaction target, select a subset of features corresponding to the interaction target from head motion features and environmental audio features.

[0071] In this embodiment, the interaction goal is to temporarily enhance environmental perception. Based on this goal, the core objective is to understand the user's spatial exploration behavior in response to sudden, potentially risky environmental sounds. Therefore, this embodiment extracts features related to "actively scanning the environment to assess safety" from head movement characteristics, such as the precise angles of turning the head left and right, peak angular velocity, and the start and end times of these movements, while filtering out irrelevant features such as slight swaying synchronized with the music rhythm. Simultaneously, from environmental audio features, the azimuth, sound pressure level, duration, and timestamps of car horns and bicycle bells are selected, potentially temporarily ignoring background music or stable wind noise. Through this guided filtering of the interaction goal, two concise, highly focused feature subsets (head movement feature subset and environmental audio feature subset) are obtained. These two feature subsets together accurately characterize the crucial information of "what kind of exploration response the user makes to what kind of sound," which is essential for achieving the current interaction goal.

[0072] In one embodiment of this application, before filtering the feature subset corresponding to the interaction target from head motion features and environmental audio features based on the interaction target, the method further includes: The temporal attention sequence is extracted from the collaborative pattern features. Each attention value in the temporal attention sequence is used to characterize the correlation strength between head movement and ambient audio at the corresponding time point. Specifically, based on the interaction target, a subset of features corresponding to the interaction target is selected from head motion features and environmental audio features, including: Based on the interaction target and the preset first mapping rule, the feature filtering mode corresponding to the current interaction target is determined; the preset first mapping rule is used to characterize the mapping between the interaction target and the feature filtering mode. Based on the feature selection mode, the time-series attention sequence is divided into time segments to obtain the divided time segments. Then, the attention values ​​within each time segment are analyzed to obtain the pattern type of the time segment. The pattern type is used to characterize the change pattern of the correlation between head movement and environmental audio within the time segment. Based on the mode type, the environmental complexity in the environmental interaction context vector, and the current working mode of the smart audio glasses, dynamic selection conditions for selecting key time segments are determined; environmental complexity is used to characterize the degree of disorder in the current environmental audio data. Based on dynamic selection criteria, key time segments are determined from the divided time segments; Based on key time segments, features corresponding to the time points are selected from head motion features as a subset of head motion features, and features corresponding to the time points are selected from environmental audio features as a subset of environmental audio features.

[0073] This embodiment extracts a temporal attention sequence from the acquired collaborative pattern features. This temporal attention sequence is a numerical sequence aligned with the time axis, where each value (an attention value that has been normalized to between 0 and 1) quantifies the strength of the association between head movement and ambient audio at the corresponding time point or within a very short time window. For example, a higher attention value indicates a closer connection between the user's action and a certain sound event in the environment at that moment, and a greater likelihood of causal or intentional association.

[0074] This embodiment queries a preset first mapping rule based on the current interaction goal (temporarily enhanced environmental awareness). This preset first mapping rule can be a mapping rule table, which defines the feature filtering modes corresponding to different interaction goals. For example, the temporary enhanced environmental awareness goal might correspond to a sudden risk sound source detection mode, which indicates that the focus should be on moments when user actions are highly correlated with sudden environmental sounds.

[0075] This embodiment intelligently divides the time-series attention sequence into time segments based on the aforementioned feature filtering mode (sudden risk sound source detection mode). For example, in the sudden risk sound source detection mode, the intervals near which attention values ​​exceed a certain threshold (such as 0.8) are divided to obtain multiple time segments. For each segment, the changing pattern of attention values ​​within it is analyzed (e.g., a sharp rise followed by a slow decline, a sustained high-level plateau, etc.) to determine the pattern type of that time segment. In this embodiment, the pattern types may include strongly correlated initial peaks, sustained correlated plateaus, and weakly correlated troughs.

[0076] This embodiment does not mechanically select all time segments corresponding to high-interest values, but rather dynamically determines the selection criteria for key time segments based on three factors: mode type, environmental complexity, and the current working mode. The specific determination steps are as follows: Regarding the selection of mode types, this embodiment prioritizes the mode type that best matches the current interaction target. For example, the mode type that matches the target of temporarily enhancing environmental perception is a strong correlation initiation peak. In this embodiment, the strong correlation initiation peak typically corresponds to the immediate response of an action to sound. The environmental complexity in this embodiment is derived from the environmental interaction context vector, used to quantify the degree of disorder in the current acoustic environment (e.g., background noise level, number of concurrent sound sources). The correlation strength is adjusted based on the determined mode type and the current environmental complexity, and the priority of time segments corresponding to different mode types is determined in conjunction with the current working mode. Dynamic selection conditions are determined based on this priority.

[0077] This embodiment can also dynamically adjust the association strength based on environmental complexity when the mode type of the interaction target is determined. When the environmental complexity is higher than a preset complexity threshold, a higher association strength is required to eliminate interference. For example, in a quiet environment, an association strength exceeding 0.5 may be considered significant; however, in a noisy environment, the screening criterion for association strength will be raised to 0.8.

[0078] This embodiment determines the dynamic selection criteria based on the three factors mentioned above, and selects key time segments that meet the dynamic selection criteria from all the divided time segments. The selected key time segments represent the cross-modal interaction moments that are most relevant to the current interaction target and have the highest confidence. Based on the time index corresponding to the selected key time segments, this embodiment extracts feature data at the corresponding time points from the complete head motion feature sequence and environmental audio feature sequence, forming head motion feature subsets and environmental audio feature subsets, respectively.

[0079] This embodiment extracts the temporal attention sequence from the collaborative pattern features and uses the attention value to characterize the correlation strength between head movement and ambient audio at the corresponding time point, clearly presenting the degree of correlation between the two at different times in a quantitative manner. Based on the interaction goal, this embodiment selects corresponding feature subsets from head movement features and ambient audio features, avoiding information redundancy and interference caused by using all features, improving the effectiveness and specificity of the features, and making subsequent analysis and decision-making more efficient and accurate.

[0080] S106: Based on the interaction target, the feature subset is fused to generate fused features.

[0081] In this embodiment, the feature subset includes a head motion feature subset and an environmental motion feature subset.

[0082] In this embodiment, the interaction goal is to temporarily enhance environmental perception. The feature subset is obtained by filtering from the complete head motion feature sequence and the environmental audio feature sequence based on key time segments. The head motion feature subset includes features such as the precise angle and peak angular velocity within the key time segment (e.g., the start and end time of two head turning actions); the environmental audio feature subset includes features such as the location and sound pressure level of specific acoustic events (e.g., horns, bells) within the same time period.

[0083] This embodiment generates a structured fusion feature unit for each cross-modal event pair based on the completed pairing (pairing of head movement with ambient audio). Each fusion feature unit may include temporal relationship features, spatial relationship features, intensity relationship features, and information on the event semantic label dimension. The temporal relationship features record the delay time of the action relative to the sound, the duration of the action, etc. The spatial relationship features record the angular difference between the direction of the action and the direction of the sound source, the location of the sound source, etc. The intensity relationship features record the peak angular velocity of the action and the peak sound pressure level of the sound event. The event semantic label is used to associate the acoustic event category label (e.g., car horn, bicycle bell).

[0084] In this embodiment, all identified fusion feature units (left turn - left horn unit, right turn - right bell unit) are organized in chronological order to form fusion features.

[0085] S107: Based on the fusion features, perform dynamic policy prediction to obtain dynamic interaction policy, and adjust the audio output and / or interaction mode of the smart audio glasses according to the dynamic interaction policy.

[0086] This embodiment analyzes the user's dual-sided exploration pattern (the user actively and directionally explores sudden sound sources from two different spatial directions, left and right, in a short period of time) represented by the fusion features, determines the dynamic interaction strategy that adapts to the pattern, and adjusts the audio output and / or interaction mode of the smart audio glasses based on the dynamic interaction strategy.

[0087] In one embodiment of this application, before performing dynamic policy prediction based on fused features to obtain a dynamic interaction policy, the method further includes: Based on the interaction objectives, determine the type of interaction objectives, and based on the type, determine the core strategy dimensions that need to be considered for strategy decision-making; the core strategy dimensions include the audio response dimension used to decide audio output parameters and the interaction logic dimension used to decide interactive control behavior. Among them, dynamic interaction strategies are obtained by predicting dynamic strategies based on fused features, including: Based on the fusion features, a subset of environmental audio features, and potential user intent, an audio output strategy is generated in the audio response dimension; the audio output strategy includes at least one of the following: audio content type, playback intensity, and spatial orientation. Based on the fusion features, a subset of head motion features, and the interaction target, an interaction control strategy is generated under the dimension of interaction logic. The interaction control strategy includes at least one of the following: the responsiveness to subsequent user actions, the activation state of the interactive interface, and the trigger command to external devices. The audio output strategy and the interactive control strategy are combined to generate multiple candidate interaction strategies; Each candidate interaction strategy is evaluated for its suitability with the environment interaction context vector to obtain the suitability score for each candidate interaction strategy. The candidate interaction strategy corresponding to the highest fit score is used as the dynamic interaction strategy.

[0088] This embodiment first determines the task type as environmental awareness and safety response based on the current interaction goal (temporarily enhancing environmental awareness). Based on this task type, the strategy decision needs to focus on two core strategy dimensions: audio response and interaction logic. The audio response dimension focuses on how to adjust the sound output to meet the user's auditory needs. Its decision object is the playback attributes of the audio signal. The interaction logic dimension focuses on how to change the device's own response rules and states to adapt to the user's behavior patterns and task requirements. Its decision object is the device's functional logic and behavior.

[0089] This embodiment derives a strategy based on fusion features, a subset of environmental audio features, and potential user intent, and generates an audio output strategy in the audio response dimension.

[0090] The method for generating an audio output strategy may include: analyzing fusion features to confirm that the user needs to clearly perceive sudden sound sources on the left and right sides; combining the direction and intensity information of the sound sources in the audio feature subset, as well as a preset strategy, to generate a specific audio output strategy. This audio output strategy may include at least one of the following: audio content type, playback intensity, and spatial orientation. Specifically, the audio content type represents pausing the currently playing music and switching to an ambient sound enhancement mode; the playback intensity represents globally reducing any audio background noise that might mask ambient sounds and specifically enhancing the intensity of ambient sounds falling within the horn and bell frequency bands; and the spatial orientation is used to dynamically adjust the gain balance of the binaural audio channels based on the sound source orientation (left and right), making the horn sound slightly more prominent in the left ear and the bell sound slightly more prominent in the right ear, thereby strengthening the spatial directivity of the sound and assisting the user in locating it.

[0091] This embodiment derives an interaction control strategy based on fusion features, a subset of head motion features, and the interaction target, within the dimension of interaction logic.

[0092] The method for generating the interaction control strategy may include: inferring that the user is currently in a high-alert state by analyzing the user's rapid and continuous probing actions in the fused features, and generating an interaction control strategy in conjunction with the interaction goal. This interaction control strategy may include at least one of the following: responsiveness to subsequent user actions, activation state of the interactive interface, and triggering commands to external devices. Specifically, the responsiveness to subsequent user actions may reflect a temporary increase in the detection sensitivity and response speed to similar rapid head-turning actions, so that the device can respond more quickly when the user probes again; the activation state of the interactive interface may activate a simple voice assistant standby state, allowing the user to quickly obtain semantic broadcasts of ambient sounds through quick commands (e.g., what sound is there) without waking up the full voice assistant; the triggering commands to external devices are used to send notifications and warnings to the paired smartphone when necessary (e.g., when the risk is judged to be extremely high).

[0093] This embodiment combines audio output strategies with interactive control strategies to form multiple complete, multi-layered candidate interaction strategies. For each candidate interaction strategy, the strategy is compared with the environmental interaction context vector to evaluate its matching degree with the current overall scenario and obtain a corresponding fit score.

[0094] In this embodiment, the adaptation scores of all candidate interaction strategies are sorted in descending order, and the candidate interaction strategy with the highest adaptation score is determined as the final dynamic interaction strategy to be executed.

[0095] As can be seen from the above, this embodiment of the application performs joint analysis of head motion features and environmental audio features based on a cross-modal association network to determine cross-modal association relationships, quantify the dynamic association between head motion and environmental audio, and thus determine potential user intentions. It also generates an environmental interaction context vector by combining the encoding vector of the current working mode of the smart audio glasses. This embodiment comprehensively considers the association of multiple data sources and the working mode of the smart audio glasses, enabling a more comprehensive and accurate representation of the current contextual semantics that integrates user action intentions and environmental states. This embodiment also filters out subsets of head motion features and environmental audio features corresponding to key time segments by determining dynamic selection conditions for key time segments. This dynamic filtering method can accurately select features highly correlated with the interaction based on different interaction goals and environmental audio conditions, avoiding interference from irrelevant features, improving feature utilization efficiency, and making decision-making more accurate.

[0096] This embodiment also generates an audio output strategy in the audio response dimension and an interaction control strategy in the interaction logic dimension. The two generated strategies are then combined to generate multiple candidate interaction strategies. Adaptability is evaluated through the environmental interaction context vector to select the final dynamic interaction strategy. This embodiment combines multiple dimensions and adaptability evaluation methods to generate the most suitable dynamic interaction strategy for the current scenario based on fusion features, potential user intent, and interaction goals, providing users with a personalized interaction experience.

[0097] Based on the same inventive concept, this application also provides an adaptive interaction device for smart audio glasses to implement the aforementioned adaptive interaction method. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the adaptive interaction device for smart audio glasses provided below can be found in the limitations of the adaptive interaction method for smart audio glasses described above, and will not be repeated here.

[0098] This application provides an adaptive interaction device for smart audio glasses, such as... Figure 2 As shown, the smart audio glasses adaptive interaction device 20 includes: a multi-source data processing module 21, a feature extraction module 22, an environmental interaction context vector determination module 23, an interaction target determination module 24, a feature subset determination module 25, a fusion feature acquisition module 26, and a dynamic interaction strategy determination module 27.

[0099] Multi-source data processing module 21 is used to acquire multi-source data, perform time synchronization processing on the multi-source data, and obtain processed multi-source data; the multi-source data includes head motion data and environmental audio data; Feature extraction module 22 is used to extract multimodal features from the processed multi-source data to obtain head motion features and environmental audio features respectively; The environmental interaction context vector determination module 23 is used to jointly analyze head motion features and environmental audio features to determine cross-modal correlations, and generate environmental interaction context vectors in combination with the current working mode of smart audio glasses. The environmental interaction context vectors represent the current contextual semantics that integrate user action intentions and environmental states. The interaction target determination module 24 is used to determine the corresponding interaction target based on the environmental interaction context vector; the interaction target represents the task category that needs to be performed in the current context semantics. The feature subset determination module 25 is used to filter the feature subset corresponding to the interaction target from head motion features and environmental audio features based on the interaction target; The feature fusion acquisition module 26 is used to fuse feature subsets based on the interaction target to generate fused features; The dynamic interaction strategy determination module 27 is used to predict the dynamic strategy based on the fusion features, obtain the dynamic interaction strategy, and adjust the audio output and / or interaction mode of the smart audio glasses according to the dynamic interaction strategy.

[0100] In one embodiment of this application, the environmental interaction context vector determination module 23, when jointly analyzing head motion features and environmental audio features to determine cross-modal correlations and generating an environmental interaction context vector in conjunction with the current working mode of the smart audio glasses, is specifically used for: Cross-modal correlation networks are used to jointly analyze head motion features and environmental audio features to obtain cross-modal correlation relationships; these relationships are then used to quantify the dynamic correlation between head motion and environmental audio. Based on cross-modal associations, determine potential user intent; Encode the current working mode to obtain the encoding vector of the current working mode; The potential user intent, head movement features, and environmental audio features are vectorized and concatenated, and then fused with the encoded vector to generate an environmental interaction context vector.

[0101] In one embodiment of this application, the environmental interaction context vector determination module 23, when determining potential user intent based on cross-modal association relationships, is specifically used for: Cross-modal correlations are analyzed to extract co-modal features that characterize the temporal and intensity-related changes of head motion and ambient audio. These co-modal features include one or more combinations of temporal following tightness, intensity change coupling, and event co-occurrence frequency between head motion and ambient audio. Based on the characteristics of the collaborative pattern, multiple candidate basic intentions are identified; Based on the current working mode of the smart audio glasses and the candidate intent priority strategy, multiple candidate basic intents are prioritized, and the candidate basic intent with the highest priority is taken as the potential user intent.

[0102] In one embodiment of this application, the feature subset determination module 25, when selecting a feature subset corresponding to the interaction target from head motion features and environmental audio features based on the interaction target, is specifically used for: The temporal attention sequence is extracted from the collaborative pattern features. Each attention value in the temporal attention sequence is used to characterize the correlation strength between head movement and ambient audio at the corresponding time point. Based on the interaction target and the preset first mapping rule, the feature filtering mode corresponding to the current interaction target is determined; the preset first mapping rule is used to characterize the mapping between the interaction target and the feature filtering mode. Based on the feature selection mode, the time-series attention sequence is divided into time segments to obtain the divided time segments. Then, the attention values ​​within each time segment are analyzed to obtain the pattern type of the time segment. The pattern type is used to characterize the change pattern of the correlation between head movement and environmental audio within the time segment. Based on the mode type, the environmental complexity in the environmental interaction context vector, and the current working mode of the smart audio glasses, dynamic selection conditions for selecting key time segments are determined; environmental complexity is used to characterize the degree of disorder in the current environmental audio data. Based on dynamic selection criteria, key time segments are determined from the divided time segments; Based on key time segments, features corresponding to the time points are selected from head motion features as a subset of head motion features, and features corresponding to the time points are selected from environmental audio features as a subset of environmental audio features.

[0103] In one embodiment of this application, the interaction target determination module 24, when determining the corresponding interaction target based on the environmental interaction context vector, is specifically used for: Obtain the scene vector set corresponding to the smart audio glasses in different working modes, where each scene vector in the scene vector set corresponds to a candidate interaction target; Determine the corresponding set of scenario vectors based on the current working mode of the smart audio glasses; Obtain the similarity between the environmental interaction context vector and each context vector in the context vector set; The maximum similarity among all similarities is selected. If the maximum similarity exceeds the preset scenario matching threshold, the candidate interaction target associated with the scenario vector corresponding to the maximum similarity is taken as the interaction target.

[0104] In one embodiment of this application, the dynamic interaction strategy determination module 27, when performing dynamic strategy prediction based on fusion features to obtain a dynamic interaction strategy, is specifically used for: Based on the interaction objectives, determine the type of interaction objectives, and based on the type, determine the core strategy dimensions that need to be considered for strategy decision-making; the core strategy dimensions include the audio response dimension used to decide audio output parameters and the interaction logic dimension used to decide interactive control behavior. Based on the fusion features, a subset of environmental audio features, and potential user intent, an audio output strategy is generated in the audio response dimension; the audio output strategy includes at least one of the following: audio content type, playback intensity, and spatial orientation. Based on the fusion features, a subset of head motion features, and the interaction target, an interaction control strategy is generated under the dimension of interaction logic. The interaction control strategy includes at least one of the following: the responsiveness to subsequent user actions, the activation state of the interactive interface, and the trigger command to external devices. The audio output strategy and the interactive control strategy are combined to generate multiple candidate interaction strategies; Each candidate interaction strategy is evaluated for its suitability with the environment interaction context vector to obtain the suitability score for each candidate interaction strategy. The candidate interaction strategy corresponding to the highest fit score is used as the dynamic interaction strategy.

[0105] In one embodiment of this application, the cross-modal association network includes a first attention subnetwork and a second attention subnetwork. The environment interaction context vector determination module 23, when performing joint analysis of head motion features and environmental audio features based on the cross-modal association network to obtain cross-modal association relationships, is specifically used for: Temporally aligned encoding is performed on head motion features and ambient audio features to obtain temporally aligned head motion codes and ambient audio codes; Head motion encoding is used as the query vector, and ambient audio encoding is used as the key vector and value vector. The query vector, key vector, and value vector are input into the first attention sub-network to obtain the first cross-modal attention map. The first cross-modal attention map is used to characterize the semantic contribution of ambient audio information to head motion at each moment. The ambient audio encoding is used as the query vector, and the head motion encoding is used as the key vector and value vector. The query vector, key vector and value vector are input into the second attention sub-network to obtain the second cross-modal attention map. The second cross-modal attention map is used to characterize the semantic contribution of head motion information to the ambient audio at each time step. The first cross-modal attention map and the second cross-modal attention map are fused to generate a cross-modal attention weight matrix; The head motion encoding and environmental audio encoding are weighted and fused based on the cross-modal attention weight matrix to obtain the cross-modal joint feature vector; Cross-modal associations are obtained based on cross-modal joint feature vectors.

[0106] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. The electronic device may be smart audio glasses. Figure 3The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of each module / unit in the above-described device embodiments, for example... Figure 2 The functions of the multi-source data processing module 21, feature extraction module 22, environment interaction context vector determination module 23, interaction target determination module 24, feature subset determination module 25, fusion feature acquisition module 26, and dynamic interaction strategy determination module 27 are shown.

[0107] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0108] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0109] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store information such as head motion features, environmental audio features, environmental interaction context vectors, fused features, and dynamic interaction strategies.

[0110] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation method described in the intelligent audio glasses adaptive interaction method provided in the embodiments of this application, or they can execute the implementation method of the electronic device described in the embodiments of this application, which will not be repeated here.

[0111] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0112] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., provided on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0113] Those skilled in the art will recognize that the modules / units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0114] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0115] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules, units, or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or modules / units, or it may be an electrical, mechanical, or other form of connection.

[0116] The modules / units described as separate components may or may not be physically separate. Similarly, the components shown as modules / units may or may not be physical modules / units; they may be located in one place or distributed across multiple network modules / units. Some or all of the modules / units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0117] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The integrated modules / units described above can be implemented in hardware or in the form of software functional modules / units.

[0118] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An adaptive interaction method for smart audio glasses, characterized in that, include: Acquire multi-source data, perform time synchronization processing on the multi-source data, and obtain processed multi-source data; the multi-source data includes head motion data and environmental audio data. Multimodal feature extraction is performed on the processed multi-source data to obtain head motion features and environmental audio features, respectively. The head motion features and the environmental audio features are jointly analyzed to determine cross-modal correlations. Combined with the current working mode of the smart audio glasses, an environmental interaction context vector is generated. The environmental interaction context vector represents the current contextual semantics that integrates the user's action intention and the environmental state. The corresponding interaction target is determined based on the environmental interaction context vector; the interaction target represents the task category that needs to be performed in the current context semantics. Based on the interaction target, a subset of features corresponding to the interaction target is selected from the head motion features and the environmental audio features; The feature subset is fused based on the interaction target to generate a fused feature; Based on the fusion features, dynamic strategy prediction is performed to obtain a dynamic interaction strategy, and the audio output and / or interaction mode of the smart audio glasses are adjusted according to the dynamic interaction strategy.

2. The adaptive interaction method for smart audio glasses as described in claim 1, characterized in that, The joint analysis of the head motion features and the environmental audio features determines cross-modal correlations, and, combined with the current operating mode of the smart audio glasses, generates an environmental interaction context vector, including: The head motion features and the ambient audio features are jointly analyzed based on a cross-modal association network to obtain cross-modal association relationships; these cross-modal association relationships are used to quantify the dynamic association between head motion and ambient audio. Based on the cross-modal associations, potential user intent is determined; The current working mode is encoded to obtain the encoding vector of the current working mode; The potential user intent, the head movement features, and the environmental audio features are vectorized and concatenated, and then fused with the encoded vector to generate the environmental interaction context vector.

3. The adaptive interaction method for smart audio glasses as described in claim 2, characterized in that, Determining potential user intent based on the cross-modal association includes: The cross-modal correlation is analyzed to extract co-modal features that characterize the temporal and intensity-related changes of head movement and ambient audio. The co-modal features include one or more combinations of the temporal following tightness of head movement and ambient audio, the coupling degree of intensity change, and the event co-occurrence frequency. Based on the characteristics of the collaborative mode, multiple candidate basic intentions are determined; Based on the current working mode of the smart audio glasses and the candidate intent priority strategy, the multiple candidate basic intents are prioritized, and the candidate basic intent with the highest priority is taken as the potential user intent.

4. The adaptive interaction method for smart audio glasses as described in claim 3, characterized in that, Before selecting a subset of features corresponding to the interaction target from the head motion features and the ambient audio features based on the interaction target, the method further includes: The temporal attention sequence is parsed from the collaborative mode features. Each attention value in the temporal attention sequence is used to characterize the correlation strength between head movement and ambient audio at the corresponding time point. The step of selecting a subset of features corresponding to the interaction target from the head motion features and the environmental audio features, based on the interaction target, includes: Based on the interaction target and a preset first mapping rule, a feature filtering mode corresponding to the current interaction target is determined; the preset first mapping rule is used to characterize the mapping between the interaction target and the feature filtering mode. Based on the feature filtering mode, the time-series attention sequence is divided into time segments to obtain the divided time segments. Pattern analysis is then performed on the attention values ​​within each divided time segment to obtain the pattern type of the time segment. The pattern type is used to characterize the change pattern of the correlation strength between head movement and environmental audio within the time segment. Based on the mode type, the environmental complexity in the environmental interaction context vector, and the current working mode of the smart audio glasses, dynamic selection conditions for selecting key time segments are determined; the environmental complexity is used to characterize the degree of disorder in the current environmental audio data. Based on the dynamic selection criteria, key time segments are determined from the divided time segments; Based on the key time segments, features corresponding to the time points are selected from the head motion features as a subset of the head motion features, and features corresponding to the time points are selected from the environmental audio features as a subset of the environmental audio features.

5. The adaptive interaction method for smart audio glasses as described in claim 1, characterized in that, Before determining the corresponding interaction target based on the environmental interaction context vector, the method further includes: Obtain the scene vector set corresponding to the smart audio glasses in different working modes, wherein each scene vector in the scene vector set corresponds to a candidate interaction target; The step of determining the corresponding interaction target based on the environmental interaction context vector includes: Determine the corresponding set of scenario vectors based on the current working mode of the smart audio glasses; Obtain the similarity between the environmental interaction context vector and each scenario vector in the scenario vector set; The maximum similarity among all similarities is selected. If the maximum similarity exceeds a preset scenario matching threshold, the candidate interaction target associated with the scenario vector corresponding to the maximum similarity is taken as the interaction target.

6. The adaptive interaction method for smart audio glasses as described in claim 4, characterized in that, Before performing dynamic policy prediction based on the fused features to obtain the dynamic interaction policy, the method further includes: The type of the interaction objective is determined based on the interaction objective, and the core strategy dimensions to be considered for strategy decision-making are determined based on the type; the core strategy dimensions include the audio response dimension for deciding audio output parameters and the interaction logic dimension for deciding interactive control behavior. The step of predicting dynamic strategies based on the fused features to obtain dynamic interaction strategies includes: Based on the fusion features, the subset of environmental audio features, and the potential user intent, an audio output strategy is generated under the audio response dimension; the audio output strategy includes at least one of the following: audio content type, playback intensity, and spatial orientation. Based on the fusion features, the subset of head motion features, and the interaction target, an interaction control strategy is generated under the interaction logic dimension; the interaction control strategy includes at least one of the following: response sensitivity to subsequent user actions, activation state of the interactive interface, and trigger command to external device. The audio output strategy and the interaction control strategy are combined to generate multiple candidate interaction strategies; Each candidate interaction strategy is evaluated for its suitability with the environmental interaction context vector to obtain the suitability score for each candidate interaction strategy. The candidate interaction strategy corresponding to the highest fit score is used as the dynamic interaction strategy.

7. The adaptive interaction method for smart audio glasses as described in claim 2, characterized in that, The cross-modal association network includes a first attention subnetwork and a second attention subnetwork. The joint analysis of the head motion features and the environmental audio features based on the cross-modal association network to obtain cross-modal association relationships includes: The head motion features and the ambient audio features are temporally aligned and encoded to obtain temporally aligned head motion codes and ambient audio codes. The head motion encoding is used as a query vector, and the environmental audio encoding is used as a key vector and a value vector. The query vector, the key vector, and the value vector are input into the first attention subnetwork to obtain a first cross-modal attention map. The first cross-modal attention map is used to characterize the semantic contribution of environmental audio information to head motion at each moment. The environmental audio encoding is used as a query vector, and the head motion encoding is used as a key vector and a value vector. The query vector, the key vector, and the value vector are input into the second attention subnetwork to obtain a second cross-modal attention map. The second cross-modal attention map is used to characterize the semantic contribution of head motion information to the environmental audio at each moment. The first cross-modal attention map and the second cross-modal attention map are fused to generate a cross-modal attention weight matrix; The head motion encoding and the environmental audio encoding are weighted and fused according to the cross-modal attention weight matrix to obtain a cross-modal joint feature vector; The cross-modal association relationship is obtained based on the cross-modal joint feature vector.

8. An adaptive interaction device for smart audio glasses, characterized in that, include: A multi-source data processing module is used to acquire multi-source data, perform time synchronization processing on the multi-source data, and obtain processed multi-source data; the multi-source data includes head motion data and environmental audio data. The feature extraction module is used to extract multimodal features from the processed multi-source data to obtain head motion features and environmental audio features, respectively. The environmental interaction context vector determination module is used to jointly analyze the head motion features and the environmental audio features to determine the cross-modal correlation, and generate an environmental interaction context vector in combination with the current working mode of the smart audio glasses. The environmental interaction context vector represents the current context semantics that integrates the user's action intention and the environmental state. An interaction target determination module is used to determine the corresponding interaction target based on the environmental interaction context vector; the interaction target represents the task category that needs to be performed in the current context semantics. The feature subset determination module is used to filter a feature subset corresponding to the interaction target from the head motion features and the environmental audio features, based on the interaction target. The feature fusion acquisition module is used to fuse the feature subset based on the interaction target to generate fused features; The dynamic interaction strategy determination module is used to predict the dynamic strategy based on the fusion features, obtain the dynamic interaction strategy, and adjust the audio output and / or interaction mode of the smart audio glasses according to the dynamic interaction strategy.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.