A voice directional transmission method and system

By using the microphone array and multi-dimensional decision-making model of smart glasses to dynamically adjust the encoding strategy, the problems of sound source direction change and dialogue rhythm fluctuation during the wearing of smart glasses are solved, achieving high-quality voice transmission and stable communication experience.

CN122637795APending Publication Date: 2026-08-25SHENZHEN ZHILIANMAO TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610859808.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing smart glasses suffer from changes in sound source direction due to head movements during wear, affecting voice signal quality. They also cannot adapt to the rhythm of real-time conversations and fluctuations in wireless communication links, resulting in unstable voice transmission quality.

Method used

By using a microphone ring array and beamforming technology to directionally acquire sound sources, combined with speech activity detection and a multi-dimensional decision model, the coding mode and preprocessing parameters are dynamically adjusted to adapt to changes in sound source direction and dialogue rhythm, and to predict link quality fluctuations, thus achieving adaptive voice transmission.

Benefits of technology

It improves the quality of voice acquisition and recognition accuracy, reduces latency and packet loss, and enhances the continuity and smoothness of communication, making it suitable for resource-constrained wearable devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122637795A_ABST
    Figure CN122637795A_ABST
Patent Text Reader

Abstract

The application provides a voice directional transmission method and system, and relates to the technical field of communication. The method comprises the following steps: collecting audio data of a sound source in a directional manner, generating sound source direction information and calculating the change rate thereof; performing voice activity detection on the audio, counting the voice frame density and the dialogue round change rate, and generating user communication state data; obtaining real-time quality parameters of a communication link and predicting a short-time prediction value of the link quality; constructing a multi-dimensional decision model, taking the short-time prediction value of the link, the real-time quality parameters, the user communication state data and the sound source direction change rate as inputs, and dynamically calculating a mixing coefficient; dynamically generating a target coding mode and its pre-processing parameters from at least two coding modes according to the mixing coefficient; and sending the coding result obtained according to the target parameters to a mobile terminal through a low-power channel. The voice acquisition and transmission are adaptively and cooperatively optimized according to the sound source dynamics, the dialogue rhythm and the link fluctuation, and the communication quality and user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and in particular to a method and system for directional voice transmission. Background Technology

[0002] In cross-language communication or real-time voice communication scenarios, smart glasses, as a wearable device, have attracted attention due to their ability to provide near-eye display and convenient voice interaction functions. In existing technologies, smart glasses are typically used in conjunction with mobile terminals: the smart glasses are responsible for collecting the user's voice audio data and sending it to the mobile terminal via a low-power transmission channel (such as Bluetooth). The mobile terminal then performs speech recognition, translation, or other voice processing, and the processing results are sent back to the smart glasses for display or playback.

[0003] However, the above-mentioned existing solutions have the following technical problems in practical applications: First, smart glasses are worn on the user's head. During speaking or listening, head movements cause the direction of the sound source relative to the glasses' microphone array to change in real time. Existing audio acquisition strategies for smart glasses are typically fixed-point or rely solely on initial sound source localization, which cannot adapt to such dynamic changes. When the user turns their head, moves around, or interacts with multiple speakers, the fixed-point acquisition method causes the target sound source to deviate from the main lobe of the acquisition, resulting in a decrease in the quality of the acquired speech signal, such as volume attenuation and a reduced signal-to-noise ratio, which in turn affects the accuracy of backend speech recognition.

[0004] Second, in real-time dialogue scenarios, the pace and intensity of interaction are dynamically changing. For example, in scenarios involving rapid question-and-answer sessions, debates, or multiple speakers taking turns, the real-time requirements for voice transmission are extremely high; while during monologues, statements, or silent reflection, the requirements for voice integrity and anti-interference capabilities are even higher. Existing technologies typically employ fixed encoding parameters (such as fixed frame length and compression rate) or only perform simple mode switching when link quality deteriorates (such as choosing between real-time mode and buffered mode). This discrete, reactive adjustment method is difficult to smoothly adapt to the continuous changes in the pace of dialogue, and is prone to performance jitter caused by frequent mode switching, or problems such as voice stuttering and sudden increases in latency occurring before mode switching.

[0005] Third, the wireless communication link (such as Bluetooth) between smart glasses and mobile terminals is susceptible to environmental interference, distance changes, and human occlusion, leading to non-stationary fluctuations in link quality (such as latency, packet loss rate, and jitter). Existing technologies typically make passive adjustments based solely on the link quality parameters at the current moment, such as switching encoding strategies or triggering retransmissions only after detecting a packet loss rate exceeding a threshold. This reactive mechanism is lagging; by the time the link quality deteriorates, voice frame loss or timeouts may have already occurred, affecting the continuity and smoothness of communication. Furthermore, a single link quality parameter (such as relying solely on the packet loss rate) is insufficient to comprehensively assess the changing trends of the link state, resulting in inaccurate adjustment decisions. Summary of the Invention

[0006] To address the aforementioned issues, this invention proposes a voice directional transmission method and system. By dynamically calculating the mixing coefficients through a multi-dimensional decision model and continuously adjusting the coding mode and preprocessing parameters, the method achieves adaptive and collaborative optimization of voice acquisition and transmission in response to sound source dynamics, dialogue rhythm, and link fluctuations, thereby improving communication quality and user experience.

[0007] The objective of this invention is achieved through the following technical solution: In a first aspect, embodiments of the present invention provide a voice-oriented transmission method, applied to smart glasses that are communicatively connected to a mobile terminal, the method comprising: The built-in audio acquisition unit collects audio data of the first sound source in a directional manner and generates corresponding sound source direction information; the rate of change of the sound source direction information is calculated based on the sound source direction information sequence. The audio data is subjected to voice activity detection, and the voice frame density and dialogue turn change rate within a preset time window are statistically analyzed to generate user communication status data representing real-time interaction intentions. Obtain the real-time quality parameter sequence of the communication link between the smart glasses and the mobile terminal, and predict the short-time predicted value of the link quality based on the real-time quality parameter sequence; A multidimensional decision model is constructed, taking the short-term predicted value of the link quality, the real-time quality parameters, the user communication status data, and the rate of change of the sound source direction information as input parameters, and dynamically calculating the mixing coefficient of the target coding mode, wherein the mixing coefficient is a continuous value between 0 and 1; Based on the mixing coefficients, a target encoding mode and its corresponding preprocessing parameters are dynamically generated from at least two preset encoding modes, wherein the at least two preset encoding modes include a real-time encoding mode and a buffered encoding mode; and the preprocessing parameters include noise reduction intensity and echo cancellation depth. The audio data is encoded according to the encoding parameters corresponding to the target encoding mode and the preprocessing parameters to obtain encoded audio data; The encoded audio data is sent to the mobile terminal via a low-power transmission channel.

[0008] On the other hand, embodiments of the present invention provide a voice-oriented transmission system, including: Smart glasses, and a mobile terminal that is communicatively connected to the smart glasses; The smart glasses include: An audio acquisition unit is used to directionally acquire audio data from a sound source; A display unit for projecting an image in front of the user; One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the voice directional transmission method described in the embodiments of the present invention.

[0009] The beneficial effects of this invention include: directional acquisition of audio data from a sound source via a built-in microphone ring array, and real-time generation of sound source direction information and its rate of change. When the rate of change of the sound source direction information exceeds a preset threshold, it is determined that the user is in a dynamic interactive environment (such as head rotation, walking, or multiple speakers alternating), and the encoding mode and preprocessing parameters can be adjusted accordingly (such as increasing noise reduction intensity and echo cancellation depth, and smoothly transitioning to a buffered encoding mode). Compared to existing acquisition strategies that rely on fixed direction or only initial sound source localization, this invention can dynamically track changes in sound source direction, effectively avoiding volume attenuation and signal-to-noise ratio reduction caused by the target sound source deviating from the acquisition main lobe, thereby significantly improving the quality of speech acquisition and the accuracy of backend speech recognition.

[0010] By detecting speech activity in audio data and statistically analyzing the speech frame density and dialogue turn change rate within a preset time window, user communication state data representing real-time interaction intentions is generated. Based on this, a multi-dimensional decision model is used to calculate continuous mixing coefficients between 0 and 1. These mixing coefficients are then used to perform linear interpolation or weighted fusion of encoding parameters for at least two preset encoding modes and preprocessing parameters. This allows for real-time and smooth adjustment of the encoding strategy according to continuous changes in dialogue rhythm (such as alternation between rapid question-and-answer and silent thinking), avoiding performance jitter and stuttering latency caused by discrete switching, and achieving a dynamic optimal balance between real-time performance and anti-interference capability.

[0011] By acquiring the real-time quality parameter sequence of the communication link between smart glasses and mobile terminals, and using Kalman filters or linear regression models to make short-term predictions of link quality, the system can proactively adapt to link fluctuations and pre-adjust the coding strategy before link quality deterioration occurs. This effectively reduces voice frame loss and latency timeout events, and improves the continuity and fluency of communication.

[0012] By inputting information from multiple dimensions, including the rate of change of sound source direction information, user communication status data, real-time quality parameters, and short-term predicted link quality values, into a multi-dimensional decision model, the mixing coefficients are dynamically calculated, and encoding and preprocessing parameters are adjusted collaboratively. This avoids performance imbalances caused by a single factor dominating the decision-making process, thereby achieving global collaborative optimization of audio acquisition, encoding, and transmission in complex dynamic interaction scenarios, significantly improving the user's overall communication experience.

[0013] Employing a rule-based explicit multidimensional decision-making model, it eliminates the need for extensive data pre-training or complex neural network inference, completing decisions through low-complexity operations such as threshold comparison, weighted correction, and linear interpolation. The entire decision-making process has extremely low computational load and power consumption, making it suitable for deployment in resource-constrained wearable devices such as smart glasses. Combined with low-power transmission channels (such as Bluetooth Low Energy), it can ensure high-quality voice transmission while maintaining device battery life. Attached Figure Description

[0014] The invention will now be further described with reference to the accompanying drawings.

[0015] Figure 1 This is a flowchart illustrating a voice directional transmission method according to the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1, see Figure 1 This invention provides a voice-directed transmission method applied to smart glasses that communicate with a mobile terminal. The method includes: The built-in audio acquisition unit collects audio data of the first sound source in a directional manner and generates corresponding sound source direction information; the rate of change of the sound source direction information is calculated based on the sound source direction information sequence. The audio data is subjected to voice activity detection, and the voice frame density and dialogue turn change rate within a preset time window are statistically analyzed to generate user communication status data representing real-time interaction intentions. Obtain the real-time quality parameter sequence of the communication link between the smart glasses and the mobile terminal, and predict the short-time predicted value of the link quality based on the real-time quality parameter sequence; A multi-dimensional decision-making model is constructed, using the short-term predicted value of the link quality, the real-time quality parameters, the user communication status data, and the rate of change of the sound source direction information as input parameters. The model dynamically calculates the mixing coefficients of the target coding mode, where each mixing coefficient is a continuous value between 0 and 1. The dynamic calculation includes: based on a preset benchmark value, generating direction adjustment amounts according to the rate of change of the sound source direction information, the user communication status data, the packet loss rate in the real-time quality parameters, and the degree to which the short-term predicted value of the link quality deviates from its corresponding threshold; accumulating all direction adjustment amounts with the benchmark value; and limiting the accumulation result to between 0 and 1 to obtain the mixing coefficients. Based on the mixing coefficients, a target encoding mode and its corresponding preprocessing parameters are dynamically generated from at least two preset encoding modes, wherein the at least two preset encoding modes include a real-time encoding mode and a buffered encoding mode; and the preprocessing parameters include noise reduction intensity and echo cancellation depth. The audio data is encoded according to the encoding parameters corresponding to the target encoding mode and the preprocessing parameters to obtain encoded audio data; The encoded audio data is sent to the mobile terminal via a low-power transmission channel.

[0018] The mobile terminal can be a smartphone, tablet, or other device with strong computing power. The smart glasses and the mobile terminal establish a wireless connection through a low-power transmission channel (such as Bluetooth Low Energy or Zigbee). The smart glasses are worn on the user's head to collect voice signals from the user's environment and send the processed audio data to the mobile terminal for further processing (such as voice recognition, translation, or forwarding to other terminals).

[0019] In a typical application scenario of this embodiment, User A wears smart glasses to engage in face-to-face cross-language communication with User B. User A needs to transmit User B's speech to a mobile terminal in real time for translation and display the translation result on the smart glasses' screen. During this process, User A may turn their head to observe the surrounding environment at any time, or may engage in conversation with multiple speakers in turn. Simultaneously, the wireless link between User A and the mobile terminal may be affected by human obstruction, changes in distance, or interference from other wireless signals, leading to fluctuations in communication quality.

[0020] Based on the above scenario, the voice directional transmission method in this embodiment includes the following steps.

[0021] Step 1: Acquire audio data in a directional manner and calculate the rate of change of sound source direction information.

[0022] Smart glasses use a built-in audio acquisition unit (such as a circular array of multiple microphones) to directionally acquire audio data from a first sound source. The first sound source typically refers to the speaker the smart glasses wearer currently wishes to hear, such as User B. The smart glasses utilize beamforming technology to calculate the sound source direction information of each sound signal based on the arrival time difference and phase difference of the sound signals received by each microphone. Based on this, they amplify the sound signal from the target direction and suppress noise interference from other directions.

[0023] During continuous data acquisition, the smart glasses record the sound source direction information corresponding to each frame of audio data, forming a sequence of sound source direction information. By analyzing the direction changes at adjacent time points in this sequence, the smart glasses can calculate the rate of change of the sound source direction information. This rate of change reflects the angular velocity of the target sound source relative to the smart glasses. For example, when user A quickly turns their head to look elsewhere, the direction of the target sound source (user B) relative to the smart glasses will change drastically, and the rate of change value will be high; when user A's head remains stable or moves slowly, the rate of change value will be low. This rate of change is one of the important input parameters for subsequent decision-making models, used to determine the dynamic level of the current interaction environment.

[0024] Step 2: Generate user communication state data that represents real-time interaction intent.

[0025] After collecting audio data, the smart glasses perform voice activity detection, that is, determine whether each frame of audio contains valid human voice, silence, or background noise. Within a preset time window (e.g., 2 seconds), the smart glasses count the voice frame density. The higher the voice frame density, the more continuous the user's speech or listening reception; the lower the density, the more pauses or silences there are.

[0026] Simultaneously, the smart glasses detect the rate of change in dialogue turns within this time window. A dialogue turn refers to the frequency of speaker alternation; for example, when user B finishes speaking and user A responds, one turn is completed. The smart glasses can count the number of turns occurring per unit of time by recognizing the start and end points of speech and speaker switching. The rate of change in dialogue turns reflects the pace of the conversation. For example, the rate of change in turns is high in fast-paced question-and-answer or debate scenarios, while it is low in monologues or lengthy statements.

[0027] Combining voice frame density and dialogue turn change rate creates user communication state data that characterizes real-time interaction intent. This data tells the system whether the current dialogue is tense and fast or relaxed and calm, and whether the user is passively listening or actively responding.

[0028] Step 3: Obtain real-time link quality parameters and predict short-term trends.

[0029] During communication with mobile terminals via low-power transmission channels, smart glasses can periodically receive transmission quality reports from the mobile terminals, or actively detect parameters such as round-trip time, jitter, and acknowledgment loss rate. These parameters constitute a real-time quality parameter sequence.

[0030] Smart glasses use predictive models (such as Kalman filters or linear regression models) to process the aforementioned real-time quality parameter sequences, predicting the trend of link quality changes over a future period (e.g., 0.5 to 1 second), thus obtaining a short-term predicted value for link quality. This predicted value can indicate whether the link quality will deteriorate, improve, or remain stable. For example, if the round-trip time gradually increases and the loss rate shows signs of rising over several consecutive detection cycles, the predictive model may determine that the link will deteriorate; conversely, if the indicators continue to improve, it is determined that the link will improve.

[0031] Step 4: Construct a multidimensional decision-making model and dynamically calculate the mixing coefficients.

[0032] The smart glasses have a built-in multi-dimensional decision model. The inputs of the model include: the rate of change of sound source direction information obtained in step one, the user communication status data obtained in step two (i.e., voice frame density and dialogue turn rate of change), the real-time quality parameters (current latency, packet loss rate, jitter) obtained in step three, and the short-term prediction value of link quality.

[0033] The multidimensional decision model comprehensively evaluates these inputs according to pre-defined logical rules, outputting a continuous value between 0 and 1, called the mixing coefficient. This coefficient reflects which end the system should lean towards between real-time encoding and buffered encoding. Generally speaking, when the mixing coefficient approaches 0, it indicates that the current scenario prioritizes low latency, making real-time encoding suitable; when the mixing coefficient approaches 1, it indicates that the current scenario prioritizes anti-interference and integrity, making buffered encoding suitable; a value in between represents a balance between the two.

[0034] For example, if the rate of change of sound source direction information is high (the user's head turns rapidly) and the rate of change of dialogue turns is also high (the question-and-answer pace is fast), even if the link quality is acceptable, the model may adjust the mixing coefficients towards 1. This is because stronger noise reduction and echo cancellation are needed to cope with the acquisition distortion caused by rapid changes in direction, prioritizing audio quality over minimum latency. Conversely, if the head is stable, the dialogue pace is slow, and the link quality is good, the model will adjust the mixing coefficients towards 0, pursuing real-time transmission with the lowest latency.

[0035] Step 5: Dynamically generate the target coding pattern and preprocessing parameters based on the mixing coefficients.

[0036] Smart glasses have at least two preset encoding modes: real-time encoding mode and cached encoding mode. Real-time encoding mode uses shorter frame lengths and lower compression rates to reduce encoding and transmission latency, but it is more sensitive to packet loss and noise. Cached encoding mode uses longer frame lengths and higher compression rates, and may enable additional redundant error correction mechanisms to enhance resilience against packet loss, but it introduces greater latency.

[0037] Based on the mixing coefficients calculated in step four, the smart glasses dynamically generate a target encoding mode. This target encoding mode is not a simple choice between two options, but rather an intermediate state under continuous adjustment of the mixing coefficients. Specifically, the smart glasses perform linear interpolation on the encoding parameters (such as frame length and compression ratio) of the real-time encoding mode and the corresponding parameters of the buffered encoding mode to obtain the encoding parameters of the target encoding mode. Simultaneously, the smart glasses also dynamically generate preprocessing parameters based on the mixing coefficients, including noise reduction intensity and echo cancellation depth. For example, when the mixing coefficients are large (biased towards the buffered mode), the noise reduction intensity is increased, and the echo cancellation depth is deepened to filter out more environmental noise and its own echo; when the mixing coefficients are small (biased towards the real-time mode), the noise reduction intensity is appropriately reduced to decrease processing latency.

[0038] Step 6: Encode the audio data according to the target parameters.

[0039] The smart glasses process the raw audio data according to the target encoding mode and its corresponding encoding parameters and preprocessing parameters obtained in step five. The processing flow is usually as follows: first, the audio signal is preprocessed by applying the specified noise reduction algorithm and echo cancellation algorithm; then, it is divided into frames according to the target frame length; finally, it is encoded and compressed using the target compression rate to obtain the encoded audio data.

[0040] Step 7: Send to the mobile terminal via a low-power transmission channel.

[0041] The smart glasses package the encoded audio data and send it to the mobile terminal with which they communicate via an established low-power transmission channel (such as BLE). After receiving the data, the mobile terminal can perform speech recognition, translation, or forwarding to the cloud or other users according to application requirements.

[0042] The effects of the above technical solution are as follows: The smart glasses can perceive the dynamic direction of the target sound source, the real-time changes in the rhythm of the conversation, and the quality fluctuations of the wireless link in real time. Based on a multi-dimensional decision model, it outputs a continuous mixing coefficient, thereby smoothly and continuously adjusting the encoding mode and preprocessing parameters. This allows the voice acquisition and transmission strategy to adaptively match the current interaction environment and network conditions. Compared to the fixed parameter or discrete binary adjustment methods in existing technologies, this embodiment significantly improves the voice communication quality in complex dynamic scenarios, reduces the impact of latency and packet loss, and improves the user experience.

[0043] In one possible implementation, the step of directionally acquiring audio data from the first sound source through a built-in audio acquisition unit and generating corresponding sound source direction information includes: Multiple sound signals are simultaneously acquired using a microphone ring array consisting of at least two microphones; Using a beamforming algorithm, a directional enhancement beam is generated based on the sound source direction information, and the sound signal is weighted and summed to obtain a directional enhancement sound signal; The directional enhanced sound signal is subjected to bandpass filtering and spectral subtraction noise reduction processing to obtain the audio data; The rate of change of the sound source direction information is calculated based on beamforming direction information at multiple consecutive time points.

[0044] The calculation of the rate of change of the sound source direction information includes: Obtain the azimuth and elevation angles corresponding to two adjacent time points in the sound source direction information sequence, and calculate the time difference between these two time points; Calculate the azimuth and elevation differences between the two adjacent time points respectively; The spatial angular displacement is obtained by adding the square of the azimuth difference to the square of the pitch difference. The instantaneous angular velocity is obtained by removing the spatial angular position by the time difference; Within a preset smoothing window, the average or maximum value of the instantaneous angular velocity corresponding to multiple time points is taken, and the calculation result is used as the rate of change of the sound source direction information. The rate of change of the sound source direction information is used to characterize the angular motion velocity of the first sound source relative to the smart glasses.

[0045] In one specific implementation of the present invention, the audio acquisition unit built into the smart glasses employs a microphone ring array consisting of at least two microphones. This ring array can be symmetrically arranged on both sides of the frame of the smart glasses or around the frame, for example, one microphone can be placed on each of the left temple, right temple, and above the bridge of the nose, forming a spatially distributed array. This arrangement allows each microphone to be in a different spatial position when the user wears the glasses, thereby utilizing the time difference and phase difference of sound reaching different microphones to achieve sound source localization and directional enhancement.

[0046] The specific process of targeted data collection is as follows: When a user wears smart glasses and activates the voice communication function, the smart glasses control a circular array of microphones to simultaneously collect sound signals from various directions. Because the distances between each microphone and the target sound source (e.g., a speaker opposite) differ, the arrival time of the same sound signal varies slightly between microphones. The smart glasses' built-in processor uses beamforming algorithms to process these simultaneously collected sound signals. The basic principle of beamforming is: based on a preset focusing direction (e.g., directly in front of the smart glasses), a time delay compensation value is calculated for the signal collected by each microphone. This ensures that the sound signal from that focusing direction is phase-aligned after compensation, thus enhancing it during superposition; while sound signals from other directions are attenuated due to phase misalignment. The processor dynamically adjusts the focusing direction based on real-time calculated sound source direction information, generating a directional enhancement beam, and weighted sums the signals from each microphone to obtain a directional enhancement sound signal. In this signal, the energy from the target sound source is significantly amplified, while environmental noise and side interference are effectively suppressed.

[0047] To further improve audio quality, the smart glasses also perform post-processing on the aforementioned directional enhanced audio signal. First, a bandpass filter is used to remove low-frequency booming and high-frequency hissing outside the human voice frequency range, retaining, for example, the main speech frequency band from 300 Hz to 3400 Hz. Next, a spectral subtraction noise reduction algorithm is employed: the processor compares the spectrum of the current audio signal with a pre-learned or real-time estimated noise spectrum, subtracting the estimated noise spectrum from the signal spectrum to further eliminate smooth background noise, such as fan noise and air conditioner noise. After these processes, clean and clear audio data is finally obtained for subsequent speech activity detection and encoding.

[0048] During the aforementioned beamforming process, the processor continuously estimates the direction of the sound source, typically expressed as azimuth and pitch angles. Azimuth refers to the angle of the sound source on the horizontal plane relative to the front of the smart glasses, while pitch refers to the angle of the sound source on the vertical plane relative to the horizontal direction. These angles change in real time as the user's head turns or the target speaker moves.

[0049] To quantify the magnitude of this change, the processor first records the direction information of the sound source at multiple consecutive time points, forming a sequence. For two adjacent time points in the sequence, the processor obtains the azimuth and elevation angles corresponding to these two moments and calculates the time difference between the two moments.

[0050] Next, the processor calculates the change in azimuth (i.e., the absolute value of the difference between the azimuth and elevation angles at two moments) and the change in elevation angle, respectively. Since the changes in azimuth and elevation angles together determine the change in the direction of the sound source in three-dimensional space, the processor combines these two changes: it calculates the sum of the squares of the change in azimuth and the squares of the change in elevation angle, and then takes the square root of this sum to obtain a comprehensive angular change value, i.e., spatial angular displacement. This value represents the angular distance that the target sound source moves on the three-dimensional sphere from one moment to the next.

[0051] The processor then removes the spatial angular position by the aforementioned time difference to obtain the instantaneous angular velocity within that time interval. The instantaneous angular velocity reflects how quickly the direction of the sound source changes per unit time, and its unit can be degrees per second. For example, if the user turns their head quickly, and the direction of the sound source changes by 30 degrees in 0.1 seconds, then the instantaneous angular velocity is 300 degrees per second; if the user's head remains stable, the instantaneous angular velocity is close to zero.

[0052] Due to noise and slight jitter in real-world scenarios, a single instantaneous angular velocity may not be stable enough. Therefore, the processor collects instantaneous angular velocity values ​​calculated at multiple consecutive time points within a preset smoothing window (e.g., 0.5 seconds or 1 second), and then takes the average or maximum value of these values ​​as the final rate of change of the sound source direction information. Taking the average provides a stable, noise-resistant indicator, suitable for measuring the overall motion intensity over a period of time; taking the maximum value can capture sudden, rapid motion and is more sensitive to sudden changes in direction. Designers can choose the appropriate method based on the specific application scenario.

[0053] The rate of change of the final sound source direction information represents the angular velocity of the first sound source (target speaker) relative to the smart glasses. The higher the value, the faster the user's head or the target sound source is moving, and the more dynamic the current interaction environment is; the lower the value, the more stable both are.

[0054] The effects of the above technical solution are as follows: First, the system utilizes a microphone ring array and beamforming technology to achieve directional enhancement of the target sound source, effectively suppressing environmental noise and improving the signal-to-noise ratio of the acquired audio. Second, bandpass filtering and spectral subtraction further purify the audio signal, providing high-quality input for subsequent speech activity detection and coding. Third, by calculating spatial angular displacement and instantaneous angular velocity and processing them through a smoothing window, a stable and reliable indicator of the sound source direction change rate is obtained. This indicator accurately reflects the motion state of the user's head or the target sound source, providing crucial quantitative evidence for subsequent multi-dimensional decision models to assess the dynamics of the interactive environment. For example, when the change rate is high, the system can appropriately increase the noise reduction intensity or adjust the coding mode to cope with the acquisition challenges brought about by rapid changes in direction, thereby ensuring the stability of speech quality in dynamic scenarios such as user head turning, walking, or multiple speakers alternating.

[0055] In one possible implementation, the statistical analysis of the voice frame density and dialogue turn change rate within a preset time window to generate user communication state data representing real-time interaction intentions includes: Set the length of the preset time window, count the number of frames that are determined to be valid speech by speech activity detection within the window and the total number of frames within the window, and use the ratio of the number of valid speech frames to the total number of frames as the speech frame density. The number of dialogue turns in the current window is detected, and the number of dialogue turns in the previous window is obtained. The time interval between the center moments of the two windows is calculated. The difference between the number of turns in the current window and the number of turns in the previous window is divided by the time interval to obtain the rate of change of the number of dialogue turns. The voice frame density and the dialogue turn change rate are combined into a feature vector, which is then output as the user communication state data to the multidimensional decision model.

[0056] In one specific implementation of the present invention, the smart glasses perform voice activity detection on the collected audio data and statistically analyze two key indicators based on a preset time window: voice frame density and dialogue turn change rate, thereby generating user communication state data representing real-time interaction intentions. This data is then input into a multi-dimensional decision model to assist in dynamically adjusting the encoding mode.

[0057] The built-in voice activity detection module in smart glasses can determine whether the current audio segment is valid human voice, silence, or background noise frame by frame. This module typically uses parameters such as sound energy, zero-crossing rate, and spectral characteristics for this determination, but can also employ a lightweight deep learning model. The length of each audio frame is usually fixed, such as 10 milliseconds, 20 milliseconds, or 30 milliseconds. The voice activity detection module outputs a flag: if the frame contains human voice, it is marked as a valid voice frame; otherwise, it is marked as an invalid voice frame. This process is real-time and continuous, and does not rely on any external information.

[0058] The smart glasses are configured with a fixed-length time window, such as 2 seconds. This window slides forward over time, updating every 0.5 seconds to ensure the system can promptly capture changes in the conversation's pace. Within each time window, the smart glasses count two values: the total number of audio frames within the window (e.g., if the frame length is 20 milliseconds, a 2-second window contains 100 frames), and the number of frames that are identified as valid speech frames by voice activity detection.

[0059] The speech frame density is the ratio obtained by dividing the number of effective speech frames by the total number of frames within the window. This ratio ranges from 0 to 1. When both parties in a conversation speak continuously with almost no pauses or only very short gaps between sentences, the speech frame density is close to 1; when there is a lot of silence or pauses in the conversation (such as thinking, listening, or looking at documents), the speech frame density is lower. Therefore, the speech frame density intuitively reflects the busyness or speaking density of the current conversation.

[0060] A turn in a conversation refers to the number of times a speaker's role changes. In a face-to-face two-person conversation, a turn is typically defined as when User B ends speaking and User A begins speaking, or vice versa. In a multi-person conversation, the turn is counted based on the switching of speaker roles.

[0061] Smart glasses need to identify different speakers, which can be achieved in several ways: one way is to use the aforementioned sound source direction information, where speakers from different directions correspond to different sound source locations, and when effective speech continuously switches from one direction to another, it is counted as a new round; another way is to use speaker recognition or clustering algorithms to distinguish different people based on the spectral characteristics of the sound. For this solution, the accuracy requirement for round counting is not stringent; the key is to capture the trend of the speed of round changes.

[0062] In its implementation, the smart glasses maintain a time window of equal duration (e.g., 2 seconds) that slides across the screen. Within each window, the system counts the total number of dialogue turns that occur, recorded as the current window turn count. Simultaneously, it stores the turn count of the previous time window (whether it overlaps with the previous window or is consecutive without overlap, depending on the design). Then, it calculates the time interval between the center moments of the two windows (e.g., if the window slides once every 0.5 seconds, the center moments of the two windows differ by 0.5 seconds). Next, it subtracts the previous window turn count from the current window turn count to obtain the turn count change, then divides this by the time interval to obtain the dialogue turn change rate. The unit of this change rate can be turns per second. If the dialogue turn change rate is positive and has a large absolute value, it indicates that the dialogue pace is accelerating, with both parties alternating more frequently; if it is close to zero or negative, it indicates that the pace is stable or slowing down.

[0063] Combining the calculated speech frame density and dialogue turn change rate creates a two-dimensional feature vector. This feature vector represents the user communication state data, indicating the real-time interaction intent. Speech frame density reflects the proportion of speaking time, while the dialogue turn change rate reflects the active trend of speaker switching. The combination of these two values ​​allows us to characterize the current dialogue dynamics from two perspectives: how many people are speaking and how quickly speakers switch. For example, in a rapid debate scenario, both speech frame density and dialogue turn change rate are high; in a scenario where one person is speaking and the other is listening, speech frame density may also be high (because one person is continuously speaking), but the turn change rate is close to zero; in a scenario where both parties are thinking and occasionally take turns speaking, speech frame density is low, and the turn change rate is also low. These different combinations of states provide crucial information for subsequent multi-dimensional decision models to determine the current interaction's emphasis on low latency and interference resistance.

[0064] Through the above implementation, this invention achieves the following effects: First, the statistical analysis of speech frame density is simple and efficient, requiring no complex calculations to determine the busyness of the dialogue in real time; second, the dialogue turn change rate can sensitively capture the acceleration or deceleration trend of speaker switching, reflecting the dynamic evolution direction of the interaction better than the absolute value of the turn; third, combining these two indicators into a feature vector enables the subsequent multi-dimensional decision model to distinguish different types of dialogue scenarios. For example, high density and high change rate tend to require lower latency to ensure smooth dialogue transitions; while high density and low change rate (one person speaking continuously) may allow for some buffering to achieve better packet loss resistance. This refined perception based on real-time interaction intent is significantly superior to the coarse-grained methods in existing technologies that rely solely on historical communication duration or coarse response frequency, providing accurate input for achieving smooth and adaptive coding strategy adjustments.

[0065] In one possible implementation, obtaining a real-time quality parameter sequence of the communication link with the mobile terminal, and predicting a short-time predicted value of the link quality based on the real-time quality parameter sequence, includes: Receive transmission quality feedback data periodically sent by the mobile terminal; Round-trip time, jitter value, and acknowledgment loss rate are extracted from the transmission quality feedback data and used as elements of the real-time quality parameter sequence; The real-time quality parameter sequence is processed using a Kalman filter or a linear regression model to predict the link quality trend within a preset time period, and the link quality trend is used as the short-term predicted value of the link quality. The link quality trend includes: worsening, improving, and remaining stable. Specifically, when using a linear regression model, the fitting slope of the most recent N points in the real-time quality parameter sequence is calculated. If the slope is greater than a preset positive slope threshold, it is determined that the quality will worsen; if the slope is less than a preset negative slope threshold, it is determined that the quality will improve; otherwise, it is determined that the quality will remain stable. When using a Kalman filter, based on the comparison between the predicted value at the next moment and the current actual value output by the filter, if the predicted value increases by more than a preset increase threshold compared to the current actual value, it is determined that the quality will worsen; if it decreases by more than a preset decrease threshold, it is determined that the quality will improve; otherwise, it is determined that the quality will remain stable.

[0066] In one specific implementation of the present invention, the smart glasses periodically acquire real-time quality parameters of the communication link through a low-power transmission channel (e.g., a Bluetooth Low Energy link) with the mobile terminal, and analyze these parameter sequences using a predictive model to obtain short-term predicted values ​​of the link quality. These predicted values ​​are used to guide the multi-dimensional decision model to adjust its coding mode in advance, achieving proactive link adaptation.

[0067] After establishing a low-power transmission channel between the smart glasses and the mobile terminal, both parties can agree on a fixed feedback period, such as issuing a quality report every 100 or 200 milliseconds. The mobile terminal, as the receiving end, continuously monitors the reception of voice transmission frames from the smart glasses and packages relevant statistical information into feedback data, periodically sending it back to the smart glasses through the same channel. Upon receiving this feedback data, the smart glasses extract three core parameters: The first category is round-trip time. Round-trip time refers to the total time from when the smart glasses send a probe packet or ordinary data frame to when they receive a corresponding acknowledgment response from the mobile terminal. It can comprehensively reflect twice the sum of the current one-way transmission delay and processing delay of the link. By measuring the round-trip time, the smart glasses can determine whether the link is becoming congested or experiencing increased latency.

[0068] The second category is jitter. Jitter refers to the fluctuation in the time interval between multiple consecutive voice transmission frames arriving at the mobile terminal. For example, if the arrival interval between frames is irregular and varies greatly, the jitter value is high; if the interval is stable, the jitter value is low. Jitter reflects the stability of the link. Smart glasses can calculate this indicator based on the timestamp information carried in the feedback data, or directly receive the jitter value calculated by the mobile terminal.

[0069] The third category is the acknowledgment loss rate. The mobile terminal reports in its feedback how many voice transmission frames were not successfully received in the most recent period; that is, the proportion of lost frames to the total number of transmitted frames. A higher loss rate indicates more severe packet loss in the link, which may be due to signal interference, excessive distance, or obstruction.

[0070] The three types of parameters mentioned above constitute the elements in the real-time quality parameter sequence. The smart glasses continuously record these parameters over multiple periods, forming a sequence that changes over time. For example, the sequence stores the average round-trip time, average jitter value, and average loss rate for each of the past 10 periods in chronological order.

[0071] The smart glasses utilize a predictive model to process the aforementioned real-time quality parameter sequence to predict the trend of link quality changes over a short future period (e.g., 0.5 to 1 second). This solution provides two exemplary prediction methods: a Kalman filter and a linear regression model. Both methods have their own characteristics and are suitable for real-time computation in embedded devices.

[0072] When using a Kalman filter, the smart glasses treat round-trip time, jitter, and packet loss rate as time-varying state variables. The Kalman filter can recursively estimate the optimal state for the current state based on historical observations, even in the presence of noise, and simultaneously predict the state at the next moment. For this solution, the smart glasses do not need to implement a complete matrix Kalman filter; instead, a one-dimensional Kalman filter can be used independently for each quality parameter. Specifically, in each feedback cycle, the smart glasses input the measured parameter value into the filter, and the filter outputs a prediction of that parameter value for the next cycle. If the predicted value shows a significant upward trend compared to the current value, the link quality is determined to deteriorate; if the predicted value decreases, the quality is determined to improve; if the predicted value fluctuates within a small range, the quality is determined to remain stable.

[0073] When using a linear regression model, the smart glasses store data points for the same quality parameter (such as round-trip time) from the most recent 5 to 10 periods in memory. Each data point includes a time number and a parameter value. The smart glasses perform a linear fit on these points to calculate the slope of the parameter value over time. If the slope is positive and exceeds a preset rising threshold, the link quality is considered to have deteriorated; if the slope is negative and its absolute value exceeds a falling threshold, the quality is considered to have improved; if the slope is close to zero, the quality is considered to have remained stable. The linear regression model is simple to implement and requires very little computation, making it ideal for resource-constrained smart glasses.

[0074] Regardless of the prediction model used, the final output short-term link quality prediction is a discrete label containing three possible states: it will worsen, it will improve, or it will remain stable. This label does not need to accurately predict specific future values; it only needs to indicate the direction and trend of change, which is sufficient to provide forward-looking information for multidimensional decision-making models.

[0075] Suppose that the smart glasses detect that the round-trip time gradually increases from 40 milliseconds to 60 milliseconds over three consecutive feedback cycles, and the packet loss rate rises from 0.1% to 0.5%. The Kalman filter predicts that the round-trip time in the next cycle will further increase to over 70 milliseconds, at which point the smart glasses determine that the link quality will deteriorate. This prediction is input into a multidimensional decision model, which adjusts the mixing coefficients in advance towards a buffered coding mode (e.g., increasing them by 0.1 to 0.2) and deepens the echo cancellation depth to enhance resilience against packet loss. When the link actually deteriorates, the system has already completed the pre-adjustment of the coding strategy, thus avoiding significant voice stuttering or frame loss during the deterioration period.

[0076] Conversely, if the link has been in a poor state (high latency, high packet loss) and the parameters have started to improve significantly in the last two cycles, the prediction model will determine that it will improve. The multidimensional decision model can then tentatively adjust towards a real-time coding mode in advance to reduce latency as soon as possible after the link is restored and improve the response speed of the interaction.

[0077] The effects of the above technical solution are as follows: Through the above implementation, the present invention achieves the following beneficial effects: First, by utilizing the periodic transmission quality feedback from the mobile terminal, the smart glasses can continuously obtain the round-trip time, jitter value, and packet loss rate of the link with extremely low additional overhead, without the need for complex active detection; Second, by using a Kalman filter or linear regression model to predict the parameter sequence in real time, the computational load is small, suitable for embedded real-time systems, and trends can be identified before the actual deterioration of the link occurs; Third, by using discrete predicted trend labels (will worsen, will improve, remain stable) as input to the multidimensional decision model, the encoding strategy can be proactively adjusted, avoiding the lag problem of switching only after packet loss in traditional solutions, effectively reducing latency spikes and packet loss rates in voice transmission, and improving the fluency and stability of voice communication.

[0078] In one possible implementation, the decision logic of the multidimensional decision model includes: When the rate of change of the sound source direction information exceeds a first preset threshold, it is determined that the user is in a dynamic interactive environment, the mixing coefficient is adjusted towards 1 to approach the buffer coding mode, and the noise reduction intensity and the echo cancellation depth are increased simultaneously. When the rate of change of the sound source direction information is lower than the second preset threshold and the rate of change of the dialogue turn is higher than the third preset threshold, it is determined that the user is in a state of high real-time demand, the mixing coefficient is adjusted towards 0 to approach the real-time encoding mode, and the noise reduction intensity is reduced simultaneously. When the packet loss rate in the real-time quality parameters exceeds the preset packet loss rate threshold, the mixing coefficient is adjusted towards 1, and the echo cancellation depth is increased simultaneously to enhance the anti-packet loss capability. When the short-term predicted link quality indicates that the communication link will deteriorate, the mixing coefficient is pre-adjusted towards 1; when the short-term predicted link quality indicates that the communication link will improve, the mixing coefficient is pre-adjusted towards 0.

[0079] In one specific implementation of this invention, a multi-dimensional decision-making model is constructed internally within the smart glasses. This model is essentially a rule-based explicit decision-making module, whose function is to comprehensively evaluate input information from multiple different dimensions and output a continuous value between 0 and 1, i.e., a mixing coefficient. The mixing coefficient is subsequently used to dynamically generate the target encoding pattern and preprocessing parameters.

[0080] The input to the multidimensional decision-making model includes four types of information, each from a different processing module described above: The rate of change of sound source direction information: calculated in real time by the audio acquisition and beamforming module, reflecting the angular velocity of the target speaker relative to the smart glasses. The higher the value, the faster the user's head or the target sound source is moving.

[0081] User communication status data includes voice frame density and dialogue turn change rate. Voice frame density reflects the "busyness" of the current conversation; dialogue turn change rate reflects the active trend of speaker switching. Together, they characterize the real-time interaction intent.

[0082] Real-time quality parameters: including the packet loss rate of the current communication link (and may also include latency and jitter values), obtained from periodic feedback from the mobile terminal.

[0083] Short-term link quality prediction: The prediction result of the future link trend by the Kalman filter or linear regression model, with values ​​of worsening, improving, or remaining stable.

[0084] The model's sole output is a continuous value between 0 and 1, the mixing coefficient. This coefficient is designed to mean that: when the mixing coefficient approaches 0, it indicates that the current scenario is more suitable for real-time encoding (for low latency); when the mixing coefficient approaches 1, it indicates that a buffered encoding mode is more suitable (for robustness and integrity); values ​​in between represent a compromise between the two modes.

[0085] For example, the basic method for calculating the mixing coefficient of the model is as follows: taking 0.5 as the baseline value, calculating a positive or negative correction amount according to the deviation of each input parameter from its respective preset threshold, accumulating all correction amounts with the baseline value, and finally limiting the accumulated result to between 0 and 1.

[0086] Specifically: The correction amount is calculated based on the rate of change of the sound source direction information: when the rate of change is lower than the second preset threshold, the correction amount is 0; when the rate of change is higher than the first preset threshold, the correction amount is the positive maximum value (e.g., +0.3); when the rate of change is between the two, the correction amount increases linearly with the increase of the rate of change (linearly increasing from 0 to the maximum value).

[0087] The correction amount is calculated based on user communication status data: this correction item only takes effect when the rate of change of the sound source direction information is lower than the second preset threshold (i.e., head stability). Under this premise, if the rate of change of dialogue turns is higher than the third preset threshold, a negative correction amount (e.g., -0.25) is generated to adjust the mixing coefficients towards the real-time coding mode; otherwise, the correction amount is 0.

[0088] The correction amount is calculated based on real-time quality parameters: when the packet loss rate exceeds the preset packet loss rate threshold, a positive correction amount is generated, and the correction amount increases with the increase of the packet loss rate (for example, it increases linearly from 0 to +0.4); when the packet loss rate does not exceed the threshold, the correction amount is 0.

[0089] The correction amount is calculated based on the short-term predicted link quality: when the predicted value is "worsening", a fixed positive pre-adjustment amount (e.g., +0.1) is added; when the predicted value is "improving", a fixed negative pre-adjustment amount (e.g., -0.1) is added; when the predicted value is stable, the correction amount is 0.

[0090] Add the four correction values ​​mentioned above to the baseline value of 0.5 to obtain the preliminary mixing coefficient. If the preliminary mixing coefficient is less than 0, the final mixing coefficient is 0; if it is greater than 1, it is 1; otherwise, the calculated value is used.

[0091] The multidimensional decision-making model dynamically adjusts the mixing coefficients according to changes in input parameters, following pre-defined rules. These rules reflect trade-off strategies for different scenarios. The following examples illustrate the model's decision-making process.

[0092] Rule 1: Responding to rapid dynamic changes in sound sources When the rate of change of the sound source direction information exceeds a first preset threshold, such as exceeding 90 degrees per second (corresponding to the user rapidly turning their head), the model determines that the user is in a dynamic interactive environment. At this time, regardless of other parameters, the model adjusts the mixing coefficient towards 1 (for example, increasing it by 0.3 from the current value, but not exceeding 1.0), making the encoding mode tend towards a buffered encoding mode. Simultaneously, the model outputs instructions to increase the noise reduction intensity and echo cancellation depth. The rationale for this is that when the head rotates rapidly, the directivity of the microphone array may temporarily deviate from the target sound source, causing the acquired signal to be mixed with more environmental noise and the user's own echo. In this case, prioritizing the clarity of audio acquisition is more important than pursuing the lowest latency.

[0093] Rule 2: Addressing the Need for Highly Real-Time Dialogue When the rate of change of sound source direction information is low (below the second preset threshold, e.g., below 20 degrees per second, indicating head stability) while the rate of change of dialogue turns is high (above the third preset threshold, e.g., exceeding 0.8 times per second, indicating a fast question-and-answer pace), the model determines that the user is in a state of high real-time demand. At this time, the model adjusts the mixing coefficients towards 0 (e.g., reducing them by 0.2 from the current value), making the encoding mode tend towards a real-time encoding mode, and simultaneously reducing the noise reduction intensity. The rationale for this is that in scenarios where the head is stable and dialogue turns alternate frequently (such as fast question-and-answer), users are extremely sensitive to low latency; any additional encoding delay will disrupt the smoothness of the dialogue. In this case, it is appropriate to sacrifice some noise reduction depth in exchange for shorter frame lengths and lower encoding latency.

[0094] Rule 3: Addressing Packet Loss on the Current Link The model monitors the packet loss rate in real-time quality parameters. When the packet loss rate exceeds a preset threshold (e.g., exceeding 5%), it indicates that significant packet loss has occurred in the current link. At this point, the model adjusts the mixing coefficients towards 1 and simultaneously increases the echo cancellation depth to enhance packet loss resilience. Increasing the echo cancellation depth is a metaphorical expression in this scheme; in actual implementation, it could correspond to enabling stronger redundant coding, increasing forward error correction strength, or increasing frame aggregation—all common packet loss resilience techniques in buffered coding schemes.

[0095] Rule 4: Forward-looking link trend adaptation The model receives short-term predictions of link quality. When the prediction indicates a worsening trend, the model pre-adjusts the mixing coefficients towards 1 (e.g., by adding a small offset, such as 0.1) before actual packet loss occurs. This proactive action allows the system to transition to buffered mode before link degradation, avoiding the lag of switching only after severe packet loss. Conversely, when the prediction indicates an improvement, the model pre-adjusts the mixing coefficients towards 0 to return to low-latency mode as quickly as possible after link recovery.

[0096] In actual operation, multiple rules may be triggered simultaneously. For example, when a user quickly turns their head (triggers rule one), the link prediction value also indicates a deterioration (triggers rule four). The multidimensional decision model accumulates the adjustments generated by each rule. In specific implementation, a fixed adjustment step size can be set for each rule, and the final mixing coefficient equals the base value plus the adjustment step sizes of all triggered rules, then clamped between 0 and 1. A more refined weighting strategy can also be used, such as dynamically determining the adjustment magnitude based on the deviation of the input parameters. Regardless of the accumulation method used, the model always maintains a smooth, continuous output value, avoiding abrupt output changes due to rule conflicts.

[0097] The calculated mixing coefficients are passed to the encoding parameter generation module. This module performs linear interpolation on predefined parameters (such as frame length and compression ratio) for both real-time and buffered encoding modes, and also interpolates the noise reduction intensity and echo cancellation depth accordingly, thereby obtaining the target encoding mode and preprocessing parameters for the current frame or current time window. The audio signal is then encoded according to these target parameters and finally transmitted to the mobile terminal through a low-power transmission channel.

[0098] In this embodiment of the invention, the multidimensional decision model involves multiple preset thresholds, including a first preset threshold (high threshold) and a second preset threshold (low threshold) corresponding to the rate of change of sound source direction information, a third preset threshold corresponding to the rate of change of dialogue rounds, and a preset packet loss rate threshold corresponding to the packet loss rate, etc. The typical numerical ranges of these thresholds and their acquisition methods are as follows: The second preset threshold (low threshold) typically ranges from 15 degrees / second to 30 degrees / second, with 20 degrees / second being preferred. When the rate of change is below this value, the user's head is considered relatively stable, and the direction of the sound source does not change drastically.

[0099] First preset threshold (high threshold): The typical value range is 80 degrees / second to 120 degrees / second, with 90 degrees / second being preferred. When the rate of change is higher than this value, it is considered that the user is in a dynamic interactive environment where the target sound source is moving rapidly or turning their head quickly.

[0100] The aforementioned thresholds can be obtained by inviting multiple testers to wear the devices during the smart glasses product development phase, simulating everyday conversation scenarios (including normal conversation, rapid question-and-answer sessions, turning heads to look at different speakers, etc.), while simultaneously recording the rate of change of sound source direction information. Through statistical analysis of a large amount of measured data, for example, taking the 95th percentile of the rate of change under a stable head position as the low threshold and the 5th percentile of the rate of change under a rapid head-turning position as the high threshold, the specific value suitable for the product can be determined. The thresholds may vary slightly under different microphone array layouts and different beamforming algorithm accuracies; these can be obtained through the aforementioned measurement calibration.

[0101] Dialogue turn change rate threshold (third preset threshold): The typical range is 0.5 times / second to 0.8 times / second, preferably 0.6 times / second. When the dialogue turn change rate is higher than this value, it is determined that the dialogue pace is fast, the speaker alternates frequently, and the user is in a state of high real-time demand.

[0102] Acquisition Method: Collect voice samples with different dialogue rhythms (such as monologues, normal conversations, debates, and rapid question-and-answer sessions), calculate the dialogue turn change rate for each sample, and combine this with subjective experience ratings (such as user questionnaire results regarding latency sensitivity) to determine a threshold that can effectively distinguish between high real-time demands and regular real-time demands. Alternatively, an adaptive calibration method can be used, collecting users' own dialogue habit data in the early stages of device use and dynamically adjusting the threshold to adapt to individual user differences.

[0103] Packet loss rate threshold: The typical range is 3% to 8%, with 5% being preferred. When the packet loss rate exceeds this value, the link quality is considered to have significantly degraded, and it is necessary to adjust to a cached encoding mode.

[0104] Acquisition Method: Subjective testing of voice transmission quality at different packet loss rates (e.g., MOS score evaluation). Typically, users can hardly perceive a quality degradation when the packet loss rate is below 3%, while perceptible sound quality deterioration or stuttering begins to appear at 5%. Therefore, 5% is used as the threshold for triggering adjustments. This threshold can also be dynamically adjusted based on the actual communication environment feedback from the mobile terminal (e.g., indoor / outdoor, user movement speed, etc.).

[0105] The effects of the above technical solution are as follows: Through the aforementioned multidimensional decision-making model, this invention achieves the following beneficial effects: First, the model can simultaneously perceive sound source dynamics, dialogue rhythm, current link quality, and future link trends, fusing multidimensional information into a continuous control quantity, avoiding bias caused by single-factor decision-making. Second, the rule setting aligns with engineering intuition and is configurable; designers can adjust the threshold and adjustment step size based on actual product test results, exhibiting good adjustability. Third, the model's output is a continuous value between 0 and 1, which, combined with subsequent parameter interpolation, achieves a smooth transition between encoding modes and preprocessing parameters, avoiding performance jitter caused by traditional discrete switching. Fourth, the proactive adjustment mechanism enables the system to respond before link deterioration occurs, significantly reducing latency spikes and packet loss rates in voice transmission, and improving the user's communication experience in complex dynamic environments.

[0106] The real-time encoding mode corresponds to a first compression ratio and a first frame length, and the cached encoding mode corresponds to a second compression ratio and a second frame length. The first compression ratio is lower than the second compression ratio, and the first frame length is shorter than the second frame length.

[0107] Real-time encoding mode employs a lower compression ratio and shorter frame length, significantly reducing encoding and transmission latency. This allows voice data to reach the other end of the channel faster, meeting the core low-latency requirement of high-real-time interactive scenarios such as rapid question-and-answer sessions and debates. Conversely, buffered encoding mode uses a higher compression ratio and longer frame length. On one hand, the higher compression ratio reduces data volume, thus lowering the pressure on link bandwidth. On the other hand, the longer frame length allows for the aggregation of more voice information and the activation of redundant error correction mechanisms, enhancing resilience against packet loss and effectively ensuring the integrity of voice transmission even in environments with fluctuating link quality or noise. This differentiated design of compression ratio and frame length allows the system to dynamically balance low latency and high robustness according to the actual scenario, avoiding the problem of inconsistent performance across all scenarios with a single fixed encoding parameter.

[0108] In one possible implementation, dynamically generating the target encoding mode and its corresponding preprocessing parameters from at least two preset encoding modes based on the mixing coefficients includes: Linear interpolation is performed on the first encoding parameters corresponding to the real-time encoding mode and the second encoding parameters corresponding to the buffered encoding mode to obtain the encoding parameters of the target encoding mode, wherein the encoding parameters include frame length and compression ratio; Linear interpolation is performed on the first preprocessing parameter corresponding to the real-time encoding mode and the second preprocessing parameter corresponding to the cached encoding mode to obtain the preprocessing parameter corresponding to the target encoding mode.

[0109] This implementation achieves a continuous and smooth transition of the encoding strategy with the mixing coefficients by linearly interpolating the encoding parameters (frame length, compression ratio) and preprocessing parameters (noise reduction intensity, echo cancellation depth) of the real-time encoding mode and the buffered encoding mode. Compared with the discrete two-way switching method in the prior art, linear interpolation avoids the performance jitter caused by mode jumps, making the voice transmission latency and packet loss resistance gradually change with the scene, making the system behavior more predictable and the user experience smoother and more natural. At the same time, this method is computationally simple and easy to implement, without the need for additional encoder fusion or complex scheduling logic, making it suitable for efficient operation on resource-constrained smart glasses.

[0110] In another possible implementation, the step of dynamically generating a target encoding mode and its corresponding preprocessing parameters from at least two preset encoding modes based on the mixing coefficients includes: A nonlinear transformation is performed on the mixing coefficient to obtain an effective mixing coefficient. The nonlinear transformation makes the rate of change of the effective mixing coefficient when the mixing coefficient is in the middle interval lower than the rate of change when the mixing coefficient is in the two end intervals. Based on the effective mixing coefficients, the first encoding parameters corresponding to the real-time encoding mode and the second encoding parameters corresponding to the cached encoding mode are weighted and fused to obtain the encoding parameters of the target encoding mode. The encoding parameters include frame length and compression ratio. In the weighted fusion process, a cooperative constraint between frame length and compression ratio is also introduced: based on the target frame length value obtained by fusion and the preset target bitrate range, the target value of compression ratio is determined so that the combination of frame length and compression ratio satisfies the target bitrate range. Based on the effective mixing coefficients, the first preprocessing parameters corresponding to the real-time coding mode and the second preprocessing parameters corresponding to the buffered coding mode are weighted and fused to obtain the preprocessing parameters corresponding to the target coding mode. The preprocessing parameters include noise reduction intensity and echo cancellation depth.

[0111] The nonlinear transformation is implemented in any of the following ways: Method 1: Use a sigmoid function for transformation: ; Where α is the mixing coefficient; α eff The effective mixing coefficient is denoted by k; k is a steepness parameter greater than 0, with a value between 5 and 20. Method 2: Use a piecewise linear function to transform the interval [0,1] into a lower interval [0,a], a middle interval (a,b), and a higher interval [b,1]. The slope of the transformation function is greater in the lower and higher intervals than in the middle interval, where 0... <a<0.5<b<1。

[0112] The mixing coefficients output by the multidimensional decision model are continuous values ​​between 0 and 1. In actual dialogue, the mixing coefficients may fluctuate slightly around a certain value, especially in the middle region (e.g., between 0.4 and 0.6). This fluctuation often stems from minor, unconscious head movements or momentary fluctuations in link quality, and does not represent actual scene switching needs. Directly using the raw mixing coefficients for parameter interpolation would lead to frequent fine-tuning of the encoding parameters, increasing computational overhead and potentially causing slight fluctuations in speech quality.

[0113] To address this, this implementation first performs a nonlinear transformation on the mixing coefficients to obtain an effective mixing coefficient. The core characteristic of this nonlinear transformation is that when the original mixing coefficients are in the middle range (e.g., between 0.3 and 0.7), the rate of change of the effective mixing coefficient is low; that is, when the original value changes within this range, the range of change in the effective value is compressed. When the original mixing coefficients are in the extreme ranges (i.e., less than 0.3 or greater than 0.7), the rate of change of the effective mixing coefficient is high; that is, a small change in the original value will trigger a large change in the effective value. This characteristic of being flat in the middle and steep at both ends can be figuratively understood as: maintaining stability in the fuzzy zone, and only responding quickly when clearly biased towards a certain mode.

[0114] For example, when the original mixing coefficient changes from 0.4 to 0.5, the effective mixing coefficient after nonlinear transformation may only change from 0.45 to 0.48, a change significantly smaller than the change in the original value. However, when the original mixing coefficient changes from 0.8 to 0.9, the effective mixing coefficient may jump rapidly from 0.85 to 0.97, thus quickly entering the extreme state of the buffered coding mode. This transformation effectively suppresses useless adjustments caused by head micro-jitter or parameter measurement noise, while ensuring a decisive response when a switch is truly needed.

[0115] After obtaining the effective mixing coefficients, the first encoding parameters (including the first frame length and the first compression ratio) corresponding to the real-time encoding mode and the second encoding parameters (including the second frame length and the second compression ratio) corresponding to the cached encoding mode are weighted and fused. The real-time encoding mode uses a shorter frame length (e.g., 20 milliseconds) and a lower compression ratio (e.g., the compressed data volume is 50% of the original data) to pursue low latency; the cached encoding mode uses a longer frame length (e.g., 60 milliseconds) and a higher compression ratio (e.g., the compressed data volume is 12.5% ​​of the original data) to enhance the ability to resist packet loss.

[0116] The fusion method is as follows: target frame length = first frame length × (1 - effective mixing coefficient) + second frame length × effective mixing coefficient; the target compression ratio is similar. When the effective mixing coefficient is close to 0, the target parameters are close to the real-time coding mode; when it is close to 1, it is close to the buffered coding mode; when it is in the middle value, the weighted average of the two is taken to achieve a continuous and smooth parameter transition.

[0117] Weighting and fusing frame length and compression rate independently may result in unreasonable parameter combinations. For example, when the effective mixing coefficient is 0.5, a target frame length of 40 milliseconds and a target compression rate of 31.25% (i.e., 31.25% of the original data) may be obtained. However, if a 40-millisecond voice frame is encoded at this compression rate, the generated data packet size may exceed the maximum transmission unit of a low-power transmission channel (such as Bluetooth Low Energy), causing it to be fragmented or dropped during transmission, thus increasing latency and the risk of packet loss.

[0118] To address this issue, this implementation introduces a collaborative constraint between frame length and compression ratio. Specifically, the system internally presets a target bitrate range (e.g., 16 kilobits per second to 64 kilobits per second). After determining the target frame length based on the effective mixing coefficients, instead of directly using the compression ratio obtained from weighted fusion, a reasonable compression ratio is inferred from the target frame length and the target bitrate range, ensuring that the combination of frame length and compression ratio satisfies the bitrate constraint. For example, if the target frame length is long, a higher compression ratio is automatically selected to ensure that the data volume of a single frame does not exceed the link's transmission capacity; if the target frame length is short, a lower compression ratio is allowed to retain more voice details.

[0119] The specific implementation process of collaborative constraints is as follows: First, determine the initial value of the target frame length based on the effective mixing coefficients; then, calculate the compression ratio range required to maintain the bitrate within a preset range at this frame length; finally, select the value closest to the original weighted fusion compression ratio from this range as the final target compression ratio. This retains the guiding role of the mixing coefficients while ensuring the engineering rationality of the parameter combination.

[0120] Changes in encoding mode should be synchronized with front-end acquisition and processing. In this implementation, preprocessing parameters, including noise reduction intensity and echo cancellation depth, are weighted and fused using the same effective mixing coefficients. In real-time encoding mode, to reduce processing latency, the noise reduction intensity is set to a lower level (e.g., using a low-order filter), and the echo cancellation depth is correspondingly reduced. In buffered encoding mode, to improve anti-interference capability, the noise reduction intensity is increased (e.g., enabling multi-band noise reduction), and the echo cancellation depth is correspondingly increased. Through weighted fusion of effective mixing coefficients, the noise reduction intensity and echo cancellation depth can transition continuously and smoothly between the two extreme modes, ensuring that the control objectives of the entire audio processing chain (from front-end noise reduction to encoded transmission) remain consistent.

[0121] Through the above implementation, the present invention achieves the following beneficial effects: First, the nonlinear transformation effectively suppresses the disturbance of coding parameters caused by minor fluctuations in the mixing coefficients in the intermediate region. Since slight, unintentional head movements of the user or instantaneous fluctuations in link quality will not cause the mixing coefficients to deviate significantly from the intermediate region, these disturbances will not change the effective mixing coefficients after nonlinear compression. This avoids frequent and unnecessary adjustments to the coding parameters, thereby improving the stability and computational efficiency of the system.

[0122] Second, cooperative constraints ensure that the combination of frame length and compression rate is always within a reasonable engineering range. Independent interpolation may result in a combination of long frame length and low compression rate, leading to excessively large single-frame data volume, exceeding the carrying capacity of the Bluetooth Low Energy transmission channel. Cooperative constraints constrain the compression rate through the bit rate range, ensuring that any parameter combination under any effective mixing coefficient can be transmitted smoothly, effectively avoiding packet loss and latency spikes caused by excessively large data packets.

[0123] Third, the joint adjustment of encoding parameters and preprocessing parameters ensures a high degree of consistency in the control objectives of the entire speech processing chain. When the system shifts towards buffered mode due to user head turning or link deterioration, not only does the encoding become more robust, but front-end noise reduction and echo cancellation are also enhanced simultaneously, achieving collaborative optimization from the acquisition end to the encoding end; when shifting towards real-time mode, the latency across the entire chain is reduced synchronously, ensuring the smoothness of the dialogue.

[0124] In one possible implementation, encoding the audio data according to the encoding parameters corresponding to the target encoding mode and the preprocessing parameters to obtain encoded audio data includes: When the mixing coefficient is greater than the preset cache mode threshold, the low bit rate encoder is invoked, and the deep noise reduction algorithm and long-tail echo cancellation algorithm are activated simultaneously to generate coded audio data with strong anti-interference ability. When the mixing coefficient is less than the preset real-time mode threshold, the high-fidelity encoder is invoked, and the filtering order of the noise reduction algorithm is reduced simultaneously to generate low-latency encoded audio data. When the mixing coefficient is between the real-time mode threshold and the cached mode threshold, the outputs of the low bit rate encoder and the high-fidelity encoder are weighted and fused, or the encoding parameters are interpolated.

[0125] In one specific implementation of this invention, the smart glasses use a segmented decision-making strategy to call different encoders and preprocessing algorithms based on the mixing coefficients calculated by the multidimensional decision model and two preset thresholds (real-time mode threshold and cached mode threshold). This strategy further optimizes the efficiency of computing resource utilization while ensuring a smooth transition, avoiding complex parameter interpolation or encoder output fusion in all cases.

[0126] The smart glasses internally store two threshold parameters: a real-time mode threshold and a cached mode threshold. The real-time mode threshold is a small value close to 0, such as 0.2; the cached mode threshold is a large value close to 1, such as 0.8. These two thresholds divide the mixing coefficient's range of 0 to 1 into three intervals: a low value interval (0 to the real-time mode threshold), a medium value interval (real-time mode threshold to the cached mode threshold), and a high value interval (cached mode threshold to 1). The specific values ​​of the two thresholds can be adjusted according to product requirements: if the system is to favor real-time encoding, the real-time mode threshold can be set larger (e.g., 0.3), and the cached mode threshold can also be set larger (e.g., 0.85); if cached mode is preferred, the values ​​are adjusted accordingly. Typical default values ​​are a real-time mode threshold of 0.25 and a cached mode threshold of 0.75.

[0127] When the blending coefficient exceeds the cached mode threshold, for example, a blending coefficient of 0.85, it indicates that the current scene clearly favors the cached encoding mode. At this point, the system no longer performs continuous parameter interpolation but directly switches to the extreme configuration of the cached mode. Specifically, the smart glasses invoke a low bitrate encoder (e.g., an encoder with a bitrate between 8 kilobits per second and 16 kilobits per second, such as the low bitrate mode of SILK or Opus). Low bitrate encoders use longer frame lengths and higher compression ratios, generating smaller amounts of data and exhibiting strong resilience against packet loss.

[0128] Simultaneously, deep noise reduction and long-tail echo cancellation algorithms are activated. Deep noise reduction algorithms can perform more refined spectral analysis and noise suppression on audio signals, such as using multi-subband Wiener filtering or noise reduction models based on recurrent neural networks, effectively filtering out stationary and non-stationary noise. Long-tail echo cancellation algorithms use longer-order adaptive filters to eliminate far-end echoes with multiple reflections (e.g., in environments with many hard walls, such as inside a car or in a conference room). Although these two algorithms have higher computational complexity, the system already accepts relatively high latency in buffered mode, thus tolerating this additional processing overhead. The resulting encoded audio data has extremely strong anti-interference capabilities and is suitable for scenarios where users turn their heads quickly, environments are noisy, or link quality is poor.

[0129] When the mixing factor is less than the real-time mode threshold, for example, a mixing factor of 0.15, it indicates that the current scene is clearly biased towards real-time encoding mode. Directly call a high-fidelity encoder (e.g., an encoder with a bitrate between 32 kilobits per second and 64 kilobits per second, such as Opus's high bitrate mode or AAC-LD). High-fidelity encoders use shorter frame lengths and lower compression rates, preserving more speech details and exhibiting extremely low encoding / decoding latency.

[0130] Simultaneously, the filter order of the noise reduction algorithm is reduced. For example, the original 256th-order noise reduction filter is reduced to 32nd order, or certain fine noise reduction processing steps are skipped, retaining only basic spectral subtraction. The echo cancellation algorithm is also switched to a shorter-order version, eliminating only direct sound and first-reflection echoes. This downsizing strategy significantly reduces the computation time required for preprocessing and the introduced algorithmic latency, allowing the speech signal to enter the encoding and transmission stages at the fastest speed. This is crucial for scenarios that are extremely sensitive to latency, such as rapid question-and-answer sessions and debates, ensuring real-time fluency of the dialogue even at the cost of some audio purity.

[0131] When the blending coefficient falls between the real-time mode threshold and the cached mode threshold, for example, a blending coefficient of 0.5, it indicates that the current scene is in a gray area and cannot be clearly biased towards either extreme mode. In this case, the system adopts one of two optional schemes to achieve a smooth transition.

[0132] The first approach involves weighted fusion of the outputs from the low-bitrate encoder and the high-fidelity encoder. Specifically, the smart glasses simultaneously feed the same audio data into both the low-bitrate encoder and the high-fidelity encoder, resulting in two encoded output streams. Then, weights are calculated based on the mixing coefficients (e.g., the weight of the high-fidelity encoder output = 1 - the mixing coefficient, and the weight of the low-bitrate encoder output = the mixing coefficient). Finally, the two encoded data streams are weighted and merged frame by frame to generate a fused encoded audio data stream. This approach fully utilizes the advantages of both encoders but requires twice the computational resources.

[0133] The second approach involves directly interpolating the encoding parameters without dual-encoder fusion. The smart glasses calculate the target frame length and compression ratio based on the mixing coefficients, then call a configurable general-purpose encoder (such as Opus, which supports continuous bitrate adjustment and frame length selection) to perform a single encoding operation directly according to the target parameters. This approach has lower computational complexity and is more suitable for smart glasses with limited resources. The system can dynamically select the appropriate approach based on the current battery level or CPU load: when the battery is sufficient, the fusion approach is used to achieve the best sound quality; when the battery is low or the system is overheating, the interpolation approach is used to reduce power consumption.

[0134] Through the above-described segmented decision-making implementation method, the present invention achieves the following effects: First, when the mixing coefficients are clearly biased towards extreme modes (less than the real-time mode threshold or greater than the cached mode threshold), the system directly calls the corresponding fixed encoder and fixed preprocessing configuration, avoiding unnecessary continuous calculations and reducing power consumption and processing latency. Compared with existing technologies that always perform interpolation or fusion, this approach saves computing resources while ensuring performance.

[0135] Second, weighted fusion or parameter interpolation is only used in the intermediate ambiguity range, achieving a smooth transition while controlling computational overhead. Since the mixing coefficients are likely to remain stable at one end for most of the actual dialogue (e.g., during long, stable user conversations), and only briefly pass through the intermediate range during scene transitions, the overall average computational load of the system is significantly reduced.

[0136] Third, the preprocessing algorithms (noise reduction and echo cancellation) and the encoder are adjusted in tandem according to the same segmented logic, ensuring the consistency of the end-to-end processing strategy: when biased towards real-time mode, preprocessing latency and encoding latency are reduced synchronously; when biased towards buffered mode, preprocessing robustness and encoding robustness are enhanced synchronously. This coordinated linkage avoids mismatches such as encoding switching to low-latency mode while noise reduction still uses high-latency algorithms, thus fully leveraging the advantages of segmented decision-making.

[0137] In one possible implementation, the method further includes: When the rate of change of the sound source direction information exceeds a preset rate of change threshold and the mixing coefficient tends toward the 0 direction, a visual cue signal is generated. The visual cue signals are displayed through the display unit of the smart glasses to prompt the user to adjust their wearing posture or confirm the sound source lock status.

[0138] In one specific implementation of the present invention, the smart glasses also provide a user interaction feedback mechanism. When a specific combination of conditions is detected, this mechanism sends a visual cue signal to the user through the display unit, guiding the user to adjust their wearing posture or confirm the current sound source lock status, thereby improving the quality of voice acquisition and the overall communication experience.

[0139] The smart glasses continuously monitor two parameters: the rate of change of sound source direction information and the mixing coefficient output by the multidimensional decision model. The rate of change of sound source direction information reflects the angular velocity of the target sound source (e.g., a speaker opposite) relative to the smart glasses. The mixing coefficient characterizes the preferred direction of the current encoding strategy: a mixing coefficient approaching 0 (e.g., less than 0.3) indicates that the system is in or is adjusting to a real-time encoding mode, pursuing low-latency transmission.

[0140] The generation of a visual cue signal is triggered when both of the following conditions are met simultaneously: Condition 1: The rate of change of the sound source direction information exceeds a preset rate of change threshold. The typical range of this threshold is 40 degrees / second to 60 degrees / second, preferably 50 degrees / second. When the rate of change exceeds this value, it indicates that the user's head is turning rapidly, or the target speaker is moving rapidly (e.g., the other party moves from the side to the front while the user is walking and talking), causing a drastic change in the sound source direction.

[0141] Condition 2: The mixing coefficient tends towards 0. Specifically, a real-time mode determination threshold can be set, for example, 0.3. When the mixing coefficient is less than this threshold, the system is considered to be currently biased towards real-time encoding mode.

[0142] It should be noted that the simultaneous triggering of the above two conditions constitutes a special contradictory scenario: the user desires low-latency real-time transmission (small mixing coefficient), but at the same time, there are drastic changes in the direction of the sound source (large rate of change). In this scenario, the short frame length and low compression rate strategy adopted by the real-time encoding mode is more sensitive to fluctuations in acquisition quality. If the sound source deviates from the main lobe of the microphone array, the audio signal-to-noise ratio will drop rapidly, while the low-latency mode lacks the redundant error correction capability of the buffered mode to compensate for this drop, resulting in a significant deterioration in speech recognition accuracy and listening quality.

[0143] When the above triggering conditions are met, the smart glasses generate a visual cue signal and present it to the user through its display unit (such as the waveguide display of augmented reality glasses or a micro OLED screen).

[0144] Visual cues can take one or more forms in combination: Icon hint: A dynamic icon appears in the corner of the display field of view (such as the upper right corner), such as a head outline with an arrow pointing to it, or a warning sign that the microphone is pointing off-center.

[0145] Text prompts: Display brief prompts such as "Keep your head steady" or "Look at the speaker."

[0146] Edge lighting effect: A specific color light effect (such as orange or red) flashes at the edge of the display field of view to attract the user's attention without interfering with the central field of view.

[0147] Semi-transparent guide line: A semi-transparent auxiliary line is superimposed in the field of vision to indicate the deviation angle between the current sound source direction and the front of the glasses, guiding the user to turn their head to align.

[0148] The notification signal can be short-lived (e.g., disappearing automatically after 2 seconds) or continuous until the triggering condition is met. To avoid frequent notifications from interfering with the user, a minimum notification interval can be set (e.g., a maximum of once every 10 seconds).

[0149] After receiving visual cues, users can take corresponding actions: for example, stabilizing their head to reduce shaking, slightly turning their head to realign themselves with the speaker, or confirming the current sound source lock status via voice commands or gestures. When the user adjusts their posture, the rate of change in sound source direction information typically decreases, the triggering condition is no longer met, and the cues automatically disappear. Simultaneously, due to improved acquisition quality, the reliability of backend speech recognition and transmission is restored, and the mixing coefficients may be adjusted accordingly (e.g., a proper callback from real-time mode to buffered mode to avoid continuing to use low-robust coding strategies in dynamic environments).

[0150] Through the above implementation method, the present invention achieves the following effects: First, in the contradictory scenario of "rapid head movement by the user" and "the system being in low-latency real-time mode," visual cues are proactively issued to the user, guiding them to adjust their posture. This improves audio acquisition quality without sacrificing low-latency transmission. This "human-machine collaboration" approach fully utilizes the user's proactive cooperation, compensating for the shortcomings of pure algorithms in extreme dynamic environments.

[0151] Secondly, visual cues are presented through the display unit of the smart glasses. Users do not need to look down at their phones or other external devices. They can obtain the cues simply through their field of vision, maintaining natural eye-to-eye interaction during voice conversations, which is in line with the usage habits of wearable devices.

[0152] Third, the prompt mechanism has clear triggering conditions, appearing only when user intervention is truly needed, thus avoiding the interference of frequent prompts. By reasonably setting the change rate threshold and the mixing coefficient threshold, it can be ensured that prompts will not be triggered in most daily conversation scenarios, but will only be activated in special dynamic scenarios such as turning one's head too quickly or talking while walking, thus balancing user experience and system performance.

[0153] Example 2: This example provides a voice directional transmission system, including: Smart glasses, and a mobile terminal that is communicatively connected to the smart glasses; The smart glasses include: An audio acquisition unit is used to directionally acquire audio data from a sound source; A display unit for projecting an image in front of the user; One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the voice directional transmission method as described in Embodiment 1.

[0154] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the present invention should still fall within the scope of the present invention.

Claims

1. A method for directional voice transmission, characterized in that, The method, applied to smart glasses that communicate with a mobile terminal, includes: The built-in audio acquisition unit collects audio data of the first sound source in a directional manner and generates corresponding sound source direction information; the rate of change of the sound source direction information is calculated based on the sound source direction information sequence. The audio data is subjected to voice activity detection, and the voice frame density and dialogue turn change rate within a preset time window are statistically analyzed to generate user communication status data representing real-time interaction intentions. Obtain the real-time quality parameter sequence of the communication link between the smart glasses and the mobile terminal, and predict the short-time predicted value of the link quality based on the real-time quality parameter sequence; A multidimensional decision model is constructed, taking the short-term predicted value of the link quality, the real-time quality parameters, the user communication status data, and the rate of change of the sound source direction information as input parameters, and dynamically calculating the mixing coefficient of the target coding mode, wherein the mixing coefficient is a continuous value between 0 and 1; Based on the mixing coefficients, a target encoding mode and its corresponding preprocessing parameters are dynamically generated from at least two preset encoding modes, wherein the at least two preset encoding modes include a real-time encoding mode and a buffered encoding mode; and the preprocessing parameters include noise reduction intensity and echo cancellation depth. The audio data is encoded according to the encoding parameters corresponding to the target encoding mode and the preprocessing parameters to obtain encoded audio data; The encoded audio data is sent to the mobile terminal via a low-power transmission channel.

2. The method according to claim 1, characterized in that, The process of acquiring audio data from the first sound source directionally through the built-in audio acquisition unit and generating corresponding sound source direction information includes: Multiple sound signals are simultaneously acquired using a microphone ring array consisting of at least two microphones; Using a beamforming algorithm, a directional enhancement beam is generated based on the sound source direction information, and the sound signal is weighted and summed to obtain a directional enhancement sound signal; The directional enhanced sound signal is subjected to bandpass filtering and spectral subtraction noise reduction processing to obtain the audio data; The rate of change of the sound source direction information is calculated based on beamforming direction information at multiple consecutive time points.

3. The method according to claim 2, characterized in that, The calculation of the rate of change of the sound source direction information includes: Obtain the azimuth and elevation angles corresponding to two adjacent time points in the sound source direction information sequence, and calculate the time difference between these two time points; Calculate the azimuth and elevation differences between the two adjacent time points respectively; The spatial angular displacement is obtained by adding the square of the azimuth difference to the square of the pitch difference. The instantaneous angular velocity is obtained by removing the spatial angular position by the time difference; Within a preset smoothing window, the average or maximum value of the instantaneous angular velocity corresponding to multiple time points is taken, and the calculation result is used as the rate of change of the sound source direction information. The rate of change of the sound source direction information is used to characterize the angular motion velocity of the first sound source relative to the smart glasses.

4. The method according to claim 1, characterized in that, The statistical analysis of voice frame density and dialogue turn change rate within a preset time window generates user communication state data representing real-time interaction intentions, including: Set the length of the preset time window, count the number of frames that are determined to be valid speech by speech activity detection within the window and the total number of frames within the window, and use the ratio of the number of valid speech frames to the total number of frames as the speech frame density. The number of dialogue turns in the current window is detected, and the number of dialogue turns in the previous window is obtained. The time interval between the center moments of the two windows is calculated. The difference between the number of turns in the current window and the number of turns in the previous window is divided by the time interval to obtain the rate of change of the number of dialogue turns. The voice frame density and the dialogue turn change rate are combined into a feature vector, which is then output as the user communication state data to the multidimensional decision model.

5. The method according to claim 1, characterized in that, The step of obtaining a real-time quality parameter sequence of the communication link between the smart glasses and the mobile terminal, and predicting a short-time predicted value of the link quality based on the real-time quality parameter sequence, includes: Receive transmission quality feedback data periodically sent by the mobile terminal; Round-trip time, jitter value, and acknowledgment loss rate are extracted from the transmission quality feedback data and used as elements of the real-time quality parameter sequence; The real-time quality parameter sequence is processed using a Kalman filter or a linear regression model to predict the link quality trend within a preset time period, and the link quality trend is used as the short-term predicted value of the link quality; the link quality trend includes: deteriorating, improving, or remaining stable.

6. The method according to claim 3, characterized in that, The decision logic of the multidimensional decision model includes: When the rate of change of the sound source direction information exceeds a first preset threshold, it is determined that the user is in a dynamic interactive environment, the mixing coefficient is adjusted towards 1 to approach the buffer coding mode, and the noise reduction intensity and the echo cancellation depth are increased simultaneously. When the rate of change of the sound source direction information is lower than the second preset threshold and the rate of change of the dialogue turn is higher than the third preset threshold, it is determined that the user is in a state of high real-time demand, the mixing coefficient is adjusted towards 0 to approach the real-time encoding mode, and the noise reduction intensity is reduced simultaneously. When the packet loss rate in the real-time quality parameters exceeds the preset packet loss rate threshold, the mixing coefficient is adjusted towards 1, and the echo cancellation depth is increased simultaneously to enhance the anti-packet loss capability. When the short-term predicted link quality indicates that the communication link will deteriorate, the mixing coefficient is pre-adjusted towards 1; when the short-term predicted link quality indicates that the communication link will improve, the mixing coefficient is pre-adjusted towards 0.

7. The method according to claim 1, characterized in that, The step of dynamically generating a target encoding mode and its corresponding preprocessing parameters from at least two preset encoding modes based on the mixing coefficients includes: A nonlinear transformation is performed on the mixing coefficient to obtain an effective mixing coefficient. The nonlinear transformation makes the rate of change of the effective mixing coefficient when the mixing coefficient is in the middle interval lower than the rate of change when the mixing coefficient is in the two end intervals. Based on the effective mixing coefficients, the first encoding parameters corresponding to the real-time encoding mode and the second encoding parameters corresponding to the cached encoding mode are weighted and fused to obtain the encoding parameters of the target encoding mode. The encoding parameters include frame length and compression ratio. In the weighted fusion process, a cooperative constraint between frame length and compression ratio is also introduced: based on the target frame length value obtained by fusion and the preset target bitrate range, the target value of compression ratio is determined so that the combination of frame length and compression ratio satisfies the target bitrate range. Based on the effective mixing coefficients, the first preprocessing parameters corresponding to the real-time coding mode and the second preprocessing parameters corresponding to the buffered coding mode are weighted and fused to obtain the preprocessing parameters corresponding to the target coding mode. The preprocessing parameters include noise reduction intensity and echo cancellation depth.

8. The method according to claim 1, characterized in that, The step of encoding the audio data according to the encoding parameters corresponding to the target encoding mode and the preprocessing parameters to obtain encoded audio data includes: When the mixing coefficient is greater than the preset cache mode threshold, the low bit rate encoder is invoked, and the deep noise reduction algorithm and long-tail echo cancellation algorithm are activated simultaneously to generate coded audio data with strong anti-interference ability. When the mixing coefficient is less than the preset real-time mode threshold, the high-fidelity encoder is invoked, and the filtering order of the noise reduction algorithm is reduced simultaneously to generate low-latency encoded audio data. When the mixing coefficient is between the real-time mode threshold and the cached mode threshold, the outputs of the low bit rate encoder and the high-fidelity encoder are weighted and fused, or the encoding parameters are interpolated.

9. The method according to claim 1, characterized in that, The method further includes: When the rate of change of the sound source direction information exceeds a preset rate of change threshold and the mixing coefficient tends toward the 0 direction, a visual cue signal is generated. The visual cue signals are displayed through the display unit of the smart glasses to prompt the user to adjust their wearing posture or confirm the sound source lock status.

10. A voice directional transmission system, characterized in that, include: Smart glasses, and a mobile terminal that is communicatively connected to the smart glasses; The smart glasses include: An audio acquisition unit is used to directionally acquire audio data from a sound source; A display unit for projecting an image in front of the user; One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the voice directional transmission method as described in any one of claims 1-9.