OWS earphone simultaneous interpretation method based on ASR algorithm
Through the combination of distributed multi-microphone arrays, knowledge graphs and sound field compensation algorithms, the problems of low speech recognition accuracy and high multilingual simultaneous interpretation delay of open headphones in complex sound field environments are solved, and the low latency and high quality simultaneous interpretation effect is achieved.
Patent Information
- Application Number
- CN202510784627.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Open headphones have low speech recognition accuracy and high multilingual simultaneous interpretation delay in complex sound field environments. The existing technology lacks effective environmental noise suppression, semantic coherence and user interaction mechanisms.
Adaptive beamforming and noise reduction are performed through distributed multi-microphone arrays, combined with dynamic path planning modules based on knowledge graphs and lightweight hybrid ASR models, the sound field compensation algorithm is used to drive the open speaker array to output translated speech directionally, and the knowledge graph weight is optimized through multimodal interaction interfaces and reinforcement learning mechanisms to form a closed-loop system.
It realizes low-latency and high-quality simultaneous interpretation in complex scenarios, improves speech recognition accuracy and translation accuracy, reduces latency and enhances the system's adaptability.
Smart Images

Figure CN120356457A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of speech recognition technology, wearable devices and real-time translation systems, and particularly relates to a simultaneous interpretation method for OWS headphones based on the ASR algorithm. Background Art
[0002] Traditional simultaneous interpretation systems mostly adopt a closed-in-ear design. Although they can physically isolate environmental noise, long-term wearing is likely to cause discomfort in the ear canal and cannot meet the safety requirements in sports scenarios. Existing open headphones generally have the problem of low signal-to-noise ratio in speech acquisition. Especially in a noisy environment, the speech recognition accuracy drops significantly. The mainstream speech recognition algorithms rely on cloud processing, resulting in a relatively high translation delay and being difficult to meet the needs of real-time conversations. Multilingual translation systems usually adopt a cascaded processing mode with a fixed path, lacking context relevance and being prone to ambiguity when facing professional terms or cultural differences. In the prior art, noise reduction algorithms are mostly based on static acoustic models and cannot adapt to the change of sound source position caused by head movement, resulting in the failure of target speech tracking. In terms of translation path optimization, traditional methods are limited to predefined rules or simple probability models and are difficult to dynamically adjust the term priority and language switching strategy. The sound field control technology of open speakers has limitations, and the translated speech is easily mixed with environmental noise, affecting the auditory clarity. The user interaction method is single, lacking a multimodal feedback mechanism and being unable to correct translation errors or adjust output preferences in real time. Although current research has made some improvements in part, such as using beamforming to improve speech acquisition quality or introducing an attention mechanism to optimize the recognition model, the systematic problem of multi-module collaborative optimization in an open sound field environment has not been solved. Most solutions still have the following defects: the acoustic noise reduction and speech recognition modules are independently designed, resulting in distorted feature transmission; the translation path planning lacks the support of a knowledge graph and has insufficient semantic coherence; the sound field compensation algorithm ignores the dynamic changes of the environment and has poor directional output stability; the user feedback mechanism is disconnected from the core algorithm, and the system's adaptability is limited. This patent constructs a full-link collaborative system from speech acquisition, semantic parsing to sound field output by integrating acoustic structure innovation and algorithm optimization, and overcomes the real-time translation problem of open headphones in complex scenarios. Summary of the Invention
[0003] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a simultaneous interpretation method for OWS headphones based on the ASR algorithm, which is used to solve the problems of low speech recognition accuracy and high multi-language simultaneous interpretation delay of open headphones in complex sound field environments. The present invention dynamically tracks the target sound source through head movement data to suppress environmental noise; constructs a dynamic path planning module based on a knowledge graph, and generates multi-language optimization signals by using the translation confidence association and context dependence relationship between semantic nodes to guide the lightweight hybrid speech recognition model to focus on key semantics; drives the open speaker array to directionally output the translated speech through a sound field compensation algorithm, and continuously optimizes the knowledge graph weights in combination with the reinforcement learning mechanism of user feedback to form a closed-loop system that links environmental perception and user behavior, and finally realizes low-latency and high-quality simultaneous interpretation in complex scenarios. The present invention provides a simultaneous interpretation method for OWS headphones based on the ASR algorithm, including: S1: Collect environmental voice signals through a distributed multi-microphone array, and use an adaptive beamforming algorithm to generate noise-reduced speech feature signals; S2: Input the speech feature signals into a dynamic path planning module based on a knowledge graph, and generate multi-language optimization signals through the translation confidence association and context dependence relationship between semantic nodes; S3: Use the optimization signals to control the lightweight hybrid ASR model, and generate a text sequence signal with path markers through end-cloud collaborative processing, where the end-side model adjusts the attention mechanism focusing range according to the dynamic weights in the knowledge graph; S4: Convert the text sequence signal into a target language text through a neural machine translation module, and drive the open speaker array to generate a directional voice signal through a sound field compensation algorithm; S5: Receive user feedback signals through a multi-modal interaction interface, and use a reinforcement learning mechanism to update the path weights and semantic association parameter update signals in the knowledge graph, where the translation confidence evaluation strategy is dynamically corrected according to the user's head pose change signal, and the parameter update signal is fed back to the dynamic path planning module in S2 to achieve closed-loop optimization.
[0004] In an embodiment of the present invention, in step S1, the distributed multi-microphone array is arranged on the edge of the headphone housing in a ring topology, and each microphone unit is equipped with an independent acoustic preprocessing module. The adaptive beamforming algorithm realizes dynamic noise reduction by iteratively optimizing the sound source localization accuracy, specifically including: using the inertial measurement unit built in the headphone to obtain head movement acceleration data, coupling and calculating the movement acceleration data with the time delay estimation value of the microphone array to generate three-dimensional space sound source tracking parameters, which are used to adjust the spatial filtering weight matrix of the beamforming algorithm in real time, so that when the user's head deflects, the main lobe direction of the beam always aligns with the azimuth of the target speaker, while suppressing reverberation noise and sudden interference sound sources from other directions.
[0005] In an embodiment of the present invention, the method for constructing the dynamic path planning module of the knowledge graph in step S2 includes: dividing the multilingual semantic nodes into three-layer structures of a core term layer, a context association layer, and a scenario adaptation layer according to the domain knowledge graph. The translation confidence association is realized through a semantic similarity propagation algorithm. Specifically: establish cross-language professional vocabulary mapping relationships in the core term layer, construct a dialogue state transition probability matrix in the context association layer, introduce environmental acoustic features as node attributes of the graph neural network in the scenario adaptation layer. When receiving a new speech feature signal, activate the corresponding sub-graph region according to the current dialogue scenario, calculate the comprehensive scores of each language path through the graph attention mechanism, and select the path with the highest score to generate an optimization signal.
[0006] In an embodiment of the present invention, the end-cloud collaborative processing process of the lightweight hybrid model in step S3 includes: deploying a compressed speech recognition basic model on the edge device. This model uses knowledge distillation technology to extract key feature representation capabilities from the cloud large model. When detecting professional domain terms or low-confidence recognition segments, automatically trigger the cloud-assisted verification mechanism. The adjustment method for the focus range of the attention mechanism is: perform importance sampling on the acoustic feature sequence in the time dimension and perform band-pass enhancement on the speech spectrum in the frequency dimension according to the dynamic weight coefficients provided by the knowledge graph, so that the model attention weight distribution changes synchronously with the term priority of the current translation path.
[0007] In an embodiment of the present invention, the implementation method of the sound field compensation algorithm in step S4 includes: establishing an acoustic radiation model of an open speaker array, obtaining the optimal phase distribution parameters by solving the Helmholtz equation. The generation process of the directional speech signal is specifically: first, perform emotional feature analysis on the target language text to determine the fundamental frequency fluctuation range, then design an adaptive filter bank in combination with the environmental noise spectrum characteristics, and finally calculate the driving signal time delay difference according to the spatial position parameters of the speaker units, so that the synthesized speech forms a constructive interference beam in the target direction and a destructive interference region in the non-target direction.
[0008] In an embodiment of the present invention, the reinforcement learning mechanism in step S5 adopts an online learning method based on policy gradients, and its policy network update process satisfies the following formula: where θ represents the policy network parameters, represents the user's interaction action at time t, is the system state vector including the current translation quality evaluation value, environmental noise level, and language switching frequency, Ψ( :T) is the cumulative reward function from time t to the termination time T. This function comprehensively considers the weighted scores of three dimensions: translation accuracy, response latency, and user satisfaction. The gradient calculated by this formula is used to update the connection weights of semantic nodes in the knowledge graph. In an embodiment of the present invention, the implementation method of the multimodal interaction interface includes: dividing a semantic perception partition in the headphone touch area, defining the mapping relationship between the sliding trajectory pattern and the translation control instruction, and at the same time combining the head pose sensor data to implement a composite interaction logic. Specifically: when it is detected that the user's finger continuously swipes in a specific partition, a language switching request signal is generated; when it is recognized that the user nods, a translation confirmation signal is generated; when it is detected that the user shakes their head, the current sentence retranslation mechanism is triggered. All interaction signals are converted into update parameters of the knowledge graph through an event fusion algorithm.
[0009] In an embodiment of the present invention, the evaluation function adopted by the dynamic path planning module when generating the multilingual optimization signal is: where represents the domain relevance weight of the i-th semantic node, is the cross-lingual alignment confidence, is the historical translation error rate, Δτ is the path delay estimate value, η is the delay sensitivity factor, and ε is a small constant to prevent division-by-zero errors. This function synchronously considers the requirements of term professionalism, context consistency, and real-time when calculating. By traversing all feasible paths in the knowledge graph, the path with the largest function value is selected as the optimization output.
[0010] In an embodiment of the present invention, the enhancement processing of the acoustic feature signal includes a multi-stage noise reduction process: first, the generalized cross-correlation algorithm is used to estimate the time difference of arrival of sound to achieve primary beamforming, then a deep neural network is used for residual noise suppression, and finally the speech naturalness is restored through a phase compensation algorithm. The deep neural network includes an encoder-decoder structure. The encoder maps multi-channel speech features to the latent space, and the decoder combines the scene prior information provided by the knowledge graph to generate a clean speech spectrogram. During network training, an adversarial training strategy is adopted to improve the noise robustness.
[0011] In an embodiment of the present invention, the working process of the closed-loop optimization system includes: loading a general domain knowledge graph in the initial stage, continuously collecting user feedback signals and environmental context information during the running process. When a professional domain conversation scenario is detected, automatically download the domain-enhanced graph from the cloud and perform incremental model fine-tuning. At the same time, establish a temporary buffer to store the session context features. At the end of the conversation, update the local knowledge graph according to the optimized path weights, and clear the temporary cache data to protect user privacy.
[0012] The simultaneous interpretation method for OWS headphones based on the ASR algorithm provided by the present invention dynamically tracks the target sound source through head movement data to suppress environmental noise; constructs a dynamic path planning module based on a knowledge graph, and generates multi-language optimized signals by using the translation confidence association and context dependency relationship between semantic nodes to guide the lightweight hybrid speech recognition model to focus on key semantics; drives the open speaker array to output the translated speech directionally through a sound field compensation algorithm, and continuously optimizes the knowledge graph weights in combination with the reinforcement learning mechanism of user feedback to form a closed-loop system that links environmental perception and user behavior, and finally realizes low-latency and high-quality simultaneous interpretation in complex scenarios. Brief Description of the Drawings
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0014] Figure 1 It is a method flow chart of the simultaneous interpretation method for OWS headphones based on the ASR algorithm. Detailed Embodiments
[0015] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0016] It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape and size of the components in actual implementation. The type, quantity and proportion of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0017] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0018] Please refer to Figure 1, shown is the simultaneous interpretation method for OWS headphones based on the ASR algorithm of the present invention. The simultaneous interpretation method for OWS headphones based on the ASR algorithm of the present invention includes five steps: S1: Collect environmental voice signals through a distributed multi-microphone array, and use an adaptive beamforming algorithm to generate a noise-reduced voice feature signal; S2: Input the voice feature signal into a dynamic path planning module based on a knowledge graph, and generate a multi-language optimized signal through the translation confidence association and context dependency relationship between semantic nodes; S3: Use the optimized signal to control a lightweight hybrid ASR model, and generate a text sequence signal with path markers through end-cloud collaborative processing, where the end-side model adjusts the attention mechanism focusing range according to the dynamic weights in the knowledge graph; S4: Convert the text sequence signal into a target language text through a neural machine translation module, and drive an open speaker array to generate a directional voice signal through a sound field compensation algorithm; S5: Receive the user feedback signal through a multi-modal interaction interface, and use a reinforcement learning mechanism to update the path weights and semantic association parameter update signals in the knowledge graph, where the translation confidence evaluation strategy is dynamically corrected according to the user's head pose change signal, and the parameter update signal is fed back to the dynamic path planning module in S2 to achieve closed-loop optimization.
[0019] Such as Figure 1As shown, the simultaneous interpretation method is based on the collaborative design of the physical structure and algorithm of an open - ear headset. Its core process starts from the signal acquisition link of a distributed multi - microphone array. The array adopts a circular topology layout, and each microphone unit is evenly distributed at equal angular intervals on the edge of the headset shell, forming a 360 - degree acoustic wave coverage area. Each microphone in the array is equipped with an independent acoustic pre - processing module, which includes a pre - emphasis filter, an automatic gain control circuit, and an analog - to - digital conversion unit to ensure the integrity of the original speech signal in both the time domain and the frequency domain. The adaptive beamforming algorithm realizes noise suppression through spatial filtering technology. Its core lies in dynamically constructing a directional beam, aligning the main lobe with the direction of the target sound source, while suppressing the interference noise in the sidelobe direction. When implementing the algorithm, first, the generalized cross - correlation method is used to calculate the time - delay difference between each microphone, and combined with the head motion acceleration data obtained by the inertial measurement unit, a three - dimensional sound source tracking model is constructed. When the user's head deflects, the system real - time calculates the new azimuth angle of the sound source and updates the beamforming weight matrix accordingly to maintain continuous tracking of the target speaker. In this process, the optimization goal of the weight matrix is to minimize the speech distortion under noise interference, and overfitting is prevented through regularization constraints, and finally, a feature signal containing the pure speech component is output. After generating the speech feature signal, the system inputs it into a dynamic path planning module based on a knowledge graph. This module is built on a three - layer graph structure: the core term layer stores the mapping relationships of cross - language professional vocabulary and domain attribute labels; the context association layer records the semantic dependency relationships and state transition probabilities in the conversation history; the scene adaptation layer integrates environmental acoustic features and user interaction behavior data. When receiving new speech features, the module first performs semantic role annotation and intent recognition to activate the sub - graph area of the relevant field. Subsequently, the confidence scores of each language path are calculated through the graph attention mechanism. The scoring elements include the term alignment accuracy rate, the context coherence index, and the real - time environmental noise impact factor. When specifically implemented, the system uses the random - walk algorithm to traverse the feasible paths, dynamically adjusting the transition probability between nodes during each walk, and preferentially selecting the edges with high - confidence associations. The finally generated optimized signal includes target language selection parameters, term library priority weights, and context association rules, and these parameters will be used as the decision basis for subsequent speech recognition and translation processes.
[0020] Furthermore, the implementation details of the distributed microphone array and the adaptive beamforming algorithm in step S1. The advantage of the array using a ring topology design is to balance the acoustic response characteristics in all directions. The spacing of each microphone unit is determined by calculating the Helmholtz resonance frequency to avoid phase cancellation phenomena in specific frequency bands. The acoustic preprocessing module includes a multi-stage signal conditioning circuit: the first stage uses a programmable gain amplifier to dynamically adjust the sensitivity of each channel to eliminate signal amplitude deviations caused by differences in wearing angles; the second stage applies an anti-aliasing filter bank to dynamically switch the cut-off frequency according to the bandwidth of the speech signal; the third stage realizes clock-synchronized sampling to ensure that the time alignment accuracy of multi-channel signals reaches the microsecond level. The adaptive beamforming algorithm achieves dynamic noise reduction through iterative optimization. Its core lies in constructing space-time joint constraint conditions: in the time domain, linear prediction residual analysis is used to distinguish speech and noise components, and in the space domain, the main noise direction is extracted through covariance matrix eigenvalue decomposition. When the algorithm runs, first calculate the estimated value of the noise covariance matrix of the current frame, construct a regularization constraint term in combination with the historical speech source estimation matrix, and solve for the optimal beam weight through convex optimization. In particular, the system introduces a head movement compensation mechanism that converts the three-axis angular velocity data output by the inertial measurement unit into a spatial rotation matrix to real-time correct the calculated value of the sound source azimuth angle, ensuring stable tracking of the target sound source even when the user turns their head quickly.
[0021] As Figure 1As shown, the construction and operation mechanism of the knowledge graph dynamic path planning module. The construction of the core term layer adopts a domain-adaptive semi-supervised learning method: extract high-frequency terms from the professional corpus to build the initial nodes, establish multilingual mapping relationships through cross-lingual word vector alignment technology, and attach domain attribute tags (such as medical, legal, engineering, etc.) and timeliness markers to each node. The context association layer is modeled by a temporal graph neural network. The nodes represent semantic units in the dialogue, and the edge weights are composed of two parts: the co-occurrence probability based on statistics and the context prediction score based on deep learning. The key innovation of the scenario adaptation layer lies in fusing multi-modal perception data: the environmental noise spectrum features are encoded as graph node attributes through Mel cepstral coefficients, and the user interaction behavior data (such as language switching frequency, term correction records) are transformed into edge weight correction factors. When new speech features are input, the system performs multi-level graph reasoning: first, perform exact matching in the core term layer to identify professional terms and their associated fields; then, perform probability reasoning in the context association layer to predict the potential development direction of the current dialogue; finally, adjust the path score by integrating environmental and user factors in the scenario adaptation layer. The specific implementation of the graph attention mechanism adopts a multi-head attention architecture, and each attention head focuses on different dimensions of associated features: the first head calculates the term professionalism score, the second head evaluates the context coherence, and the third head quantifies the degree of real-time environmental interference. The system obtains the final path score by weighted fusion of the outputs of each head and selects the path with the highest comprehensive score to generate an optimization signal. The lightweight hybrid model adopts an edge-cloud collaborative architecture design. The compressed model deployed on the edge side is constructed based on knowledge distillation technology: first, train a basic model with a deep bidirectional Transformer structure in the cloud, then extract the key feature representation ability through the inter-layer attention transfer method, and finally obtain a lightweight model with 80% less parameters. This model adopts dynamic computational graph technology and adaptively adjusts the network depth according to the complexity of the input speech: when detecting common words with clear pronunciation, only activate the shallow network; when encountering fuzzy pronunciation or professional terms, activate deeper network structures layer by layer. The adjustment strategy for the focus range of the attention mechanism includes two dimensions: in the time dimension, perform importance sampling on the speech frames according to the term priority coefficient provided by the knowledge graph, and give priority to processing the acoustic features in the high-weight interval; in the frequency dimension, dynamically generate a band-pass filter in combination with the environmental noise spectrum features to strengthen the energy distribution of the key frequency bands of the speech. The cloud-assisted verification mechanism adopts a differential privacy protection data transmission protocol. When the edge-side model detects a low-confidence recognition segment, it automatically extracts the encrypted feature vector and uploads it to the cloud. The cloud large model returns the correction result and updates the parameters of the edge-side model at the same time. This collaborative mechanism improves the recognition accuracy of professional terms by about 40% while ensuring privacy and security.
[0022] In an embodiment of the present invention, the dynamic path planning module and the speech recognition module form a two-way data stream: on the one hand, the path optimization signal guides the ASR model to adjust the attention distribution and focus on the key semantic units of the current translation path; on the other hand, the recognition result output by the ASR is fed back to the knowledge graph to update the state transition probability of the context association layer. This closed-loop interaction mechanism enables the system to adapt to the dynamic changes of the dialogue scenario. For example, when it is detected that the user frequently switches languages, the delay-sensitive factor in the path evaluation function is automatically increased; when the environmental noise suddenly increases, the decision weight of the core term layer is temporarily increased. Experiments show that the semantic coherence score of this architecture is 35% higher than that of traditional methods in a noisy environment, and the translation delay is reduced to within 800 milliseconds.
[0023] Furthermore, the co - design of a distributed microphone array and an adaptive algorithm breaks through the limitations of traditional noise reduction solutions. By integrating physical acoustic modeling and deep learning technologies, the system achieves noise suppression at three levels: at the physical level, beamforming is used to eliminate spatially separated noise; at the algorithmic level, a deep neural network is used to suppress residual steady - state noise; at the application level, semantic - level noise filtering is performed by combining knowledge graph information. A specially designed phase compensation algorithm can restore the naturalness of speech lost during noise reduction processing. The principle is to analyze the phase spectrum characteristics of clean speech, construct a phase reconstruction model, so that the output speech maintains the prosodic characteristics of the original pronunciation. The directional output technology of the open - type speaker array uses a parametric sound field synthesis method. By solving the wave equation, the optimal driving signal parameters are obtained, realizing the formation of a high - energy speech beam in the target direction while forming a sound wave cancellation region in other directions, effectively reducing the aliasing interference between the translated speech and environmental noise. The simultaneous interpretation method adopts a cooperative working mechanism of a sound field compensation algorithm and a neural machine translation module in the speech synthesis and output links. The neural machine translation module adopts an encoder - decoder architecture based on Transformer. After the encoder receives the text sequence signal with path tags, it first analyzes the term priority and context association rules in the path tags and inputs them as extended features of the position encoding into the model. When the decoder generates the target - language text, it dynamically adjusts the attention distribution by combining the domain - adaptation parameters provided by the knowledge graph, and gives priority to focusing on the mapping relationship of professional vocabulary in the core term layer. The core of the sound field compensation algorithm lies in constructing the sound radiation model of the open - type speaker array, and obtaining the optimal phase distribution by solving the boundary - condition problem of the Helmholtz equation. In specific implementation, the system first analyzes the spectral characteristics of environmental noise, extracts the main interference frequency - band information, and then determines the fundamental - frequency fluctuation range according to the emotional characteristics of the target speech (such as speech rate, intonation fluctuation), generating an initial speech signal with natural prosody. The driving - signal generation process of the speaker array includes three key steps: First, calculate the acoustic wave propagation path difference of each unit according to the polar - coordinate position parameters of the array units; Second, design an adaptive filter bank in combination with the environmental noise spectrum to enhance the energy of specific frequency bands of the target speech; Third, adjust the phase of the driving signals of each unit through a time - delay control module, so that the synthesized acoustic waves form a constructive - interference beam in the target direction while forming a destructive - interference region in the lateral and backward directions. During this process, the system monitors the changes in the environmental sound field in real time and dynamically updates the phase compensation parameters using a feedback - type learning algorithm. For example, when sudden high - frequency noise is detected, the energy distribution of the speech signal in the 2 - 4 kHz frequency band is automatically enhanced, and the beam width of the speaker array is adjusted to improve the anti - interference ability. Experiments show that this sound field compensation technology can improve the speech clarity in the target direction by more than 50%, while reducing the sound pressure level in non - target directions below the environmental noise level.
[0024] such as Figure 1As shown, the reinforcement learning mechanism uses the asynchronous policy gradient algorithm to achieve continuous optimization of the knowledge graph. The policy network architecture consists of two parallel sub-networks: the state feature extraction network and the action value evaluation network. The state feature extraction network receives multi-dimensional input vectors, including real-time translation quality scores (calculated based on semantic similarity), environmental noise levels (obtained through spectrum analysis), user interaction frequencies (such as the number of language switches), and historical path selection records. This network extracts spatio-temporal correlation features through multiple layers of convolution and gated recurrent units, generating a 128-dimensional state embedding vector. The action value evaluation network calculates the probability distribution of optional actions based on the current state vector, where the action space is defined as the adjustment direction and amplitude of the connection weights of the knowledge graph nodes. The cumulative reward function design in the policy gradient update formula includes three dimensions: the translation accuracy reward is calculated based on the edit distance between the recognition result and the manual annotation, the response latency reward uses the negative exponential function to map the time difference, and the user satisfaction reward is quantified by analyzing interaction behavior patterns (such as the confirmation signal frequency, the number of retranslation requests). The system performs policy updates at the end of each conversation, using importance sampling techniques to balance the relationship between exploration and exploitation, ensuring that the knowledge graph can quickly adapt to new scenarios without deviating too much from the verified effective paths. For example, in a medical consultation scenario, when the system detects that the user repeats a specific medical term multiple times, it automatically increases the weight of the relevant term node in the core layer and strengthens the confidence score of its cross-language mapping relationship. The multi-modal interaction interface uses a hierarchical event processing architecture to achieve accurate parsing of user intentions. The touch area is divided into three functional partitions: the front area maps the language switch instruction, the rear area is associated with the translation mode selection, and the top area is bound to the volume adjustment function. Each partition surface is covered with a pressure-sensitive sensor array, which can detect the speed, direction, and contact area of the sliding trajectory. The determination of the interaction logic uses a spatio-temporal joint analysis method: in the time dimension, the long short-term memory network is used to identify the temporal pattern of continuous gestures; in the space dimension, the convolutional neural network is used to analyze the geometric features of the touch trajectory. When processing the data fusion of the head pose sensor, the quaternion representation method is used to calculate the three-dimensional rotation angle of the head, and the Kalman filter is combined to eliminate the motion noise. When the system detects a nodding action (the pitch angle changes by more than 15 degrees), it triggers the translation confirmation process, marking the current translation as a high-confidence sample for model fine-tuning; when a shaking action (the yaw angle swings by more than 20 degrees) is recognized, a three-level retranslation mechanism is started: the first level calls the local cache to regenerate the translation, the second level requests cloud-assisted translation, and the third level activates the manual review channel. All interaction events are managed through a priority queue, and emergency instructions (such as mistranslations of safety warning terms) can interrupt the current processing flow for priority response. The core of the event fusion algorithm lies in constructing a multi-modal signal correlation matrix, calculating the confidence weights of each signal source through the attention mechanism, and the final updated parameters of the knowledge graph include the semantic node weight correction amount, the edge connection strength adjustment factor, and the scene adaptation coefficient.The dynamic path evaluation function adopts a multi-objective optimization framework to balance the translation quality and real-time requirements. The term professional weight ω_i in the function is comprehensively calculated based on the TF-IDF value of the term in the domain knowledge base, the recent usage frequency, and the importance degree annotated by the user. Among them, for the domain relevance part, a representation learning method based on graph neural network is used to embed the term nodes into a low-dimensional space and then calculate the cosine similarity with the current dialogue scenario. The evaluation of the cross-lingual alignment confidence ξ_i includes two parts: static and dynamic. The static part comes from the pre-trained multi-language word vector alignment score, and the dynamic part is adjusted by monitoring the co-occurrence probability of bilingual terms in the context window in real time. The historical translation error rate γ_i is calculated using the sliding window exponential decay method, with more recent error records having a higher weight. At the same time, an error correction feedback gain factor is introduced. When the user actively corrects an error, the error rate value of the relevant term will decay super-linearly. The modeling of the path delay estimate value Δτ considers the computing resource allocation status and network transmission quality, and uses the M / M / c model in queuing theory to predict the time consumption of each processing link. The dynamic adjustment strategy of the delay-sensitive factor η is divided into a basic mode and an emergency mode: it grows slowly according to an exponential law in regular conversations, and switches to a linear fast growth mode when it is detected that the user's speech speed increases or the interaction frequency rises. During the function calculation process, the system maintains a priority queue to store candidate paths, uses the branch and bound method to quickly prune low-score paths, and at the same time retains the sub-optimal solution as a disaster recovery backup. When the comprehensive score of the optimal path is lower than the safety threshold, the cross-modal verification process is automatically triggered to coordinate the speech recognition, environment perception, and user feedback modules for joint decision-making.
[0025] Furthermore, the design of the speaker array breaks through the physical limitations of traditional open headphones. The array unit adopts an irregular diaphragm structure, and the diaphragm shape is optimized through finite element analysis to enable it to have a flat frequency response characteristic in the 300 Hz - 5 kHz frequency band. The drive circuit design adopts a combination scheme of a digital power amplifier and a brain-inspired computing chip. The digital power amplifier is responsible for basic signal amplification, and the brain-inspired chip realizes real-time parameter adjustment of pulse neural network control. The innovation of the phase compensation algorithm lies in introducing a recursive prediction mechanism for the ambient sound field: First, the residual noise signal is collected through an auxiliary microphone, then the improved LMS adaptive filter is used to estimate the sound field transfer function, and finally, the phase of the drive signal is adjusted in advance based on the prediction model. This forward-looking compensation strategy can improve the response speed of sound wave cancellation to the 10-millisecond level, effectively coping with rapidly changing ambient noise. Experimental data shows that in a 70 dB background noise environment, this technology can maintain the speech signal-to-noise ratio in the target direction above 15 dB. The optimization of the neural machine translation module is reflected in the combination of domain adaptation and real-time learning capabilities. The module is built with a dual-buffer mechanism: the main buffer stores the context feature vectors of the current conversation, and the secondary buffer retains the translation results of recent high-frequency terms. When a professional domain term is detected, the term priority processing process is started: First, the recent usage records are retrieved from the secondary buffer. If there is a high-confidence translation, it is directly called; otherwise, the domain adaptation sub-network is activated, and this sub-network improves the accuracy of terms by fine-tuning the key-value matrix of the attention layer. During the translation process, a confidence heat map is generated synchronously, and low-confidence segments automatically trigger a multi-strategy fallback mechanism, including synonym replacement, word order adjustment, and context association expansion. The data interaction between the sound field compensation algorithm and the translation module is realized through shared hidden layer representations. The emotional feature vectors output by the translation module directly guide the generation of sound field synthesis parameters. For example, the upward amplitude of the fundamental frequency curve is automatically increased at the end of an interrogative sentence.
[0026] The design of the reinforcement learning system emphasizes safety and interpretability. When training the policy network, a constraint optimization framework is introduced. Through the Lagrange multiplier method, it is ensured that the key connection weights of the knowledge graph do not mutate. The design of the reward function includes a regularization term to prevent over-optimizing short-term rewards and damaging the system stability. The network update adopts a dual-objective optimization strategy: the main objective is to maximize the cumulative reward, and the auxiliary objective is to minimize the policy fluctuation amplitude. The version management of the knowledge graph uses blockchain technology, and each update operation generates an immutable record, which is convenient for tracing the source of errors. When a malignant feedback loop is detected (such as continuously reducing the weight of a certain term resulting in semantic breakage), it automatically rolls back to the nearest stable version and issues a maintenance alert. The fault tolerance mechanism of the multi-modal interaction interface includes a multi-level verification process. When parsing touch instructions, an ensemble learning method is used to synthesize the judgment results of multiple sensors: the pressure distribution pattern of the pressure-sensitive sensor, the change trend of the contact area of the capacitive sensor, and the proximity detection data of the infrared sensor are jointly input into the random forest classifier to reduce the false touch rate. The head pose recognition introduces a kinematic model constraint to compare the original sensor data with the biomechanical characteristics of the human neck movement and filter out unreasonable abnormal actions. The event fusion algorithm adopts a federated learning framework. On the premise of protecting user privacy, it improves the accuracy of interactive intention recognition through multi-device collaborative training. The system is also designed with an emergency bypass channel. When a continuous abnormal interaction signal is detected, it automatically switches to the basic translation mode and starts the self-check program.
[0027] Specifically, the computational efficiency of the dynamic path evaluation function is significantly improved through hardware acceleration. The system is equipped with a dedicated graph computing coprocessor, and a parallel edge traversal algorithm is used to accelerate the search process of the knowledge graph. For large-scale graph scenarios, a hierarchical pruning strategy is developed: the first layer quickly filters out irrelevant subgraphs based on domain labels, the second layer applies an approximation algorithm to estimate the upper limit of the path score, and the third layer performs precise calculations on the candidate paths. The calculation of the exponential term in the function is optimized by combining a lookup table and Taylor expansion, reducing the operation time by 60% while ensuring accuracy. The real-time guarantee is also reflected in the incremental update mechanism. When new voice features arrive, only the affected part of the path is re-evaluated instead of traversing the entire graph.
[0028] The simultaneous interpretation method of the OWS headset based on the ASR algorithm of the present invention dynamically tracks the target sound source through head movement data to suppress environmental noise; constructs a dynamic path planning module based on the knowledge graph, and uses the translation confidence association and context dependency relationship between semantic nodes to generate multi-language optimization signals to guide the lightweight hybrid speech recognition model to focus on key semantics; drives the open speaker array to output the translated speech directionally through the sound field compensation algorithm, and continuously optimizes the knowledge graph weight in combination with the reinforcement learning mechanism of user feedback to form a closed-loop system of environmental perception and user behavior linkage, and finally realizes low-latency and high-quality simultaneous interpretation in complex scenarios.
[0029] Therefore, through the OWS headset simultaneous interpretation method based on the ASR algorithm of the present invention, the problems of low speech recognition accuracy and high multi-language simultaneous interpretation delay of open headphones in a complex sound field environment can be solved.
[0030] The above embodiments are only illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A simultaneous interpretation method for OWS headphones based on the ASR algorithm, characterized in that Including: S1: Collect ambient voice signals through a distributed multi-microphone array, and generate a noise-reduced voice feature signal using an adaptive beamforming algorithm; S2: Input the voice feature signal into a dynamic path planning module based on a knowledge graph, and generate a multi-language optimized signal through the translation confidence association and context dependency relationship between semantic nodes; S3: Use the optimized signal to control a lightweight hybrid ASR model, and generate a text sequence signal with path markers through end-cloud collaborative processing, where the end-side model adjusts the attention mechanism focusing range according to the dynamic weights in the knowledge graph; S4: Convert the text sequence signal into a target language text through a neural machine translation module, and drive an open speaker array to generate a directional voice signal through an acoustic field compensation algorithm; S5: Receive the directional voice signal through a multi-modal interaction interface, and update the path weights and semantic association parameter update signals in the knowledge graph using a reinforcement learning mechanism, where the translation confidence evaluation strategy is dynamically corrected according to the user's head pose change signal, and the parameter update signal is fed back to the dynamic path planning module in S2 to achieve closed-loop optimization.
2. The method for simultaneous interpretation of the OWS headset based on the ASR algorithm according to claim 1, wherein In step S1, the distributed multi-microphone array is arranged on the edge of the earphone shell in a ring topology, and each microphone unit is equipped with an independent acoustic pre-processing module. The adaptive beamforming algorithm realizes dynamic noise reduction by iteratively optimizing the sound source localization accuracy, specifically including: obtaining head movement acceleration data using the inertial measurement unit built in the earphone, coupling and calculating the movement acceleration data with the time delay estimation value of the microphone array to generate three-dimensional space sound source tracking parameters, which are used to adjust the spatial filtering weight matrix of the beamforming algorithm in real time, so that when the user's head deflects, the main lobe direction of the beam always aligns with the azimuth of the target speaker, while suppressing reverberation noise and sudden interference sound sources from other directions.
3. The method for simultaneous interpretation of OWS headphones based on the ASR algorithm according to claim 1, wherein, The construction method of the dynamic path planning module of the knowledge graph in step S2 includes: dividing multi-language semantic nodes into three-layer structures of a core term layer, a context association layer, and a scene adaptation layer according to the domain knowledge graph. The translation confidence association is realized through a semantic similarity propagation algorithm, specifically: establishing a cross-language professional vocabulary mapping relationship in the core term layer, constructing a dialogue state transition probability matrix in the context association layer, and introducing environmental acoustic features as graph neural network node attributes in the scene adaptation layer. When a new voice feature signal is received, the corresponding sub-graph area is activated according to the current dialogue scene, and the comprehensive scores of each language path are calculated through a graph attention mechanism, and the path with the highest score is selected to generate an optimized signal.
4. The simultaneous interpretation method of the OWS headset based on the ASR algorithm according to claim 1, characterized in that, The end-cloud collaborative processing flow of the lightweight hybrid ASR model described in step S3 includes: deploying the compressed speech recognition basic model on the edge device, which uses knowledge distillation technology to extract key feature representation capabilities from the cloud large model. When detecting professional domain terms or low-confidence recognition segments, the cloud-assisted verification mechanism is automatically triggered. The adjustment method for the focusing range of the attention mechanism is as follows: according to the dynamic weight coefficients provided by the knowledge graph, importance sampling is performed on the acoustic feature sequence in the time dimension, and band-pass enhancement is performed on the speech spectrum in the frequency dimension, so that the model attention weight distribution changes synchronously with the term priority of the current translation path.
5. The method for simultaneous interpretation of the OWS headset based on the ASR algorithm according to claim 1, wherein, The implementation method of the sound field compensation algorithm described in step S4 includes: establishing the sound radiation model of the open speaker array, and obtaining the optimal phase distribution parameters by solving the Helmholtz equation. The generation process of the directional speech signal is specifically as follows: first, analyze the emotional feature of the target language text to determine the fundamental frequency fluctuation range, then design an adaptive filter bank in combination with the environmental noise spectrum feature, and finally calculate the driving signal time delay difference according to the spatial position parameters of the speaker unit, so that the synthesized speech forms a constructive interference beam in the target direction, and at the same time forms a destructive interference region in the non-target direction.
6. The simultaneous interpretation method for OWS headphones based on the ASR algorithm according to claim 1, characterized in that, In step S5, the reinforcement learning mechanism adopts an online learning method based on policy gradient, and the policy network update process satisfies the following formula: where θ represents the policy network parameters, represents the interaction action of the user at time t, is the system state vector containing the current translation quality evaluation value, environmental noise level, and language switching frequency, Ψ( :T) is the cumulative reward function from time t to the termination time T. This function comprehensively considers the weighted scores of three dimensions: translation accuracy, response latency, and user satisfaction. The gradient calculated through this formula is used to update the connection weights of semantic nodes in the knowledge graph.
7. The method for simultaneous interpretation of the OWS headset based on the ASR algorithm according to claim 1, wherein, The implementation method of the multi-modal interaction interface includes: dividing the semantic perception partition in the headphone touch area, defining the mapping relationship between the sliding trajectory mode and the translation control instruction, and at the same time implementing the composite interaction logic in combination with the head pose sensor data. Specifically, when it is detected that the user's finger continuously slides in a specific partition, a language switching request signal is generated; when the user's nodding action is recognized, a translation confirmation signal is generated; when the user's shaking head action is detected, the current sentence retranslation mechanism is triggered. All interaction signals are converted into update parameters of the knowledge graph through the event fusion algorithm.
8. The method for simultaneous interpretation of the OWS headset based on the ASR algorithm according to claim 1, wherein, The evaluation function adopted by the dynamic path planning module when generating multilingual optimization signals is as follows: where represents the domain relevance weight of the i-th semantic node, is the cross-lingual alignment confidence, is the historical translation error rate, Δτ is the path delay estimation value, η is the delay sensitivity factor, and ε is a small constant to prevent division-by-zero errors. When calculating this function, the requirements of term professionalism, context consistency, and real-time are considered synchronously. By traversing all feasible paths in the knowledge graph, the path with the largest function value is selected as the optimization output.
9. The method for simultaneous interpretation of the OWS headset based on the ASR algorithm according to claim 1, wherein The enhancement processing of the speech feature signal includes a multi-stage noise reduction process: first, use the generalized cross-correlation algorithm to estimate the time difference of arrival of sound to achieve primary beamforming, then use a deep neural network for residual noise suppression, and finally restore the speech naturalness through a phase compensation algorithm. The deep neural network includes an encoder-decoder structure. The encoder maps the multi-channel speech features to the latent space, and the decoder generates a clean speech spectrogram in combination with the scene prior information provided by the knowledge graph. The adversarial training strategy is adopted during network training to improve the noise robustness.
10. The method for simultaneous interpretation of OWS headphones based on the ASR algorithm according to claim 1, wherein, The working process of the closed-loop optimization system includes: loading the general domain knowledge graph in the initial stage, continuously collecting directional speech signals and environmental context information during the operation process. When detecting a professional domain dialogue scene, automatically download the domain-enhanced graph from the cloud and perform incremental model fine-tuning, and at the same time establish a temporary buffer to store the session context features. At the end of the dialogue, update the local knowledge graph according to the optimized path weights, and clear the temporary cache data to protect user privacy.
Citation Information
Patent Citations
System and methods for maintaining speech-to-speech translation in the field
CN102084417A
Multi-mode intelligent question answering and recommending system supporting emotional speech output
CN119739840A
System, apparatus, and method for using a chatbot
US20240321279A1
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1
Contrastive learning with adversarial data for robust speech translation
US20240419927A1
Cited By
Multi-language voice content recognition method and system
CN121662048A
A method and system for multilingual speech content recognition
CN121662048B