OWS headphone simultaneous interpretation method based on ASR algorithm
By combining a distributed multi-microphone array and knowledge graph in open-back headphones, dynamically tracking the sound source and optimizing the speech recognition and translation paths, combined with sound field compensation and user feedback, the problems of low speech recognition accuracy and high delay in multilingual interpretation in complex sound field environments of open-back headphones are solved, achieving low-latency, high-quality simultaneous interpretation.
Patent Information
- Application Number
- CN202510784627.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Open-back headphones have low speech recognition accuracy in complex sound field environments and high delay in multilingual simultaneous interpretation. In addition, existing technologies lack multimodal feedback mechanisms and system adaptability, resulting in insufficient translation accuracy and real-time performance.
Through the distributed multi-microphone array combined with the adaptive beamforming algorithm, the sound source is dynamically tracked, a dynamic path planning module based on the knowledge graph is constructed, and the translation confidence association and context dependency between semantic nodes are used to generate multilingual optimized signals. The open speaker array is driven by the sound field compensation algorithm to output the translated speech in a direction, and the knowledge graph weight is optimized through the reinforcement learning mechanism of user feedback to form a closed-loop system.
It achieves low-latency, high-quality simultaneous interpretation in complex scenarios, improves speech recognition accuracy and translation accuracy, reduces latency, and enhances the system's adaptability and the real-time nature of user interaction.
Smart Images

Figure CN120356457B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of speech recognition technology, wearable devices and real-time translation systems, and in particular to an OWS headphone simultaneous interpretation method based on an ASR algorithm. Background Art
[0002] Traditional simultaneous interpretation systems often use a closed-back, in-ear design. While this design physically isolates ambient noise, long-term wear can cause ear discomfort and fail to meet safety requirements during exercise. Existing open-back headphones generally suffer from a low signal-to-noise ratio in voice acquisition, significantly reducing speech recognition accuracy in noisy environments. Mainstream speech recognition algorithms rely on cloud-based processing, resulting in high translation latency and difficulty meeting the demands of real-time conversations. Multilingual translation systems typically employ a fixed-path cascade processing model, lacking contextual awareness and prone to ambiguity when dealing with specialized terminology or cultural differences. Existing noise reduction algorithms are often based on static acoustic models and cannot adapt to changes in sound source position caused by head movement, resulting in target speech tracking failures. Regarding translation path optimization, traditional methods are limited to predefined rules or simple probabilistic models, making it difficult to dynamically adjust term priority and language switching strategies. The limitations of open-back speaker sound field control technology mean that the translated speech easily blends with ambient noise, affecting auditory clarity. User interaction methods are limited, lacking multimodal feedback mechanisms, making it impossible to correct translation errors or adjust output preferences in real time. Although current research has made some local improvements, such as using beamforming to improve voice collection quality, or introducing an attention mechanism to optimize the recognition model, it has not solved the systemic problem of multi-module collaborative optimization in an open sound field environment. Most solutions still have the following defects: the acoustic noise reduction and speech recognition modules are designed independently, resulting in feature transfer distortion; the translation path planning lacks knowledge graph support and lacks semantic coherence; the sound field compensation algorithm ignores dynamic changes in the environment, and the directional output has poor stability; the user feedback mechanism is disconnected from the core algorithm, and the system's adaptability is limited. This patent integrates acoustic structure innovation and algorithm optimization to build a full-link collaborative system from voice collection, semantic analysis to sound field output, to overcome the problem of real-time translation of open headphones in complex scenarios. Summary of the Invention
[0003] In view of the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide an OWS headphone simultaneous interpretation method based on the ASR algorithm, which is used to solve the problems of low speech recognition accuracy and high delay in multilingual simultaneous interpretation of open headphones in complex sound field environments. The present invention dynamically tracks the target sound source through head movement data to suppress environmental noise; constructs a dynamic path planning module based on the knowledge graph, and uses the translation confidence association and context dependency between semantic nodes to generate multilingual optimization signals to guide the lightweight hybrid speech recognition model to focus on key semantics; drives the open speaker array to output the translated speech in a directional manner through the sound field compensation algorithm, and continuously optimizes the knowledge graph weights in combination with the reinforcement learning mechanism of user feedback to form a closed-loop system that links environmental perception with user behavior, and finally achieves low-latency and high-quality simultaneous interpretation in complex scenarios. The present invention provides an OWS headphone simultaneous interpretation method based on the ASR algorithm, comprising:
[0004] S1: A distributed multi-microphone array is used to collect ambient speech signals and an adaptive beamforming algorithm is used to generate noise-reduced speech feature signals.
[0005] S2: Input the speech feature signal into the knowledge graph-based dynamic path planning module, and generate multilingual optimization signals through the translation confidence association and context dependency between semantic nodes;
[0006] S3: Uses optimized signals to control a lightweight hybrid ASR model, generating path-labeled text sequence signals through end-cloud collaborative processing. The end-side model adjusts the focus of the attention mechanism based on the dynamic weights in the knowledge graph.
[0007] S4: The text sequence signal is converted into target language text through the neural machine translation module, and the open speaker array is driven by the sound field compensation algorithm to generate directional speech signals;
[0008] S5: Receives user feedback signals through a multimodal interaction interface and uses a reinforcement learning mechanism to update path weights and semantic association parameter update signals in the knowledge graph. The translation confidence assessment strategy is dynamically modified based on the user's head posture change signal, and the parameter update signal is fed back to the dynamic path planning module of S2 to achieve closed-loop optimization.
[0009] In one embodiment of the present invention, in step S1, the distributed multi-microphone array is arranged on the edge of the earphone housing using a ring topology structure, each microphone unit is equipped with an independent acoustic pre-processing module, and the adaptive beamforming algorithm achieves dynamic noise reduction by iteratively optimizing the sound source positioning accuracy, specifically including: using the inertial measurement unit built into the earphone to obtain head motion acceleration data, coupling the motion acceleration data with the time delay estimation value of the microphone array to generate three-dimensional spatial sound source tracking parameters, which are used to adjust the spatial filtering weight matrix of the beamforming algorithm in real time, so that when the user's head is deflected, the main lobe direction of the beam is always aligned with the target speaker's direction, while suppressing reverberation noise and sudden interference sound sources from other directions.
[0010] In one embodiment of the present invention, the method for constructing a dynamic path planning module of the knowledge graph in step S2 includes: dividing multilingual semantic nodes into a three-layer structure of a core terminology layer, a context association layer, and a scene adaptation layer according to the domain knowledge graph, and realizing translation confidence association through a semantic similarity propagation algorithm, specifically: establishing a cross-language professional vocabulary mapping relationship in the core terminology layer, constructing a dialogue state transition probability matrix in the context association layer, introducing environmental acoustic features as graph neural network node attributes in the scene adaptation layer, and when a new speech feature signal is received, activating the corresponding sub-graph area according to the current dialogue scene, calculating the comprehensive score of each language path through the graph attention mechanism, and selecting the path with the highest score to generate an optimization signal.
[0011] In one embodiment of the present invention, the end-cloud collaborative processing flow of the lightweight hybrid model in step S3 includes: deploying a compressed speech recognition basic model on the end-side device, which uses knowledge distillation technology to extract key feature representation capabilities from the large cloud model; when professional domain terms or low-confidence recognition fragments are detected, the cloud-assisted verification mechanism is automatically triggered; the focus range of the attention mechanism is adjusted by performing importance sampling on the acoustic feature sequence in the time dimension and bandpass enhancement on the speech spectrum in the frequency dimension according to the dynamic weight coefficient provided by the knowledge graph, so that the model attention weight distribution keeps pace with the term priority of the current translation path.
[0012] In one embodiment of the present invention, the implementation method of the sound field compensation algorithm in step S4 includes: establishing an acoustic radiation model of an open speaker array, obtaining the optimal phase distribution parameters by solving the Helmholtz equation, and the generation process of the directional speech signal is specifically as follows: first, performing an emotional feature analysis on the target language text to determine the fundamental frequency fluctuation range, and then designing an adaptive filter group in combination with the spectral characteristics of the ambient noise, and finally calculating the driving signal delay difference based on the spatial position parameters of the speaker unit, so that the synthesized speech forms a constructive interference beam in the target direction and a destructive interference area in the non-target direction.
[0013] In one embodiment of the present invention, the reinforcement learning mechanism in step S5 adopts an online learning method based on policy gradient, and its policy network update process satisfies the following formula: Where θ represents the policy network parameters, represents the user's interactive action at time t, is the system state vector including the current translation quality evaluation value, environmental noise level and language switching frequency, Ψ( :T) is the cumulative reward function from time t to the end time T, which comprehensively considers the weighted scores of the three dimensions of translation accuracy, response delay and user satisfaction. The gradient calculated by this formula is used to update the connection weights of semantic nodes in the knowledge graph. In one embodiment of the present invention, the implementation method of the multimodal interactive interface includes: dividing the semantic perception partition in the touch area of the headset, defining the mapping relationship between the sliding trajectory pattern and the translation control instruction, and combining the head posture sensor data to realize the composite interaction logic, specifically: when it is detected that the user's finger is continuously sliding in a specific partition, a language switching request signal is generated; when the user nods, a translation confirmation signal is generated; when the user shakes his head, a current sentence retranslation mechanism is triggered, and all interaction signals are converted into update parameters of the knowledge graph through an event fusion algorithm.
[0014] In one embodiment of the present invention, the evaluation function used by the dynamic path planning module when generating the multilingual optimization signal is: in represents the domain relevance weight of the i-th semantic node, is the cross-language alignment confidence, is the historical translation error rate, Δτ is the estimated path delay, η is the delay sensitivity factor, and ε is a small constant to prevent division by zero errors. This function takes into account terminology expertise, context consistency, and real-time requirements during calculation. By traversing all feasible paths in the knowledge graph, it selects the path with the largest function value as the optimization output.
[0015] In one embodiment of the present invention, the enhancement processing of acoustic feature signals includes a multi-stage noise reduction process: first, the generalized cross-correlation algorithm is used to estimate the time difference of arrival to achieve primary beamforming, then a deep neural network is used to suppress residual noise, and finally, a phase compensation algorithm is used to restore the naturalness of the speech, wherein the deep neural network includes an encoder-decoder structure, the encoder maps multi-channel speech features to the latent space, and the decoder combines the scene prior information provided by the knowledge graph to generate a pure speech spectrogram, and an adversarial training strategy is adopted during network training to improve noise robustness.
[0016] In one embodiment of the present invention, the workflow of the closed-loop optimization system includes: loading a general domain knowledge graph in the initial stage, continuously collecting user feedback signals and environmental context information during operation, automatically downloading the domain enhancement graph from the cloud and performing incremental model fine-tuning when a professional domain dialogue scenario is detected, and establishing a temporary cache area to store session context features. At the end of the dialogue, the local knowledge graph is updated according to the optimized path weight, and the temporary cache data is cleared to protect user privacy.
[0017] The OWS headphone simultaneous interpretation method based on the ASR algorithm provided by the present invention dynamically tracks the target sound source through head movement data and suppresses environmental noise; constructs a dynamic path planning module based on the knowledge graph, utilizes the translation confidence association and context dependency between semantic nodes to generate multilingual optimization signals, and guides the lightweight hybrid speech recognition model to focus on key semantics; drives the open speaker array to output the translated speech in a directional manner through the sound field compensation algorithm, and continuously optimizes the knowledge graph weights in combination with the reinforcement learning mechanism of user feedback, forming a closed-loop system that links environmental perception with user behavior, and ultimately achieves low-latency and high-quality simultaneous interpretation in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 The figure is a flow chart of the OWS headphone simultaneous interpretation method based on the ASR algorithm. DETAILED DESCRIPTION
[0020] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.
[0021] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.
[0022] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.
[0023] See Figure 1 , which shows the OWS headphone simultaneous interpretation method based on ASR algorithm of the present invention. The OWS headset simultaneous interpretation method based on the ASR algorithm of the present invention includes five steps: S1: collecting environmental voice signals through a distributed multi-microphone array, and generating a noise-reduced voice feature signal using an adaptive beamforming algorithm; S2: inputting the voice feature signal into a dynamic path planning module based on a knowledge graph, and generating a multilingual optimization signal through the translation confidence association and context dependency between semantic nodes; S3: using the optimization signal to control a lightweight hybrid ASR model, and generating a text sequence signal with path markings through end-cloud collaborative processing, wherein the end-side model adjusts the focus range of the attention mechanism according to the dynamic weight in the knowledge graph; S4: converting the text sequence signal into target language text through a neural machine translation module, and driving an open speaker array to generate a directional voice signal through a sound field compensation algorithm; S5: receiving the user feedback signal through a multimodal interaction interface, and updating the path weight and semantic association parameter update signal in the knowledge graph using a reinforcement learning mechanism, wherein the translation confidence evaluation strategy is dynamically corrected according to the user's head posture change signal, and the parameter update signal is fed back to the dynamic path planning module of S2 to achieve closed-loop optimization.
[0024] like Figure 1As shown, the simultaneous interpretation method is based on the co-design of the physical structure and algorithms of open-ear headphones. Its core process begins with signal acquisition using a distributed multi-microphone array. The array adopts a ring topology, with microphone units spaced equidistantly around the edge of the headphone housing, creating a 360-degree sound coverage area. Each microphone in the array is equipped with an independent acoustic pre-processing module, including a pre-emphasis filter, automatic gain control circuitry, and an analog-to-digital converter, to ensure the integrity of the original speech signal in both the time and frequency domains. The adaptive beamforming algorithm achieves noise suppression through spatial filtering techniques. Its core is to dynamically construct a directional beam, aligning the main lobe toward the target sound source while suppressing interfering noise in the sidelobe direction. The algorithm first uses the generalized cross-correlation method to calculate the time delay difference between each microphone. This is combined with head acceleration data acquired by the inertial measurement unit to construct a three-dimensional sound source tracking model. When the user's head deflects, the system calculates the new sound source azimuth in real time and updates the beamforming weight matrix accordingly to maintain continuous tracking of the target speaker. During this process, the weight matrix is optimized to minimize speech distortion under noise interference. Regularization constraints are used to prevent overfitting, ultimately outputting a feature signal containing pure speech components. After generating the speech feature signal, the system inputs it into a knowledge graph-based dynamic path planning module. This module is built on a three-layer graph structure: a core terminology layer stores cross-language vocabulary mappings and domain attribute labels; a contextual association layer records semantic dependencies and state transition probabilities from conversation history; and a scene adaptation layer integrates environmental acoustic features and user interaction data. Upon receiving new speech features, the module first performs semantic role labeling and intent identification, activating subgraph regions in the relevant domain. A graph attention mechanism then calculates a confidence score for each language path. The scoring factors include term alignment accuracy, contextual coherence index, and real-time environmental noise impact factors. In implementation, the system uses a random walk algorithm to traverse feasible paths, dynamically adjusting transition probabilities between nodes during each walk, prioritizing edges with high-confidence connections. The final optimized signal includes target language selection parameters, terminology library priority weights, and context association rules. These parameters will serve as the decision basis for subsequent speech recognition and translation processes.
[0025] Further details are provided for the implementation of the distributed microphone array and adaptive beamforming algorithm in step S1. The array's ring topology design offers the advantage of balanced acoustic response characteristics in all directions. The spacing between each microphone unit is determined by calculating the Helmholtz resonance frequency to avoid phase cancellation in specific frequency bands. The acoustic preprocessing module comprises a multi-stage signal conditioning circuit: the first stage uses a programmable gain amplifier to dynamically adjust the sensitivity of each channel to eliminate signal amplitude deviation caused by varying wearing angles; the second stage applies an anti-aliasing filter bank that dynamically switches the cutoff frequency based on the speech signal bandwidth; and the third stage implements clock-synchronized sampling to ensure time alignment of multi-channel signals with microsecond accuracy. The adaptive beamforming algorithm achieves dynamic noise reduction through iterative optimization. Its core lies in establishing joint spatial and temporal constraints: linear prediction residual analysis is used in the time domain to distinguish speech and noise components, and eigendecomposition of the covariance matrix is used in the spatial domain to extract the dominant noise direction. During algorithm execution, the noise covariance matrix estimate for the current frame is first calculated. Regularization constraints are then constructed using the historical speech source estimate matrix. The optimal beam weights are then solved through convex optimization. In particular, the system introduces a head motion compensation mechanism, which converts the three-axis angular velocity data output by the inertial measurement unit into a spatial rotation matrix, and corrects the calculated azimuth angle of the sound source in real time, ensuring that the target sound source can still be stably tracked when the user turns his head quickly.
[0026] like Figure 1Figure 2 shows the construction and operation mechanism of the knowledge graph dynamic path planning module. The core terminology layer is constructed using a domain-adaptive semi-supervised learning approach: high-frequency terms are extracted from a professional corpus to construct initial nodes. Multilingual mappings are established using cross-lingual word vector alignment technology. Each node is labeled with domain attributes (such as medical, legal, engineering, etc.) and a timeliness marker. The context association layer is modeled using a temporal graph neural network. Nodes represent semantic units in a conversation, and edge weights are composed of two components: statistical co-occurrence probabilities and deep learning-based context prediction scores. The key innovation of the scene adaptation layer lies in the integration of multimodal perception data: ambient noise spectral features are encoded as graph node attributes using Mel-frequency cepstral coefficients, and user interaction behavior data (such as language switching frequency and terminology revision history) is converted into edge weight correction factors. When new speech features are input, the system performs multi-level graph reasoning: first, precise matching is performed at the core terminology layer to identify professional terms and their associated domains; then, probabilistic reasoning is performed at the context association layer to predict the potential direction of the current conversation; finally, the scene adaptation layer adjusts the path score based on environmental and user factors. The graph attention mechanism is implemented using a multi-head attention architecture, with each attention head focusing on related features across different dimensions: the first head calculates a term's expertise score, the second assesses contextual coherence, and the third quantifies the degree of real-time environmental interference. The system then weights the outputs of each head to derive a final path score and selects the path with the highest overall score to generate the optimization signal. The lightweight hybrid model utilizes a cloud-end collaborative architecture. The compression model deployed on-device is built using knowledge distillation technology: a base model with a deep bidirectional Transformer architecture is first trained on the cloud. Inter-layer attention transfer is then used to extract key feature representation capabilities, resulting in a lightweight model with an 80% parameter reduction. This model utilizes dynamic computational graph technology to adaptively adjust network depth based on the complexity of the input speech: when clearly pronounced common words are detected, only shallow network layers are activated; when ambiguous pronunciations or specialized terminology are encountered, deeper layers are activated layer by layer. The attention mechanism's focus adjustment strategy encompasses two dimensions: In the temporal dimension, importance sampling is performed on speech frames based on the term priority coefficients provided by the knowledge graph, prioritizing acoustic features in high-weighted intervals. In the frequency dimension, a bandpass filter is dynamically generated based on the spectral characteristics of the ambient noise, enhancing the energy distribution in key speech frequency bands. The cloud-assisted verification mechanism utilizes a differentially private data transmission protocol. When the on-device model detects a low-confidence recognition segment, it automatically extracts and uploads encrypted feature vectors to the cloud. The cloud-based model then returns the correction results and updates the on-device model's parameters. This collaborative mechanism improves the recognition accuracy of specialized terminology by approximately 40% while ensuring privacy.
[0027] In one embodiment of the present invention, a dynamic path planning module and a speech recognition module form a bidirectional data flow: on the one hand, the path optimization signal guides the ASR model to adjust its attention distribution, focusing on the key semantic units of the current translation path; on the other hand, the recognition results output by the ASR are fed back into the knowledge graph to update the state transition probabilities of the context association layer. This closed-loop interaction mechanism enables the system to adapt to dynamic changes in the conversational scenario. For example, when the user frequently switches languages, the delay sensitivity factor in the path evaluation function is automatically increased; when the environmental noise suddenly increases, the decision weight of the core term layer is temporarily increased. Experiments have shown that this architecture improves the semantic coherence score by 35% compared to traditional methods in noisy environments, and reduces translation latency to less than 800 milliseconds.
[0028] Furthermore, the collaborative design of a distributed microphone array and an adaptive algorithm overcomes the limitations of traditional noise reduction solutions. By integrating physical acoustic modeling with deep learning techniques, the system achieves noise suppression at three levels: the physical layer utilizes beamforming to eliminate spatially separated noise; the algorithmic layer uses deep neural networks to suppress residual steady-state noise; and the application layer integrates knowledge graph information for semantic-level noise filtering. A specially designed phase compensation algorithm restores speech naturalness lost during noise reduction. This algorithm analyzes the phase spectrum characteristics of pure speech and constructs a phase reconstruction model to ensure that the output speech retains the rhythmic characteristics of the original pronunciation. The directional output technology of the open speaker array utilizes a parametric sound field synthesis method. By solving wave equations to obtain the optimal driving signal parameters, this method simultaneously forms a high-energy speech beam in the target direction and creates a sound wave cancellation zone in other directions, effectively reducing the aliasing interference between the translated speech and ambient noise. Simultaneous interpretation methods employ a collaborative working mechanism between the sound field compensation algorithm and a neural machine translation module in the speech synthesis and output stages. The neural machine translation module uses a Transformer-based encoder-decoder architecture. After receiving a text sequence signal with path markers, the encoder first analyzes the term priority and contextual association rules in the path markers and inputs them into the model as extended features of the positional encoding. When the decoder generates the target language text, it dynamically adjusts the attention distribution based on the domain adaptation parameters provided by the knowledge graph, prioritizing the professional vocabulary mapping relationships at the core term layer. The core of the sound field compensation algorithm lies in constructing an acoustic radiation model of an open speaker array and obtaining the optimal phase distribution by solving the boundary conditions of the Helmholtz equation. In specific implementation, the system first analyzes the spectral characteristics of the ambient noise and extracts the main interfering frequency band information. Then, based on the emotional characteristics of the target speech (such as speaking rate and intonation), it determines the fundamental frequency fluctuation range to generate an initial speech signal with a natural rhythm. The drive signal generation process for a loudspeaker array involves three key steps: first, calculating the sound wave propagation path difference for each element based on the polar coordinate position parameters of the array elements; second, designing an adaptive filter bank based on the ambient noise spectrum to enhance the energy of specific frequency bands of the target speech; and third, adjusting the phase of the drive signal for each element through a delay control module so that the synthesized sound waves form a constructive interference beam in the target direction and destructive interference regions to the sides and rear. During this process, the system monitors changes in the ambient sound field in real time and dynamically updates the phase compensation parameters using a feedback learning algorithm. For example, when a sudden high-frequency noise is detected, the system automatically enhances the energy distribution of the speech signal in the 2-4kHz frequency band and adjusts the beamwidth of the loudspeaker array to improve interference rejection. Experiments have shown that this sound field compensation technology can improve speech intelligibility in the target direction by over 50% while reducing the sound pressure level in non-target directions to below the ambient noise level.
[0029] like Figure 1As shown, the reinforcement learning mechanism uses an asynchronous policy gradient algorithm to continuously optimize the knowledge graph. The policy network architecture consists of two parallel sub-networks: a state feature extraction network and an action value evaluation network. The state feature extraction network receives a multi-dimensional input vector, including real-time translation quality scores (based on semantic similarity), ambient noise levels (obtained through spectral analysis), user interaction frequency (such as the number of language switches), and historical path selection records. This network extracts spatiotemporal correlation features through multiple layers of convolution and gated recurrent units, generating a 128-dimensional state embedding vector. The action value evaluation network calculates the probability distribution of possible actions based on the current state vector, where the action space is defined as the direction and magnitude of the adjustment of the connection weights of knowledge graph nodes. The cumulative reward function in the policy gradient update formula is designed to include three dimensions: a translation accuracy reward based on the edit distance between the recognition result and the human annotation; a response latency reward using a negative exponential function to map processing time differences; and a user satisfaction reward quantified by analyzing interaction behavior patterns (such as confirmation signal frequency and number of retranslation requests). The system performs a policy update at the end of each conversation, using importance sampling to balance exploration and exploitation, ensuring that the knowledge graph can quickly adapt to new scenarios while not deviating too much from proven effective paths. For example, in a medical consultation scenario, when the system detects that a user repeatedly repeats a specific disease term, it automatically increases the weight of the relevant term node in the core layer and strengthens the confidence score of its cross-language mapping relationship. The multimodal interactive interface uses a layered event processing architecture to accurately interpret user intent. The touch area is divided into three functional zones: the front zone maps language switching instructions, the back zone is associated with translation mode selection, and the top zone is associated with volume adjustment. Each zone is covered with an array of pressure-sensitive sensors that detect the speed, direction, and contact area of the sliding trajectory. The interaction logic is determined using a joint spatiotemporal analysis method: in the temporal dimension, a long short-term memory network is used to identify the temporal pattern of continuous gestures; in the spatial dimension, a convolutional neural network is used to analyze the geometric features of the touch trajectory. When fusion processing the head posture sensor data, a quaternion representation is used to calculate the three-dimensional head rotation angle, combined with a Kalman filter to eliminate motion noise. When the system detects a nodding motion (a pitch angle change exceeding 15 degrees), the translation confirmation process is triggered, marking the current translation as a high-confidence sample for model fine-tuning. When a head shake motion is identified (a yaw angle swing greater than 20 degrees), a three-level retranslation mechanism is initiated: the first level calls the local cache to regenerate the translation, the second level requests cloud-assisted translation, and the third level activates the manual review channel. All interactive events are managed through a priority queue, and urgent instructions (such as mistranslation of safety warning terms) can interrupt the current processing flow and receive priority response. The core of the event fusion algorithm lies in constructing a multimodal signal correlation matrix and calculating the confidence weight of each signal source through an attention mechanism. The final knowledge graph update parameters include semantic node weight correction, edge connection strength adjustment factor, and scenario adaptation coefficient.The dynamic path evaluation function uses a multi-objective optimization framework to balance translation quality and real-time requirements. The term expertise weight ω_i in this function is calculated based on the term's TF-IDF value in the domain knowledge base, its recent frequency of use, and the importance of user annotations. The domain relevance component utilizes a graph neural network-based representation learning method, embedding term nodes in a low-dimensional space and calculating their cosine similarity with the current conversation scenario. The cross-language alignment confidence ξ_i is evaluated in two parts: the static component is derived from the pre-trained multilingual word vector alignment score, while the dynamic component is adjusted by real-time monitoring of the co-occurrence probability of bilingual terms in the context window. The historical translation error rate γ_i is calculated using a sliding window exponential decay method, with recent errors given higher weight. A correction feedback gain factor is also introduced, so that when users actively correct errors, the error rate of related terms decays superlinearly. The path delay estimate Δτ is modeled based on computing resource allocation and network transmission quality, using the M / M / C model from queuing theory to predict the time consumption of each processing step. The dynamic adjustment strategy for the latency sensitivity factor η is divided into a basic mode and an emergency mode. During regular conversations, it increases slowly according to an exponential law. When the user's speech speed or interaction frequency increases, it switches to a linear, rapid growth mode. During function calculation, the system maintains a priority queue to store candidate paths and uses branch-and-bound to quickly prune low-scoring paths while retaining suboptimal solutions as disaster recovery backups. When the combined score of the optimal path falls below a safety threshold, a cross-modal verification process is automatically triggered, coordinating the speech recognition, environmental perception, and user feedback modules to make a joint decision.
[0030] Furthermore, the speaker array design transcends the physical limitations of traditional open-back headphones. The array unit utilizes a contoured diaphragm structure, optimized through finite element analysis to achieve a flat frequency response across the 300Hz-5kHz frequency range. The driver circuit design combines a digital amplifier with a brain-inspired computing chip. The digital amplifier provides basic signal amplification, while the brain-inspired chip implements real-time parameter adjustment via a pulsed neural network. The phase compensation algorithm's innovation lies in the recursive prediction of the ambient sound field: First, an auxiliary microphone captures residual noise signals. Then, an improved LMS adaptive filter is used to estimate the sound field transfer function. Finally, the driving signal phase is proactively adjusted based on the predicted model. This proactive compensation strategy accelerates the response speed of acoustic wave cancellation to the 10-millisecond level, effectively addressing rapidly changing ambient noise levels. Experimental data shows that, in a 70dB background noise environment, this technology maintains a speech signal-to-noise ratio (SNR) above 15dB in the target direction. The optimization of the neural machine translation module is reflected in its combination of domain adaptation and real-time learning capabilities. The module incorporates a built-in dual buffering mechanism: a primary buffer stores the contextual feature vectors of the current conversation, while a secondary buffer retains translations of recent high-frequency terms. When a specialized domain term is detected, a term-prioritization process is initiated: first, the most recently used record is retrieved from the secondary cache. If a high-confidence translation exists, it is directly called; otherwise, the domain adaptation sub-network is activated, which improves terminology accuracy by fine-tuning the key-value matrix of the attention layer. A confidence heat map is generated synchronously during the translation process, and low-confidence segments automatically trigger a multi-strategy fallback mechanism, including synonym replacement, word order adjustment, and contextual association expansion. Data interaction between the sound field compensation algorithm and the translation module is achieved through a shared hidden layer representation. The emotional feature vector output by the translation module directly guides the generation of sound field synthesis parameters, such as automatically increasing the upward amplitude of the fundamental frequency curve at the end of a question sentence.
[0031] The reinforcement learning system design emphasizes security and explainability. A constrained optimization framework is introduced during policy network training, using the Lagrange multiplier method to ensure that the weights of key connections in the knowledge graph do not mutate suddenly. The reward function design includes a regularization term to prevent over-optimization for short-term gains that could compromise system stability. Network updates employ a dual-objective optimization strategy: the primary objective is to maximize cumulative rewards, while the secondary objective is to minimize policy volatility. Blockchain technology is used for version management of the knowledge graph, with each update generating an immutable record to facilitate error tracing. When a vicious feedback loop is detected (e.g., continuously reducing the weight of a term, resulting in a semantic gap), the system automatically rolls back to the most recent stable version and issues a maintenance alert. The fault-tolerance mechanism for the multimodal interactive interface includes a multi-level verification process. When parsing touch commands, an ensemble learning approach integrates the results of multiple sensors: the pressure distribution pattern of the pressure sensor, the contact area change trend of the capacitive sensor, and the proximity detection data of the infrared sensor are fed into a random forest classifier to reduce false touches. Head posture recognition incorporates kinematic model constraints, comparing the raw sensor data with the biomechanical characteristics of human neck movement to filter out unreasonable and abnormal movements. The event fusion algorithm utilizes a federated learning framework, improving the accuracy of interaction intent recognition through multi-device collaborative training while protecting user privacy. The system also features an emergency bypass channel. When persistent abnormal interaction signals are detected, it automatically switches to basic translation mode and initiates a self-diagnosis procedure.
[0032] Specifically, the computational efficiency of the dynamic path evaluation function has been significantly improved through hardware acceleration. The system is equipped with a dedicated graph computing coprocessor and uses a parallel edge traversal algorithm to accelerate the search process of the knowledge graph. For large-scale graph scenarios, a hierarchical pruning strategy is developed: the first layer quickly filters irrelevant subgraphs based on domain labels, the second layer applies an approximate algorithm to estimate the upper limit of the path score, and the third layer accurately calculates the candidate paths. The calculation of the exponential term in the function is optimized by combining a lookup table with a Taylor expansion, which reduces the computational time by 60% while ensuring accuracy. Real-time performance is also reflected in the incremental update mechanism. When new speech features arrive, only the affected part of the path is re-evaluated instead of traversing the entire graph.
[0033] The OWS headphone simultaneous interpretation method based on the ASR algorithm of the present invention dynamically tracks the target sound source through head movement data to suppress environmental noise; constructs a dynamic path planning module based on the knowledge graph, utilizes the translation confidence association and context dependency between semantic nodes to generate multilingual optimization signals, and guides the lightweight hybrid speech recognition model to focus on key semantics; drives the open speaker array to output the translated speech in a directional manner through the sound field compensation algorithm, and continuously optimizes the knowledge graph weights in combination with the reinforcement learning mechanism of user feedback, forming a closed-loop system that links environmental perception with user behavior, and ultimately achieves low-latency and high-quality simultaneous interpretation in complex scenarios.
[0034] Therefore, the OWS headphone simultaneous interpretation method based on the ASR algorithm of the present invention can solve the problems of low speech recognition accuracy and high delay of multilingual simultaneous interpretation of open headphones in complex sound field environments.
[0035] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.
Claims
1. The OWS headphone simultaneous interpretation method based on ASR algorithm is characterized by: include: S1: A distributed multi-microphone array is used to collect ambient speech signals and an adaptive beamforming algorithm is used to generate noise-reduced speech feature signals. S2: Inputting the speech feature signal into a dynamic path planning module based on a knowledge graph, and generating a multilingual optimization signal through the translation confidence association and context dependency between semantic nodes; S3: Use the optimized signal to control the lightweight hybrid ASR model, and generate a text sequence signal with path labels through end-cloud collaborative processing. The end-side model adjusts the focus range of the attention mechanism based on the dynamic weights in the knowledge graph. S4: converting the text sequence signal into a target language text through a neural machine translation module, and driving an open speaker array to generate a directional speech signal through a sound field compensation algorithm; S5: Receive the directional speech signal through the multimodal interaction interface, and use the reinforcement learning mechanism to update the path weight and semantic association parameter update signal in the knowledge graph. The translation confidence assessment strategy is dynamically modified according to the user's head posture change signal. The parameter update signal is fed back to the dynamic path planning module of S2 to achieve closed-loop optimization. The reinforcement learning mechanism adopts an online learning method based on policy gradient, and its policy network update process satisfies the following formula: Where θ represents the policy network parameters, represents the user's interactive action at time t, is the system state vector including the current translation quality evaluation value, environmental noise level and language switching frequency, Ψ( :T) is the cumulative reward function from time t to the termination time T, which comprehensively considers the weighted scores of three dimensions: translation accuracy, response delay, and user satisfaction. The gradient calculated by this formula is used to update the connection weights of semantic nodes in the knowledge graph.
2. The OWS headphone simultaneous interpretation method based on the ASR algorithm according to claim 1, characterized in that: The distributed multi-microphone array described in step S1 is arranged on the edge of the earphone housing in a ring topology structure, and each microphone unit is equipped with an independent acoustic pre-processing module. The adaptive beamforming algorithm achieves dynamic noise reduction by iteratively optimizing the sound source positioning accuracy, specifically including: using the inertial measurement unit built into the earphone to obtain head motion acceleration data, coupling the motion acceleration data with the time delay estimation value of the microphone array to generate three-dimensional spatial sound source tracking parameters, which are used to adjust the spatial filtering weight matrix of the beamforming algorithm in real time, so that when the user's head deflects, the main lobe direction of the beam is always aligned with the target speaker's direction, while suppressing reverberation noise and sudden interference sound sources from other directions.
3. The OWS earphone simultaneous interpretation method based on the ASR algorithm according to claim 1, characterized in that: The method for constructing a dynamic path planning module of the knowledge graph described in step S2 includes: dividing multilingual semantic nodes into a three-layer structure of a core terminology layer, a context association layer, and a scene adaptation layer according to the domain knowledge graph; the translation confidence association is realized through a semantic similarity propagation algorithm, specifically: establishing a cross-language professional vocabulary mapping relationship in the core terminology layer, constructing a dialogue state transition probability matrix in the context association layer, introducing environmental acoustic features as graph neural network node attributes in the scene adaptation layer; when a new speech feature signal is received, activating the corresponding sub-graph area according to the current dialogue scene, calculating the comprehensive score of each language path through the graph attention mechanism, and selecting the path with the highest score to generate an optimization signal.
4. The OWS earphone simultaneous interpretation method based on the ASR algorithm according to claim 1, characterized in that: The end-cloud collaborative processing flow of the lightweight hybrid ASR model described in step S3 includes: deploying a compressed speech recognition basic model on the end-side device, which uses knowledge distillation technology to extract key feature representation capabilities from the large cloud model. When professional domain terms or low-confidence recognition fragments are detected, the cloud-assisted verification mechanism is automatically triggered. The method for adjusting the focus range of the attention mechanism is: according to the dynamic weight coefficient provided by the knowledge graph, the acoustic feature sequence is importance sampled in the time dimension, and the speech spectrum is band-pass enhanced in the frequency dimension, so that the model attention weight distribution keeps pace with the term priority of the current translation path.
5. The OWS earphone simultaneous interpretation method based on ASR algorithm according to claim 1, characterized in that: The implementation method of the sound field compensation algorithm described in step S4 includes: establishing a sound radiation model of an open speaker array, obtaining the optimal phase distribution parameters by solving the Helmholtz equation, and the specific process of generating the directional speech signal is: first, performing an emotional feature analysis on the target language text to determine the fundamental frequency fluctuation range, then designing an adaptive filter group based on the spectral characteristics of the ambient noise, and finally calculating the driving signal delay difference based on the spatial position parameters of the speaker unit, so that the synthesized speech forms a constructive interference beam in the target direction and a destructive interference area in the non-target direction.
6. The OWS earphone simultaneous interpretation method based on the ASR algorithm according to claim 1, characterized in that: The implementation method of the multimodal interaction interface includes: dividing the semantic perception partition in the earphone touch area, defining the mapping relationship between the sliding trajectory pattern and the translation control instruction, and combining the head posture sensor data to realize the composite interaction logic. Specifically, when it is detected that the user's finger is continuously sliding in a specific partition, a language switching request signal is generated; when the user nods, a translation confirmation signal is generated; when the user shakes his head, a current sentence retranslation mechanism is triggered. All interaction signals are converted into update parameters of the knowledge graph through an event fusion algorithm.
7. The OWS earphone simultaneous interpretation method based on the ASR algorithm according to claim 1, characterized in that: The evaluation function used by the dynamic path planning module when generating multilingual optimization signals is: in represents the domain relevance weight of the i-th semantic node, is the cross-language alignment confidence, is the historical translation error rate, Δτ is the estimated path delay, η is the delay sensitivity factor, and ε is a small constant to prevent division by zero errors. This function takes into account terminology expertise, context consistency, and real-time requirements during calculation. By traversing all feasible paths in the knowledge graph, it selects the path with the largest function value as the optimization output.
8. The OWS earphone simultaneous interpretation method based on the ASR algorithm according to claim 1, characterized in that: The enhanced processing of the speech feature signal includes a multi-stage noise reduction process: first, the generalized cross-correlation algorithm is used to estimate the time difference of arrival to achieve primary beamforming, then a deep neural network is used to suppress residual noise, and finally the naturalness of the speech is restored through a phase compensation algorithm. The deep neural network includes an encoder-decoder structure, the encoder maps multi-channel speech features to a latent space, and the decoder combines the scene prior information provided by the knowledge graph to generate a pure speech spectrogram. An adversarial training strategy is used during network training to improve noise robustness.
9. The OWS earphone simultaneous interpretation method based on the ASR algorithm according to claim 1, characterized in that: The workflow of the closed-loop optimization system includes: loading a general domain knowledge graph in the initial stage, continuously collecting directional voice signals and environmental context information during operation, automatically downloading the domain enhancement graph from the cloud and performing incremental model fine-tuning when a professional domain dialogue scenario is detected, and establishing a temporary cache area to store conversation context features. At the end of the conversation, the local knowledge graph is updated according to the optimized path weights, and the temporary cache data is cleared to protect user privacy.
Citation Information
Patent Citations
Multi-mode intelligent question answering and recommending system supporting emotional speech output
CN119739840A
System, apparatus, and method for using a chatbot
US20240321279A1