A textless cross-language interaction method and system based on the inherent sound wave characteristics of human voice
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-11
AI Technical Summary
针对现有跨语种交互技术依赖语义解析、算法推理、文本中转的行业痛点,为解决传统技术延迟高、算力消耗大、保密性弱、环境抗干扰能力差、用户存在外语能力门槛的缺陷,本发明提出全新技术方案
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of physical acoustic signal processing, biological auditory native communication, and medium-free cross-language real-time interactive technology. This invention is a completely independent and original technical solution, which does not rely on, reuse, or continue any prior patented technical logic, parameter system, or framework structure. This invention completely abandons the traditional translation methods that rely on text encoding, semantic analysis, artificial intelligence models, and corpus databases. Instead, it uses the physical morphology substitution principle based on the inherent sound wave characteristics of human voices to achieve a pure physical direct conversion of sound waves from any foreign language to equivalent sound waves in the listener's native language. This represents a completely new track of underlying acoustic interaction innovation. Background Technology
[0002] Currently, there are only two main approaches to cross-language voice communication technology globally, both of which suffer from inherent flaws in their underlying logic that cannot be fixed: The first type is the traditional multi-level translation architecture, which relies on a chain of "speech-to-text—semantic parsing—text translation—speech synthesis" to complete the conversion. Its core drawback is that semantic understanding must be completed before speech reconstruction, resulting in problems such as a long chain, high latency, large signal loss, strong dependence on computing power and corpus, easy failure in noisy environments, and poor confidentiality due to the retention of interactive data. The second category is end-to-end speech-to-text technology. Although it eliminates explicit text-based translation, it still employs AI semantic acoustic coding, neural network probability fitting, and context vector matching logic at its underlying level. Essentially, it is still a secondary generation of speech after the machine parses the meaning of language, rather than a native physical replacement of sound waves. It generally suffers from defects such as model errors, limited context adaptation, poor accent adaptation, and insufficient coverage of less common languages. Existing cross-language voice interaction technologies all follow the inherent technical logic of "understanding semantics first, then converting language." There are currently no publicly available technologies, patents, or literature disclosing a technical solution that achieves cross-language conversion without relying on semantic parsing, but solely through the mapping of sound wave physical features. This invention effectively fills the core technological gap in this field. Summary of the Invention
[0003] 1. Purpose of the invention To address the industry pain points of existing cross-language interaction technologies that rely on semantic parsing, algorithmic reasoning, and text relay, and to solve the shortcomings of traditional technologies such as high latency, high computing power consumption, weak confidentiality, poor environmental interference resistance, and the foreign language proficiency barrier for users, this invention proposes a brand-new technical solution. This invention provides a purely physical sound wave closed-loop substitution scheme that achieves a complete closed loop of sound wave input from any foreign language, pure acoustic feature mapping conversion, and equivalent native sound wave output in the listener's native language. The entire process requires no semantic parsing, no text relay, no artificial intelligence computation, and no traditional translation logic, enabling the human ear to directly and natively understand foreign language speech. 2. Core Complete Closed-Loop Technology Principle The communication barrier between different languages is not a semantic difference, but rather a mismatch between the physical forms of sound waves in different languages and human auditory recognition habits. Each language possesses a stable and unique inherent sound wave morphology; gender, age, speech rate, emotion, and accent of a human voice are all temporary fluctuations and do not change the inherent acoustic characteristics of the language. The core closed-loop logic of this invention is as follows: acquire original sound wave signals of any foreign language, remove individual vocal interference, purify stable inherent acoustic base features of the language, accurately map and replace the acoustic base features of the foreign language with the standard acoustic features of the listener's preset mother tongue, physically reshape and generate an equivalent sound wave waveform of the mother tongue that is adapted to human ear recognition, and finally amplify the sound wave to achieve native human ear recognition without translation or semantic understanding. The feature replacement only changes the physical form of the sound wave to match the listener's native language hearing habits, fully preserving all the information content carried by the sound wave, without changing the information attributes of the original speech. This invention supports a two-way interactive closed loop: when an external user speaks, the device replaces the external language sound wave with the local user's native language sound wave output; when a local user speaks, the device replaces the local language sound wave with the corresponding external user's native language sound wave output. Both two-way interactions employ a purely physical waveform replacement logic from external sound waves to the listener's native language equivalent sound wave, with zero semantic involvement throughout the entire process. 3. Implementation steps of complete closed-loop technology S1, Acquisition and purification of human voice signals in the sound field Using sound wave acquisition devices (including but not limited to external hardware or the built-in microphone of a general-purpose computing device), natural human voices in any foreign language in the environment are acquired in real time. A precise filtering algorithm is then used to remove environmental noise, background noise, and various invalid interferences, preserving the pure and complete original sound wave signal of the foreign language. This process does not involve signal transcoding, generating text data, or extracting and parsing semantic information. S2. Individual interference removal and purification of language-specific acoustic base features The collected pure external sound waves are normalized and regularized to remove the floating interference caused by individual vocalizations such as gender, age, speech rate, emotion, and accent, and to extract stable and semantically independent language-specific acoustic base features. S3, Cross-linguistic Acoustic Feature Closed-Loop Mapping Permutation Based on a pre-built multilingual acoustic feature matching library, the inherent acoustic base features of foreign languages are accurately converted, mapped, and replaced into standard acoustic feature parameters of the listener's preset native language. The multilingual acoustic feature matching library is obtained by collecting standard pronunciation samples of multiple languages, statistically analyzing the distribution of inherent acoustic features of each language, and constructing feature space mapping relationships between languages, without the need for semantic annotation throughout the process. S4, Original Reconstruction of Sound Waves Equivalent to the Listener's Native Language Based on the standard acoustic characteristic parameters of the native language, the equivalent sound wave waveform that conforms to the listener's native language hearing habits is reconstructed at the physical level, generating a natural and realistic human voice vibration pattern without machine synthesis distortion problems. S5, native acoustic wave direct output, closed-loop interaction Through the directional sound wave output module, the high-fidelity external amplifier reconstructs the equivalent sound wave of the native language, enabling the human ear to directly recognize the information carried by the sound wave based on its native hearing. This achieves real-time recognition of foreign language speech and completes a closed loop of real-time cross-language sound wave interaction without text, semantics, or intelligent computation. 4. System Overall Architecture This system consists of five main functional modules. The system does not include text processing, semantic analysis, or artificial intelligence model computation units. The specific modules are as follows: 1. Multilingual sound wave acquisition and purification module: used to acquire sound wave signals in any external language, filter out environmental interference, and purify the original sound waves; 2. Language-specific acoustic feature purification module: used to remove individual vocal interference and accurately locate stable language-specific acoustic features; 3. Cross-language native language mapping and replacement module: used to achieve accurate conversion of acoustic features from foreign languages to standard acoustic features of the listener's native language; 4. Native Language Equivalent Sound Wave Reconstruction Module: Used to generate native language sound waves that can be directly recognized by the human ear through physical waveform synthesis technology; 5. Directional native sound wave output module: used for external playback of reconstructed native language sound waves to complete the auditory interaction loop. 5. Beneficial effects of the technology • Low-latency synchronous interaction: The entire process involves no semantic computation, no model inference, and no text encoding or decoding operations. It only completes sound wave feature replacement and waveform reconstruction, achieving a response speed at the level of human voice synchronization, thus realizing real-time interaction. • Strong anti-interference capability in all scenarios: Based on the stable inherent acoustic characteristics of the language, it is not affected by noisy environments, multi-person sound fields, human speech speed and accents, and its operation stability is significantly better than traditional AI translation models. • Seamless and secure interaction throughout: The physical sound wave replacement is completed only instantaneously, without generating, storing, or transmitting any text, semantic, or interaction log data, making it suitable for communication needs in high-end business and confidential scenarios. • Zero language barrier: Users do not need to have foreign language knowledge. All foreign language speech can be converted into the user's native language sound waves, enabling passive real-time understanding. • Low computational power and offline operation: It does not rely on large models, cloud corpora and GPU computing power resources. It can run stably with only conventional DSP acoustic hardware and supports a fully offline working mode. • Breaking through the traditional semantic translation paradigm, it constructs a purely physical cross-linguistic interaction path that directly replaces exogenous sound waves with native language sound waves. • Supports emotionally relatable interaction: By distinguishing between fixed language base features and dynamic emotional fluctuation features, a dual working mode is set up, which can take into account both standardized communication and immersive emotional interaction, and realize all-round interaction from semantic understanding to emotional perception. 6. The essential differences from existing technologies This invention differs fundamentally from existing mainstream cross-language voice interaction technologies (such as direct speech translation based on deep learning models and formant bending technology based on speaker identity transformation). Existing technologies either rely on deep learning models for probabilistic fitting and semantic understanding, or aim to preserve the identity features of the target speaker. This invention neither employs any deep learning or neural network models for reasoning, nor does it aim to preserve any speaker timbre or identity features. This invention focuses on the physical extraction and deterministic mathematical mapping of the inherent acoustic base features of the language, achieving sound wave substitution through deterministic mathematical methods such as lookup tables, linear interpolation, or polynomial fitting. Therefore, it differs fundamentally from existing technologies in its technical approach, technical objectives, and implementation methods. Detailed Implementation The initial phase involved collecting standard acoustic wave samples from major global languages, extracting and solidifying the inherent acoustic base features of each language, and establishing a precise mapping system from multiple languages to the mother tongue. Specifically, the construction method involved collecting a large number of pure human voice acoustic wave samples from different speakers and with different content for each language, extracting stable acoustic parameters from the samples, including the statistical distribution extrema of Mel-frequency cepstral coefficients (MFCC) and linear predictive coding coefficients (LPC) in the corresponding language, and constructing an acoustic feature space specific to each language. A deterministic mapping function between the feature spaces of different languages was then established using a statistical alignment algorithm. [Enhanced transparency] The acoustic feature mapping of this invention achieves deterministic numerical calculation based on the statistical distribution of acoustic parameters, requiring no deep learning model training parameters and not relying on large-scale parallel bilingual corpora. The feature permutation process is a mathematical transformation of physical acoustic parameters, which differs from traditional data-driven probabilistic fitting methods.
Mapping Method Example
[0004] Scenario B (Daily Social Scenarios): Activate Emotion Recreation Mode. While completing the language acoustic feature conversion, the device retains and replicates the original sound wave's pitch fluctuations, rhythm strength, and other emotional parameters, outputting voice with native tone and emotion to recreate the atmosphere of a real conversation. Attached Figure Description
[0005] Figure 1A complete physical closed-loop process diagram illustrates the complete signal flow from external sound wave acquisition and purification (S1) to language base purification and simultaneous extraction of emotional parameters (S2) to mother tongue mapping and replacement (S3) to mother tongue equivalent sound wave reshaping and simultaneous replication of emotional features (S4) to directional auditory output (S5). Each step corresponds one-to-one with the technical features in claims 1 and 11. Specifically, step S2, while removing individual interference and purifying the inherent acoustic base features of the language, simultaneously acquires emotionally related fluctuation parameters such as pitch variation, volume intensity, and breath rhythm. Step S4, while reshaping the mother tongue equivalent sound wave, simultaneously replicates the emotionally related fluctuation parameters, enabling the output sound wave to simultaneously carry language information and emotional features. Figure 2 Schematic diagram of language-specific acoustic basis feature purification: It shows the processing logic based on MFCC / LPC feature extraction, specifically demonstrating the process of removing interference and extracting stable language-specific acoustic basis feature vectors from the original sound waves containing individual interference such as emotion and speech rate through normalization and regularization. Figure 3 Schematic diagram of closed-loop mapping of multilingual sound waves to equivalent sound waves in the mother tongue: The dashed box in the figure is a pre-constructed multilingual acoustic feature mapping library, which shows the deterministic transformation relationship of acoustic feature spaces of different languages, as well as the waveform comparison effect before and after the physical reshaping of sound waves.
Claims
1. A method for text-free cross-lingual real-time interaction based on human voice inherent acoustic wave features, characterized in that, Includes the following steps: S1. Through sound wave acquisition equipment, natural human voice sound wave signals of any foreign language in the environment are acquired in real time, filtered and purified to remove environmental noise and invalid interference, and retain pure original sound wave signals of foreign languages. S2. Normalize and regularize the pure original sound wave signal of the foreign language, remove the temporary fluctuation interference caused by individual vocalization, and extract the stable inherent acoustic base features of the foreign language that are independent of semantics. S3. Based on a pre-built multilingual acoustic feature mapping library, the stable language-specific acoustic base features of the foreign language are replaced with standard acoustic feature parameters corresponding to the listener's preset mother tongue. S4. Based on the standard acoustic feature parameters, physically reconstruct and generate a native language equivalent sound wave waveform that the listener can directly recognize; S5. The equivalent sound wave of the mother tongue is amplified by the directional sound wave output module, and the listener can directly identify the information carried by the sound wave by relying on their native hearing. The feature permutation process is a deterministic mapping based on physical acoustic parameters, which does not involve reasoning based on statistical probability deep learning models, nor does it generate any intermediate semantic representations.
2. The method according to claim 1, characterized in that: The temporary fluctuation interference includes one or more of the following: gender differences, age differences, speech rate changes, emotional fluctuations, and accent deviations.
3. The method according to claim 1, characterized in that: The construction process of the multilingual acoustic feature mapping library includes: collecting standard pronunciation samples of multiple languages, statistically analyzing the distribution of inherent acoustic basis features of each language, and establishing feature space mapping relationships between languages; this process does not involve semantic annotation and does not rely on deep learning training of parallel bilingual corpora.
4. The method according to claim 1, characterized in that: The native language equivalent sound wave waveform reconstruction adopts physical waveform synthesis technology and outputs a continuous sound wave signal that can be recognized by the natural human ear; the reconstruction is based on matching the acoustic statistical distribution of the listener's native language, and does not aim to preserve the timbre or identity characteristics of the source speaker.
5. The method according to claim 1, characterized in that: By supplementing multilingual acoustic samples, the range of language compatibility can be continuously expanded, and the accuracy of equivalent sound wave reproduction in the native language can be optimized.
6. A text-free cross-lingual real-time interaction system based on human voice inherent acoustic wave features, characterized in that, include: The sound wave acquisition and purification module is used to acquire human voice sound wave signals of any foreign language in the environment in real time and perform filtering and purification. The language-specific acoustic base feature purification module is used to remove individual interference fluctuations and extract stable language-specific acoustic base features. The native language mapping and substitution module is used to convert the acoustic features of foreign languages into standard acoustic feature parameters of the listener's preset native language. The native language equivalent sound wave reconstruction module is used to physically generate native language equivalent sound wave waveforms that can be directly recognized by the listener; A directional sound wave output module is used to amplify the equivalent sound wave of the mother tongue. The system consists of a pure hardware processing unit, without text processing unit, semantic parsing unit or artificial intelligence computing unit, and the output end is only connected to the sound wave playback device, without outputting any language classification tags or text data.
7. The system according to claim 6, characterized in that: The system does not generate, store, or transmit any textual or semantic interactive data during operation.
8. The system according to claim 6, characterized in that: The system supports offline operation and is compatible with embedded terminal devices.
9. The system of claim 6, wherein: The system is encapsulated in a portable intercom device, a vehicle-mounted voice terminal, a wearable hearing aid interactive device, or a general computing peripheral device. The general computing peripheral device includes, but is not limited to, a USB interface translation device, a smartphone headphone jack adapter, a laptop audio expansion device, or other audio processing devices that are connected to the computing device through a general interface. And the system includes at least the following at the physical level: A microphone is used to enable the audio input function of the sound wave acquisition and purification module; The processor is used to run the language-specific acoustic basis feature purification module, the mother tongue mapping and replacement module, and the mother tongue equivalent sound wave reconstructing module; The memory is used to store the multilingual acoustic feature mapping library; A loudspeaker is used to realize the audio output function of the directional sound wave output module; The processor is a digital signal processor (DSP) or a microcontroller (MCU) with a hardware floating-point unit. The processor is connected to the memory via an I2S bus. An anti-aliasing filter circuit is connected between the microphone and the processor. A power amplifier and a low-pass filter circuit are connected between the processor and the speaker.
10. The system according to claim 9, characterized in that: The system does not include a wireless communication module, or interacts with external general-purpose computing devices directly through a physical interface.
11. The method according to claim 1, characterized in that: While extracting the inherent acoustic base features of the language, emotion-related fluctuation parameters in the sound wave signal are simultaneously collected. These emotion-related fluctuation parameters include one or more of pitch fluctuation, volume intensity, and breath rhythm. When physically reshaping the sound wave output, the emotion-related fluctuation parameters are simultaneously replicated, so that the output sound wave carries both language information and emotional characteristics.
12. The method according to claim 11, characterized in that: The system can switch between two working modes: the first mode is the standard communication mode, which filters out emotional fluctuations and outputs regular information speech; the second mode is the emotion restoration mode, which retains emotional parameters and outputs immersive speech with added tone and emotion.