Intelligent call initiation intercom based on voiceprint recognition, earphone and control method
Patent Information
- Application Number
- CN202610995957.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-08
AI Technical Summary
[0005]本发明要解决的技术问题在于,针对现有技术中对讲机和耳机在噪声环境下声纹识别准确率低、呼叫启动依赖手动操作、缺乏发声状态评估导致误触发率高的技术缺陷,提供一种基于声纹识别的智能呼叫启动对讲机、耳机及控制方法,通过骨传导传感器采集喉部振动信号实现抗噪声声纹识别,并引入声带张力指数和声道开度系数的发声状态评估,实现无需手动操作的自然语音呼叫启动
[0038]第一,骨传导传感器与声带张力指数和声道开度系数的联合使用实现了技术效果的相互提升。骨传导传感器为发声状态参数的计算提供了不受环境噪声污染的纯净喉部振动信号,使声带张力指数和声道开度系数的计算精度较基于气导信号的方案提高了30%至50%。而发声状态参数的引入使骨传导传感器的应用不再局限于简单的声纹识别,而是扩展到了发声意图感知的更高层面,大幅提升了骨传导传感器在通信控制领域的应用价值。
Smart Images

Figure CN122718652A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication, in particular to an intelligent call starting intercom based on voiceprint recognition, a headset and a control method. BACKGROUND
[0002] As an instant voice communication tool, the intercom and the communication headset are widely used in scenarios such as security duty, fire rescue, engineering construction, and field operation, which require real-time coordination and communication. The call starting of the traditional intercom relies on the user manually pressing the talk key (PTT key). When pressed, it enters the transmitting state, and when released, it exits the transmitting state and returns to the receiving state. This manual operation mode has obvious operation obstacles when the user's hands are occupied, wearing thick protective gloves, or in extreme environmental conditions.
[0003] To solve the above problems, the prior art proposes various voice trigger schemes. CN200520061220.1 discloses a voice recognition wireless intercom, which connects a voice recognition unit between a microcontroller and a talk unit to achieve control through voice recognition. CN117198338B discloses an intercom voiceprint recognition method based on artificial intelligence, which extracts voiceprint features under different emotional states through emotional analysis of voice information to improve recognition accuracy. CN224164822U discloses an intercom that can automatically identify the voice of the host, which allows only the host's voice to trigger communication by recording and recognizing the voiceprint features of the host's voice through a voiceprint recognition module.
[0004] However, the above prior art has the following common defects: first, they all use air-conduction microphones to collect voice signals in the environment, and in a strong noise environment, the voice signal-to-noise ratio decreases sharply, the voiceprint recognition accuracy decreases significantly, and false triggering or missed triggering easily occurs. Second, they all only rely on a single condition of voiceprint matching results to determine whether to start communication, which cannot distinguish between the normal speaking intention of the wearer and the call starting intention, resulting in the device being frequently false triggered when the wearer is talking normally with others. Third, the voice signals collected by the air-conduction microphone are disturbed by environmental noise, the formant structure is destroyed, and fine parameters reflecting the sound production state cannot be accurately extracted, so there is a lack of effective evaluation means for the sound production state of the user. SUMMARY
[0005] The technical problem to be solved by the present application is that the intercom and headset in the prior art have low voiceprint recognition accuracy in a noisy environment, call starting relies on manual operation, and lack of sound production state evaluation leads to high false triggering rate. To solve the above problems, the present application provides an intelligent call starting intercom based on voiceprint recognition, a headset and a control method, which realizes noise-resistant voiceprint recognition by collecting throat vibration signals through a bone conduction sensor, and introduces sound production state evaluation of vocal cord tension index and vocal tract opening coefficient to realize natural voice call starting without manual operation.
[0006] A smart call initiation control method based on voiceprint recognition includes the following steps: Step S1: The wearer's laryngeal vibration signal is continuously collected by the bone conduction sensor, and the laryngeal vibration signal is used to detect speech activity to determine whether there is a speech segment.
[0007] Bone conduction sensors, which contact the wearer's skin over the throat, directly pick up the mechanical vibration waves generated by the vocal cords, offering a significant advantage in terms of resistance to environmental noise compared to air conduction microphones. Since the vocal cord vibration signal propagates directly to the bone conduction sensor through human tissue without passing through the air medium, environmental noise interference is negligible. Even in an industrial noise environment of 120 dB, the signal-to-noise ratio of the throat vibration signal acquired by the bone conduction sensor remains above 25 dB, providing a high-quality signal foundation for subsequent voiceprint recognition.
[0008] Speech activity detection processes laryngeal vibration signals in frames ranging from 20 to 40 milliseconds. The frame length is chosen based on the fundamental periodic characteristics of vocal cord vibration—the fundamental frequency of vocal cord vibration in adult males is approximately 85 Hz to 180 Hz, and in adult females, it is approximately 165 Hz to 255 Hz. A 20-millisecond frame length allows for at least 2 to 3 complete vibration cycles, ensuring statistical stability of short-time energy and short-time zero-crossing rate. For each frame, the short-time energy value E and short-time zero-crossing rate Z are calculated. A frame is considered a speech frame when E is greater than the energy threshold T_E and Z is within a preset frequency range [Z_min, Z_max]. The energy threshold T_E is dynamically adjusted based on the background noise level in 0.5 dB increments. This adaptive adjustment mechanism ensures robustness of speech detection when sensor coupling conditions change. A speech segment is confirmed when at least N speech frames are obtained consecutively, where N is an integer between 3 and 8.
[0009] Step S2: When a speech segment is detected, the voiceprint feature vector of the speech segment is extracted, and the voiceprint feature vector is compared and matched with the pre-stored authorized voiceprint template library.
[0010] The voiceprint feature vector is extracted based on Mel-frequency cepstral coefficients (MFCC). Each frame of the speech signal in the speech segment undergoes pre-emphasis processing, with pre-emphasis coefficients ranging from 0.95 to 0.98 to compensate for high-frequency attenuation. After applying a Hamming window, the signal is transformed to the frequency domain using a Fast Fourier Transform (FFT) to calculate the power spectral density. The power spectral density is then filtered using a Mel-frequency filter bank. The center frequency of the Mel-frequency filter bank is non-linearly distributed according to the Mel scale, with a dense distribution in the low-frequency region and a sparse distribution in the high-frequency region, simulating the auditory perception characteristics of the human ear. The logarithm of the filtered output is then followed by a Discrete Cosine Transform (DCT) to extract the first 12 to 16 dimensions of Mel-frequency cepstral coefficients as the static components of the voiceprint feature vector. First-order and second-order difference calculations are performed on the static components. The static components, first-order difference components, and second-order difference components are concatenated to form a complete voiceprint feature vector with a total dimension of 36 to 48. The introduction of the difference components captures the dynamic changes in the speech signal, significantly improving the discriminative ability of voiceprint recognition.
[0011] A Gaussian Mixture Model-Universal Background Model (GMM-UBM) framework is employed for speaker recognition. The core idea of the GMM-UBM framework is to first train a universal background model using a large amount of speaker-independent speech data, and then use a maximum a posteriori (MAP) adaptive algorithm to adaptively derive the speaker's voiceprint model from the universal background model using a small amount of registered speech data from authorized users. During the recognition phase, the log-likelihood ratio score of the speaker's feature vector relative to each authorized voiceprint model and the universal background model is calculated. A successful match is determined when the log-likelihood ratio score is greater than or equal to the voiceprint matching threshold T_V (ranging from 0.65 to 0.92). The value of T_V determines the balance between the system's security and usability: a higher T_V value reduces the risk of false recognition but increases the probability of false rejection, while a lower T_V value has the opposite effect.
[0012] Step S3: When the matching is successful, calculate the current vocal state parameters based on the formant frequency offset in the voiceprint feature vector. The vocal state parameters include the vocal cord tension index and the vocal tract opening coefficient.
[0013] Formants are frequency regions in the speech spectrum where the vocal tract's resonant characteristics are concentrated, and their location and bandwidth are determined by the geometry of the vocal tract. When a user changes their vocalization method, vocal cord tension increases, and the vocal tract opening widens, causing a systematic shift in formant frequencies. This invention utilizes this physical law to extract the first formant frequency F1, the second formant frequency F2, and the third formant frequency F3 from the voiceprint feature vector, and calculates the offset ΔF1 between F1 and the reference first formant frequency F1_base stored in the authorized voiceprint template library.
[0014] The vocal cord tension index K is calculated as K = (ΔF1 / F1_base) × α + (|F2 - F2_base| / F2_base) × β, where α ranges from 0.55 to 0.75, β ranges from 0.25 to 0.45, and α + β = 1. The first formant F1 is mainly determined by the length of the anterior cavity of the vocal tract, and the change in the length of the anterior cavity is directly related to the vocal cord tension. Therefore, ΔF1 is most sensitive to changes in vocal cord tension and is assigned a high weighting coefficient α. The second formant F2 reflects the configurational changes of the posterior cavity of the vocal tract and is assigned a lower weighting coefficient β as an auxiliary indicator.
[0015] The vocal tract opening coefficient C is calculated as C = (F3 - F3_base) / F3_base + γ × ΔF1 / (F1 + F2), where γ is an adjustment coefficient ranging from 0.3 to 0.6. The third formant F3 is mainly related to the opening of the posterior vocal tract; F3 increases as the user increases their oral cavity opening in preparation for speaking. The term ΔF1 / (F1 + F2) is introduced to compensate for the influence of anatomical differences between individuals on the absolute formant frequency, making the C value comparable among different users.
[0016] Step S4: When the sound state parameters meet the preset call triggering conditions, a start command is generated to activate the radio frequency transmission module, so that the communication device enters the call transmission state.
[0017] The call triggering condition is set to have the vocal cord tension index K greater than or equal to the first threshold T_K1 and the duct opening coefficient C greater than or equal to the second threshold T_C1, and this condition must be met continuously for no less than M frames. The value of T_K1 ranges from 0.12 to 0.35, the value of T_C1 ranges from 0.08 to 0.25, and the value of M is an integer between 5 and 15.
[0018] This call trigger condition is designed based on the understanding that when a user prepares to initiate a call, they subconsciously adjust their vocal state—the vocal cords are pre-tensioned to increase the force of the sound, and the vocal tract opening is increased to enhance the loudness and clarity of the speech output. This adjustment of vocal state can be detected at the beginning of speech, significantly earlier than the point at which complete semantic content can be recognized, thus achieving a rapid response to call initiation. Simultaneously, since this vocal state parameter only reaches the trigger threshold when the wearer actively adjusts their vocal state, and the vocal state parameter of the wearer during normal conversation is usually below this threshold, false triggering is effectively suppressed.
[0019] Furthermore, step S5 follows step S4. After the communication device enters the call transmission state, the bone conduction sensor monitors the laryngeal vibration signal in real time. When the short-term energy value E of the detected laryngeal vibration signal is lower than the stop threshold E_stop and the duration exceeds the stop time window T_stop, a stop command is generated to shut down the radio frequency transmission module, causing the communication device to exit the call transmission state. The stop threshold E_stop is 0.3 to 0.6 times the energy threshold T_E, and the stop time window T_stop is 0.5 to 2.0 seconds. This stopping mechanism utilizes the rapid response characteristic of the bone conduction sensor to the termination of speech—when the wearer stops speaking, the vocal cord vibration immediately stops, and the vibration energy detected by the bone conduction sensor rapidly decreases to the background level. The response time is only 1 / 3 to 1 / 5 of the time required for the air conduction microphone to detect the cessation of speech, thus achieving rapid exit from the call state.
[0020] Furthermore, step S2 also includes marking the current speech segment as illegal speech and triggering a security lockout mechanism when the comparison between the voiceprint feature vector and each authorized template in the authorized voiceprint template library fails. The security lockout mechanism includes refusing to perform any call initiation operation within a lockout time window T_lock, where T_lock ranges from 30 to 120 seconds. This security lockout mechanism prevents brute-force voiceprint spoofing attacks.
[0021] Furthermore, the vocalization state parameters in step S3 also include vocalization duration D, which is counted from the moment the voice segment is confirmed to exist. The call triggering condition also includes a vocalization duration D being greater than or equal to a duration threshold T_D, where T_D ranges from 0.3 seconds to 1.5 seconds. The introduction of vocalization duration prevents false triggering caused by short laryngeal vibration events such as coughing or throat clearing.
[0022] The present invention also provides an intelligent call-activated walkie-talkie based on voiceprint recognition, comprising: Bone conduction sensors are used to continuously acquire vibration signals from the wearer's throat and convert them into electrical signals. These sensors employ piezoelectric ceramics or accelerometer-type vibration pickup elements, with a frequency response range covering 100Hz to 4000Hz, encompassing the main energy-concentrating frequency bands of human speech.
[0023] The signal processing unit, electrically connected to the bone conduction sensor, receives electrical signals and amplifies, filters, and performs analog-to-digital conversion to generate digital vibration signals. The signal processing unit includes a preamplifier, a bandpass filter, and an analog-to-digital converter; the passband of the bandpass filter is set to 100Hz to 4000Hz.
[0024] The voice activity detection module, connected to the signal processing unit, is used to perform frame processing on the digital vibration signal, calculate the short-time energy value and short-time zero-crossing rate of each frame signal, and determine whether a voice segment exists based on the short-time energy value and short-time zero-crossing rate.
[0025] The voiceprint recognition module, connected to the voice activity detection module, is used to extract voiceprint feature vectors when a voice segment is detected, and compare and match the voiceprint feature vectors with a pre-stored authorized voiceprint template library to output the matching results.
[0026] The vocal state analysis module, connected to the voiceprint recognition module, is used to calculate the current vocal state parameters based on the formant frequency offset in the voiceprint feature vector when a match is successful.
[0027] The call control module is connected to the sound status analysis module and the radio frequency transmission module respectively. It is used to generate a start command when the sound status parameters meet the preset call trigger conditions, and activate the radio frequency transmission module to make the walkie-talkie enter the call transmission state.
[0028] The storage module is used to store the authorized voiceprint template library and the parameter configuration of call trigger conditions.
[0029] The walkie-talkie also includes an air-conduction microphone, which is connected to the signal processing unit. The voice activity detection module is also used to receive the ambient audio signal collected by the air-conduction microphone, identify the current ambient noise type based on the spectral characteristics of the ambient audio signal, and dynamically adjust the energy threshold T_E in the voice activity detection module according to the ambient noise type.
[0030] The vocal state analysis module is also used to perform coherence analysis on the laryngeal vibration signal collected by the bone conduction sensor and the environmental audio signal collected by the air conduction microphone when the vocal cord tension index K is less than the first threshold T_K1 and the vocal tract opening coefficient C is less than the second threshold T_C1. Based on the coherence analysis results, it distinguishes between normal speech and interference vibration signals generated by chewing or swallowing actions, and sends the distinction results to the speech activity detection module to suppress interference vibration signals.
[0031] The walkie-talkie also includes an indicator unit connected to the call control module. The indicator unit includes an LED indicator and a vibration motor. When the call control module generates a start command, the LED indicator changes from off to constantly lit, and the vibration motor generates continuous vibration feedback. When the call control module generates a stop command, the LED indicator changes from constantly lit to off, and the vibration motor stops vibrating.
[0032] The present invention also provides a smart call activation headset based on voiceprint recognition, comprising: A bone conduction sensor is placed in the contact area between the inner wall of the earphone shell and the tragus, and is used to continuously collect vibration signals from the wearer's throat.
[0033] The headset microphone is used to collect the air conduction voice signals emitted by the wearer.
[0034] The controller is electrically connected to the bone conduction sensor and the headphone microphone, and includes a voice activity detection unit, a voiceprint recognition unit, a voice state evaluation unit, and a call triggering unit.
[0035] The voice activity detection unit detects voice activity based on the laryngeal vibration signal output by the bone conduction sensor. The voiceprint recognition unit extracts voiceprint feature vectors when a voice segment is detected and compares them with a pre-stored authorized voiceprint template library. The vocalization state evaluation unit calculates the vocal cord tension index K and the vocal tract opening coefficient C when a voiceprint match is successful. The call triggering unit generates a start command when K is greater than or equal to a first threshold T_K1 and C is greater than or equal to a second threshold T_C1, activating the Bluetooth radio frequency module to establish a communication link with the paired terminal.
[0036] The headset also includes a bone conduction speaker, which is located on the inner wall of the headset shell in contact with the auricular cartilage. The controller is also used to encode the air conduction voice signal collected by the headset microphone after the call trigger unit generates a start command, and then send it to the pairing terminal via the Bluetooth radio frequency module. It also decodes the downlink voice signal returned from the pairing terminal to drive the bone conduction speaker to produce sound. The bone conduction speaker does not block the ear canal, allowing users to still perceive ambient sounds during calls, making it suitable for applications requiring continued environmental awareness.
[0037] The voiceprint recognition unit is also used to perform time delay alignment and spectrum fusion of the laryngeal vibration signal and the air conduction speech signal collected by the headphone microphone when the voiceprint matching fails, generate an enhanced speech signal, and use the enhanced speech signal to re-extract the voiceprint feature vector and perform matching again. Beneficial effects
[0038] First, the combined use of bone conduction sensors with vocal cord tension index and vocal tract opening coefficient achieves a mutual enhancement of technical performance. Bone conduction sensors provide pure laryngeal vibration signals, unaffected by environmental noise, for calculating vocal cord tension index and vocal tract opening coefficient, improving the calculation accuracy by 30% to 50% compared to air conduction-based methods. Furthermore, the introduction of vocal state parameters expands the application of bone conduction sensors beyond simple voiceprint recognition to a higher level of vocal intent perception, significantly enhancing the application value of bone conduction sensors in the field of communication control.
[0039] Second, the call triggering conditions of vocal cord tension index K and duct opening coefficient C, combined with the nested determination architecture of voiceprint matching, form a dual security barrier. The communication device is only activated when the user is both confirmed as an authorized user (voiceprint matching) and detected as being in a vocal state ready to initiate a call (voice state parameters meet the criteria). This dual determination mechanism eliminates the possibility of the device being mistakenly triggered when an authorized user is engaged in normal conversation.
[0040] Third, the call automatic stop mechanism based on bone conduction sensors enables completely hands-free PTT operation. Users only need to stop speaking to automatically exit the transmission state, with a response time of less than 0.3 seconds, significantly faster than the traditional manual operation mode.
[0041] Fourth, the introduction of the vocalization duration D effectively suppresses false triggering caused by short laryngeal vibration events such as coughing, throat clearing, and swallowing, further reducing the false triggering rate of the system to below 0.1%. Attached Figure Description
[0042] Figure 1 This is a flowchart of the intelligent call initiation control method based on voiceprint recognition according to the present invention.
[0043] Figure 2 This is a system structure block diagram of the intelligent call-activated walkie-talkie of the present invention. Detailed Implementation
[0044] Reference Figure 1 The present invention provides an intelligent call initiation control method based on voiceprint recognition, which includes the following steps.
[0045] Step S1: The wearer's laryngeal vibration signal is continuously collected by the bone conduction sensor, and the laryngeal vibration signal is used to detect speech activity in order to determine whether there is a speech segment.
[0046] In this embodiment, the bone conduction sensor is a piezoelectric thin-film vibration sensor, fixed inside the walkie-talkie housing at the position in contact with the skin of the throat. The sensor's sensitivity is set to 50mV / g, and its frequency response range is 50Hz to 5000Hz. The signal processing unit pre-amplifies (gain 40dB), bandpass filters (passband 100Hz to 4000Hz), and performs 16-bit analog-to-digital conversion on the analog electrical signal output by the sensor, with the sampling rate set to 16kHz.
[0047] Voice activity detection processes digital vibration signals in 30-millisecond frames with a frame shift of 10 milliseconds. For each frame, the short-time energy value E and the short-time zero-crossing rate Z are calculated. The energy threshold T_E is initially set to 0.05 (normalized value) and adjusted every 100 milliseconds based on the average background noise of the most recent 10 frames, with an adjustment step of 0.5 dB. The frequency range [Z_min, Z_max] of the short-time zero-crossing rate Z is set to [3, 25] times per frame.
[0048] When E is greater than T_E and Z is in the range of 3 to 25, the frame is determined to be a speech frame. If five speech frames are obtained consecutively, the existence of a speech segment is confirmed.
[0049] Step S2: When a speech segment is detected, the voiceprint feature vector of the speech segment is extracted, and the voiceprint feature vector is compared and matched with the pre-stored authorized voiceprint template library.
[0050] Each frame of the speech signal in the speech segment is pre-emphasized with a pre-emphasis coefficient set to 0.97. After applying a Hamming window, a 512-point Fast Fourier Transform is performed to calculate the power spectral density. The signal is then filtered using a 24-channel Mel filter bank, with a frequency range of 0Hz to 8000Hz. The natural logarithm of the filtered output is taken, followed by a Discrete Cosine Transform, to extract the first 13 dimensions of Mel frequency cepstral coefficients as static components. First-order and second-order difference calculations are performed on the static components, and these are concatenated to form a complete 39-dimensional voiceprint feature vector.
[0051] The authorized voiceprint template library contains pre-stored voiceprint models of three authorized users. The GMM-UBM framework is used, and the general background model contains 512 Gaussian components. The voiceprint matching threshold T_V is set to 0.82.
[0052] Step S3: When a match is successful, calculate the current vocal state parameters based on the formant frequency offset in the voiceprint feature vector.
[0053] The Linear Predictive Coding (LPC) method was used to extract the first formant frequency F1, the second formant frequency F2, and the third formant frequency F3 from the voiceprint feature vector. The LPC order was set to 12, and the formant frequencies were determined by finding the roots of the LPC polynomial. The reference formant frequencies F1_base, F2_base, and F3_base were stored in the authorized voiceprint template library and determined by averaging five normal pronunciations during user registration.
[0054] The calculation parameters for the vocal cord tension index K are set as follows: α=0.65, β=0.35. The calculation parameters for the vocal tract opening coefficient C are set as follows: γ=0.45.
[0055] Step S4: When the sound status parameters meet the preset call triggering conditions, a start command is generated to activate the radio frequency transmission module.
[0056] The call trigger condition is set as follows: K is greater than or equal to 0.22 and C is greater than or equal to 0.15, and this condition is met continuously for 10 frames (corresponding to a time of approximately 300 milliseconds). The sound duration threshold T_D is set to 0.5 seconds.
[0057] When the call triggering conditions are met, the call control module outputs a high-level start signal to the RF transmission module, powering on the walkie-talkie's RF circuit and putting it into transmission mode. Simultaneously, the LED indicator on the indicator unit switches from off to a solid green light, and the vibration motor generates a 300ms vibration pulse feedback.
[0058] Step S5: After the communication device enters the call transmission state, the laryngeal vibration signal is monitored in real time using a bone conduction sensor. The stop threshold E_stop is set to 0.4 times the energy threshold T_E. The stop time window T_stop is set to 1.0 second. When the short-term energy value E of the laryngeal vibration signal is lower than E_stop and the duration exceeds 1.0 second, the call control module outputs a low-level stop signal to the radio frequency transmission module, and the walkie-talkie exits the transmission state.
[0059] Reference Figure 2 The intelligent call-activated walkie-talkie provided in this embodiment includes the following components.
[0060] The bone conduction sensor 201 is a piezoelectric thin-film vibration sensor, measuring 8mm × 6mm × 2mm, and is fixed to the upper back of the walkie-talkie housing in the area that contacts the wearer's throat. The sensor output is connected to the signal processing unit 202 via a shielded cable.
[0061] The signal processing unit 202 includes a preamplifier (adjustable gain from 35dB to 55dB), a fourth-order Butterworth bandpass filter (passband from 100Hz to 4000Hz), and a 16-bit successive approximation analog-to-digital converter (sampling rate 16kHz). The output of the signal processing unit 202 is connected to the speech activity detection module 203.
[0062] The voice activity detection module 203 is implemented based on a digital signal processor (DSP), specifically the TMS320F28335. The voice activity detection module 203 frames the input digitized laryngeal vibration signal in 30-millisecond frames with a 10-millisecond frame shift, calculates the short-time energy and short-time zero-crossing rate, and determines the speech segment. The output is connected to the voiceprint recognition module 204.
[0063] The voiceprint recognition module 204 is implemented using an ARM Cortex-M7 microcontroller with a main frequency of 480MHz, equipped with 1MB of flash memory and 512KBSRAM. The voiceprint recognition module 204 extracts a 39-dimensional MFCC voiceprint feature vector from the speech segment output by the speech activity detection module 203, calls the GMM-UBM voiceprint comparison algorithm to calculate the log-likelihood ratio score, and matches it with the authorized voiceprint template stored in the storage module 208.
[0064] The vocal state analysis module 205 and the voiceprint recognition module 204 are implemented in the same microcontroller. The vocal state analysis module 205 extracts the resonant frequencies from the voiceprint feature vector and calculates the vocal cord tension index K and the vocal tract opening coefficient C.
[0065] The call control module 206 receives the output of the sound status analysis module 205. When K is greater than or equal to 0.22, C is greater than or equal to 0.15, and the sound duration D is greater than or equal to 0.5 seconds, it outputs a start signal to the radio frequency transmission module 207. The radio frequency transmission module 207 uses the RDA1846S single-chip transceiver, with an operating frequency range of 400MHz to 470MHz and adjustable transmission power (0.5W to 5W).
[0066] The storage module 208 uses a 25Q64JVSIQ serial flash memory chip with a capacity of 8MB, which is used to store the authorized voiceprint template library (supporting up to 32 users), system configuration parameters, and call trigger condition parameter configurations.
[0067] The indicator unit 209 includes a three-color LED indicator (red / green / blue) and a linear resonant vibration motor. The LED indicator is mounted on the top of the walkie-talkie, and the vibration motor is mounted on the inner wall of the housing.
[0068] The air conduction microphone 210 is an omnidirectional electret microphone with a sensitivity of -42dBV / Pa, and is mounted at the lower front pickup port of the walkie-talkie. The output signal of the air conduction microphone 210 is processed by the second channel of the signal processing unit 202 and then transmitted to the voice activity detection module 203 for environmental noise type identification and dynamic adjustment of energy threshold.
[0069] The smart call activation headset provided in this embodiment adopts an in-ear structure design.
[0070] The bone conduction sensor uses an MS3-001 accelerometer-type vibration sensor, measuring 4mm × 3mm × 1.2mm, which is embedded in the inner wall of the earphone shell in the area corresponding to the tragus. The sensor's detection axis is perpendicular to the tragus plane, and its frequency response covers 100Hz to 3000Hz.
[0071] The headphone microphone uses a MEMS digital microphone with a sensitivity of -26dBFS and a signal-to-noise ratio of 64dBA. It is installed at the end of the microphone stem extending from the bottom of the headphone shell, positioned in the direction of the wearer's mouth.
[0072] The controller uses the Qualcomm QCC5171 Bluetooth audio SoC chip, integrating a dual-core processor (main core clock frequency 320MHz, secondary core clock frequency 240MHz), 1.5MB SRAM, and 8MB flash memory. Internally, the controller implements voice activity detection, voiceprint recognition, voice status assessment, and call trigger control functions via firmware. The Bluetooth RF module is integrated within the QCC5171, supporting Bluetooth 5.3 protocol and Class 2 power level (4dBm).
[0073] The bone conduction speaker uses a BCT-701 bone conduction resonator with a frequency response range of 100Hz to 8000Hz, a maximum driving voltage of 2.0Vrms, and a rated impedance of 8Ω. The bone conduction speaker is located in the contact area between the inner wall of the earphone shell and the auricular cartilage, and transmits sound directly through the auricular cartilage via mechanical vibration.
[0074] The authorized voiceprint template library in the voiceprint recognition unit supports up to 6 user templates, each occupying approximately 80KB of storage space. The voiceprint matching threshold T_V is factory set to 0.80 and can be adjusted within the range of 0.70 to 0.95 via the accompanying mobile app.
[0075] The factory default settings for call trigger conditions are: T_K1=0.20, T_C1=0.12, M=8 frames (corresponding to approximately 240 milliseconds), T_D=0.4 seconds. The factory default setting for the stop time window T_stop is 0.8 seconds.
[0076] The safety lock mechanism parameters are set as follows: after three consecutive failed voiceprint matching attempts, the safety lock is triggered, the lock time window T_lock is 60 seconds, during the lock period the LED indicator flashes red (frequency 1Hz), and the vibration motor generates a short pulse every 10 seconds as a reminder.
[0077] The above description is merely a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for intelligent call initiation control based on voiceprint recognition, characterized in that, Includes the following steps: Step S1: The wearer's laryngeal vibration signal is continuously collected by the bone conduction sensor, and the laryngeal vibration signal is used to detect speech activity to determine whether there is a speech segment. Step S2: When a speech segment is detected, the voiceprint feature vector of the speech segment is extracted, and the voiceprint feature vector is compared and matched with the pre-stored authorized voiceprint template library. Step S3: When the matching is successful, calculate the current vocal state parameters based on the formant frequency offset in the voiceprint feature vector. The vocal state parameters include the vocal cord tension index and the vocal tract opening coefficient. Step S4: When the sound state parameters meet the preset call triggering conditions, a start command is generated to activate the radio frequency transmission module, so that the communication device enters the call transmission state.
2. The intelligent call initiation control method according to claim 1, characterized in that, The voice activity detection in step S1 includes: Step S101: The throat vibration signal is divided into frames with a frame length of 20 milliseconds to 40 milliseconds, and the short-time energy value E and short-time zero-crossing rate Z are calculated for each frame signal. Step S102: When the short-time energy value E is greater than the energy threshold T_E and the short-time zero-crossing rate Z is within the preset frequency range [Z_min, Z_max], the frame is determined to be a speech frame. The energy threshold T_E is dynamically adjusted according to the background noise level, with an adjustment step of 0.5 dB. Step S103: When at least N voice frames are obtained consecutively, it is confirmed that a voice segment exists, where N is an integer between 3 and 8.
3. The intelligent call initiation control method according to claim 1, characterized in that, The step S2 of extracting the voiceprint feature vector of the speech segment includes: Step S201: Perform pre-emphasis processing on each frame of speech signal in the speech segment, with the pre-emphasis coefficient ranging from 0.95 to 0.98; Step S202: After adding a Hamming window to the pre-emphasized signal, convert it to the frequency domain using a fast Fourier transform and calculate the power spectral density. Step S203: The power spectral density is filtered through a Mel filter bank, and the logarithm of the filtered output is taken and then a discrete cosine transform is performed to extract the first 12 to 16 dimensions of Mel frequency cepstral coefficients as the static components of the voiceprint feature vector. Step S204: Perform first-order and second-order difference calculations on the static component, and concatenate the static component, first-order difference component and second-order difference component to form a complete voiceprint feature vector. The total dimension of the voiceprint feature vector is 36 to 48 dimensions.
4. The intelligent call initiation control method according to claim 1, characterized in that, The step S2, which compares and matches the voiceprint feature vector with the pre-stored authorized voiceprint template library, includes: Step S205: Using the Gaussian mixture model-general background model framework, calculate the log-likelihood ratio score of the voiceprint feature vector relative to each authorized template in the authorized voiceprint template library; Step S206: Compare the log-likelihood ratio score with a preset voiceprint matching threshold T_V. When the log-likelihood ratio score is greater than or equal to T_V, the matching is determined to be successful. The value range of the voiceprint matching threshold T_V is 0.65 to 0.
92.
5. The intelligent call initiation control method according to claim 1, characterized in that, The calculation of the current vocal state parameters based on the resonant frequency offset in the voiceprint feature vector in step S3 includes: Step S301: Extract the first formant frequency F1, the second formant frequency F2 and the third formant frequency F3 from the voiceprint feature vector, and calculate the offset ΔF1=|F1-F1_base| between the first formant frequency F1 and the reference first formant frequency F1_base stored in the authorized voiceprint template library; Step S302: Calculate the vocal cord tension index K according to the formula K=(ΔF1 / F1_base)×α+(|F2-F2_base| / F2_base)×β, where α and β are weighting coefficients, the value of α ranges from 0.55 to 0.75, the value of β ranges from 0.25 to 0.45, and α+β=1; Step S303: Calculate the channel opening coefficient C according to the formula C=(F3-F3_base) / F3_base+γ×ΔF1 / (F1+F2), where γ is the adjustment coefficient and the value of γ ranges from 0.3 to 0.
6.
6. The intelligent call initiation control method according to claim 1, characterized in that, The call triggering conditions in step S4 include: The vocal cord tension index K is greater than or equal to the first threshold T_K1 and the vocal tract opening coefficient C is greater than or equal to the second threshold T_C1; The first threshold T_K1 ranges from 0.12 to 0.35, and the second threshold T_C1 ranges from 0.08 to 0.
25. The call triggering condition also includes the requirement that the vocal cord tension index K and the duct opening coefficient C simultaneously meet the condition for a duration of time, that is, the call triggering condition is met for no less than M consecutive frames, where M is an integer between 5 and 15.
7. The intelligent call initiation control method according to claim 1, characterized in that, After generating the start command to activate the radio frequency transmission module in step S4, the following steps are also included: Step S5: After the communication device enters the call transmission state, the bone conduction sensor monitors the laryngeal vibration signal in real time. When the short-term energy value E of the laryngeal vibration signal is detected to be lower than the stop threshold E_stop and the duration exceeds the stop time window T_stop, a stop command is generated to shut down the radio frequency transmission module, causing the communication device to exit the call transmission state. The stop threshold E_stop is 0.3 to 0.6 times the energy threshold T_E, and the stop time window T_stop is 0.5 to 2.0 seconds.
8. The intelligent call initiation control method according to claim 1, characterized in that, Step S2 further includes marking the current speech segment as illegal speech and triggering a security lock mechanism when the comparison between the voiceprint feature vector and each authorized template in the authorized voiceprint template library fails. The security lock mechanism includes refusing to perform any call initiation operation within the lock time window T_lock, and the value of the lock time window T_lock is 30 seconds to 120 seconds.
9. The intelligent call initiation control method according to claim 1, characterized in that, The vocalization status parameter in step S3 also includes a vocalization duration D, which is counted from the start time when the voice segment is confirmed to exist. The call triggering condition also includes that the vocalization duration D is greater than or equal to a duration threshold T_D, which ranges from 0.3 seconds to 1.5 seconds.
10. A smart call-activated walkie-talkie based on voiceprint recognition, characterized in that, include: Bone conduction sensors are used to continuously collect vibration signals from the wearer's throat and convert them into electrical signals for output. The signal processing unit, electrically connected to the bone conduction sensor, is used to receive the electrical signal and amplify, filter, and perform analog-to-digital conversion on the electrical signal to generate a digital vibration signal; The voice activity detection module, connected to the signal processing unit, is used to perform frame-by-frame processing on the digital vibration signal, calculate the short-time energy value and short-time zero-crossing rate of each frame signal, and determine whether a voice segment exists based on the short-time energy value and short-time zero-crossing rate. The voiceprint recognition module, connected to the voice activity detection module, is used to extract voiceprint feature vectors when a voice segment is detected, and compare and match the voiceprint feature vectors with a pre-stored authorized voiceprint template library, and output the matching result; A vocal state analysis module, connected to the voiceprint recognition module, is used to calculate the current vocal state parameters based on the resonant frequency offset in the voiceprint feature vector when a match is successful. The vocal state parameters include the vocal cord tension index and the vocal tract opening coefficient. The call control module is connected to the sound state analysis module and the radio frequency transmission module respectively. It is used to generate a start command when the sound state parameters meet the preset call trigger conditions, and activate the radio frequency transmission module to make the walkie-talkie enter the call transmission state. The storage module is used to store the authorized voiceprint template library and the parameter configuration of the call triggering conditions.
11. The intelligent call-activated walkie-talkie according to claim 10, characterized in that, The walkie-talkie also includes an air conduction microphone, which is connected to the signal processing unit. The voice activity detection module is also used to receive the ambient audio signal collected by the air conduction microphone, identify the current ambient noise type based on the spectral characteristics of the ambient audio signal, and dynamically adjust the energy threshold T_E in the voice activity detection module according to the ambient noise type.
12. The intelligent call-activated walkie-talkie according to claim 11, characterized in that, The vocal state analysis module is further configured to perform coherence analysis on the laryngeal vibration signal collected by the bone conduction sensor and the environmental audio signal collected by the air conduction microphone when the vocal cord tension index K is less than the first threshold T_K1 and the vocal tract opening coefficient C is less than the second threshold T_C1. Based on the coherence analysis results, the module distinguishes between normal speech and interference vibration signals generated by chewing or swallowing actions, and sends the distinction results to the speech activity detection module to suppress interference vibration signals.
13. The intelligent call-activated walkie-talkie according to claim 10, characterized in that, The walkie-talkie also includes an indicator unit connected to the call control module. The indicator unit includes an LED indicator and a vibration motor. When the call control module generates a start command, the LED indicator switches from an off state to a constantly lit state and the vibration motor generates continuous vibration feedback. When the call control module generates a stop command, the LED indicator switches from a constantly lit state to an off state and the vibration motor stops vibrating.
14. A smart call activation headset based on voiceprint recognition, characterized in that, include: Bone conduction sensors are located in the contact area between the inner wall of the earphone shell and the tragus, and are used to continuously collect vibration signals from the wearer's throat. The headset microphone is used to collect the air conduction voice signals emitted by the wearer; The controller is electrically connected to the bone conduction sensor and the earphone microphone, and the controller includes a voice activity detection unit, a voiceprint recognition unit, a vocal state evaluation unit, and a call triggering unit. The speech activity detection unit is used to detect speech activity based on the laryngeal vibration signal output by the bone conduction sensor. The voiceprint recognition unit is used to extract voiceprint feature vectors when a speech segment is detected and compare them with a pre-stored authorized voiceprint template library; The vocal state evaluation unit is used to calculate the vocal cord tension index K and the vocal tract opening coefficient C when the voiceprint matching is successful. The call triggering unit is used to generate a start command when K is greater than or equal to the first threshold T_K1 and C is greater than or equal to the second threshold T_C1, thereby activating the Bluetooth radio frequency module to establish a communication link with the paired terminal.
15. The intelligent call activation headset according to claim 14, characterized in that, The earphone also includes a bone conduction speaker, which is disposed on the inner side wall of the earphone shell in the area where it contacts the auricular cartilage. The controller is also used to encode the air conduction voice signal collected by the earphone microphone and send it to the pairing terminal through the Bluetooth radio frequency module after the call triggering unit generates a start command, and to drive the bone conduction speaker to produce sound after decoding the downlink voice signal returned by the pairing terminal.
16. The intelligent call activation headset according to claim 14, characterized in that, The voiceprint recognition unit is also used to perform time delay alignment and spectrum fusion of the laryngeal vibration signal and the air conduction speech signal collected by the earphone microphone when the voiceprint matching fails, generate an enhanced speech signal, and use the enhanced speech signal to re-extract the voiceprint feature vector and perform matching again.
Citation Information
Patent Citations
A walkie-talkie voiceprint recognition method and system based on artificial intelligence
CN117198338B
Interphone capable of automatically identifying sound of owner
CN224164822U
Wireless intercom with voice identification function
CN2802848Y