Intercom, earphone and control method for preventing false triggering of voice wake-up call initiation

CN122845989APending Publication Date: 2026-09-29QUANZHOU HENGLIDA TEL EQUIP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610995965.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]本发明旨在提供一种防误触发语音唤醒呼叫启动的对讲机、耳机及控制方法,以解决现有技术中耳机非佩戴状态或佩戴松动状态下因环境噪声和非佩戴者语音导致的误触发问题

Benefits of technology

[0010]第三方面,本发明提供一种防误触发语音唤醒呼叫启动的耳机,耳机与对讲机主机配合使用,耳机包括耳机壳体、压电薄膜传感器阵列、骨传导传感器、气导麦克风、加速度传感器和控制器。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845989A_ABST
    Figure CN122845989A_ABST
Patent Text Reader

Abstract

The application discloses a talkback machine, an earphone and a control method for preventing false triggering of voice wake-up call starting. The talkback machine comprises a host, a wearable earphone and a control circuit connected with the host. A piezoelectric film sensor array is embedded on the surface of the earphone shell for detecting the contact pressure distribution between the earphone and the human skin and generating a fit signal. A bone conduction sensor and an air conduction microphone are arranged inside the earphone for collecting the bone conduction vibration signal and the air conduction voice signal of the wearer respectively. A microprocessor is configured in the control circuit, and the microprocessor executes the following control logic: when the fit signal reaches a first threshold value, a primary wake-up state is started; in the primary wake-up state, activity detection is performed on the bone conduction vibration signal, and after detecting voice activity, a secondary wake-up state is started; in the secondary wake-up state, the bone conduction vibration signal and the air conduction voice signal are subjected to time synchronization verification and frequency coherence analysis, and after the verification, a wake-up instruction is generated to start the call function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication equipment technology, and specifically to a walkie-talkie, headset, and control method for preventing accidental triggering of voice wake-up call initiation. Background Technology

[0002] Walkie-talkies, as wireless communication devices, are widely used in scenarios such as police duty, fire rescue, construction, logistics scheduling, and outdoor sports. With the development of voice recognition technology, voice-activated walkie-talkies and voice-activated headsets are gradually replacing the traditional PTT button triggering method, allowing operators to initiate call functions via voice commands without freeing their hands.

[0003] Existing voice-activated walkie-talkies primarily trigger transmission via volume threshold detection or wake-word recognition. In noisy environments, ambient noise levels can easily exceed the set threshold, leading to false triggers. In multi-person communication scenarios, voice signals from non-wearers may also trigger wake-word recognition, resulting in communication channel occupancy. Some solutions employ voiceprint recognition technology to verify user identity, but voiceprint recognition requires pre-recording voice samples in a silent environment, and its accuracy significantly decreases when ambient noise overlaps with the user's voice. Other solutions incorporate bone conduction sensors in the earpiece to detect the wearer's voice activity via bone conduction signals. However, bone conduction sensors still collect ambient vibration signals when the earpiece is not in close contact with the skin or is loosely worn, making it impossible to distinguish between the wearer's and non-wearer's voices.

[0004] Existing technologies fail to address the following key issues: when the headset is not worn or is loosely worn, ambient voice signals collected by the air conduction microphone may still trigger the wake-up function; and there is a lack of effective means to verify the correlation between vibration signals collected by the bone conduction sensor when it is not in contact with the ambient voice signals collected by the air conduction microphone. These two deficiencies result in a high rate of false triggering for existing voice-activated walkie-talkies and headsets in complex usage environments. Summary of the Invention

[0005] The present invention aims to provide a walkie-talkie, headset, and control method for preventing accidental triggering of voice wake-up call initiation, in order to solve the problem of accidental triggering caused by environmental noise and non-wearer's voice when the headset is not worn or is loosely worn in the prior art.

[0006] To achieve the above objectives, the present invention provides the following technical solutions.

[0007] In a first aspect, the present invention provides a method for preventing accidental triggering of voice wake-up call initiation control, applied to a communication device including an earpiece host, comprising the following steps: Acquire contact pressure distribution data collected by the piezoelectric thin film sensor array, and generate fit characteristic values ​​based on the contact pressure distribution data; The fitting feature value is compared with a first threshold. When the fitting feature value reaches the first threshold, the communication device is controlled to enter a primary wake-up state. In the initial wake-up state, the bone conduction sensor is activated to collect bone conduction vibration signals, and the bone conduction vibration signals are used to detect voice activity. When voice activity is detected, the communication device is controlled to enter a secondary wake-up state, and the air conduction microphone is activated to collect air conduction voice signals. In the secondary wake-up state, the time-domain feature sequence of the bone conduction vibration signal and the time-domain feature sequence of the air conduction speech signal are extracted, the time cross-correlation function between the two time-domain feature sequences is calculated, and the peak cross-correlation delay value is obtained. Extract the frequency domain feature vector of the bone conduction vibration signal and the frequency domain feature vector of the air conduction speech signal, and calculate the amplitude coherence coefficient of the two frequency domain feature vectors within a set frequency band. When the peak cross-correlation delay value is less than the delay threshold and the amplitude coherence coefficient is greater than the coherence threshold, a wake-up command is generated and the call transmission function is started.

[0008] The control method of this invention constitutes a multi-level protection system through three progressively advancing discrimination steps: fit detection, bone conduction speech activity detection, and joint verification of bone conduction-air conduction dual-mode signals. The first discrimination step eliminates environmental noise interference when the earphone is not worn at the physical contact level; the second discrimination step confirms that the wearer is speaking at the vibration conduction level; and the third discrimination step confirms that the bone conduction vibration and air conduction speech originate from the same sound source at the signal correlation level. The three discrimination levels are interconnected, and if any step is not satisfied, a call will not be triggered, fundamentally solving the problem of false triggering caused by non-wearing state and non-wearer speech.

[0009] Secondly, based on the above control method, this invention provides a walkie-talkie for preventing accidental triggering of voice wake-up call initiation, comprising a main unit and an earpiece. The main unit internally includes a radio frequency transceiver module, a control circuit, and a power supply module. The earpiece is connected to the main unit via a cable or wireless link; a piezoelectric thin-film sensor array is embedded on the surface of the earpiece shell, and a bone conduction sensor and an air conduction microphone are installed inside the earpiece. A microprocessor is included in the control circuit, and the microprocessor executes the above control method.

[0010] Thirdly, the present invention provides an earphone for preventing accidental triggering of voice wake-up call activation. The earphone is used in conjunction with a walkie-talkie host. The earphone includes an earphone shell, a piezoelectric thin film sensor array, a bone conduction sensor, an air conduction microphone, an accelerometer, and a controller.

[0011] The beneficial effects of this invention are reflected in the following aspects: First, a closed-loop protection system is formed by fit detection and bone conduction-air conduction dual-mode signal verification. Fit detection ensures that the voice wake-up function is activated only when the earphone reaches sufficient contact area and pressure with the human skin; bone conduction-air conduction dual-mode signal verification verifies the homology between bone conduction vibration signals and air conduction voice signals from two dimensions: time-domain cross-correlation and frequency-domain coherence. The two modules form a positive reinforcement relationship: fit detection enables the voice verification channel, and the voice verification result inversely optimizes the fit determination parameters. This mutually reinforcing closed-loop control reduces the false trigger rate by more than two orders of magnitude compared to existing voice control solutions.

[0012] Secondly, the three-level progressive discrimination architecture ensures reliable protection against accidental touches while minimizing power consumption. In the initial wake-up state, only the piezoelectric film sensor array is activated, consuming negligible power; the bone conduction sensor is activated only after the fit is satisfactory to detect voice activity; and only after voice activity is detected are the most power-intensive air conduction microphone and dual-mode verification calculation activated. This step-by-step wake-up mechanism reduces the device's average power consumption in standby mode to less than one-tenth of that of traditional voice control solutions.

[0013] Third, dual-mode signal verification does not rely on pre-recorded fingerprint templates, so users do not need to register their voice before first use. The time cross-correlation and frequency domain coherence analysis of bone conduction vibration signals and air conduction speech signals are based on the physical correlation of the same phonation time, eliminating the possibility of the same speech content being generated at different times and from different sound sources, thus possessing natural anti-counterfeiting capabilities. Attached Figure Description

[0014] Figure 1 This is a flowchart of the control method according to an embodiment of the present invention. Detailed Implementation Example 1

[0015] This embodiment provides a walkie-talkie that prevents accidental triggering of voice wake-up call initiation.

[0016] The walkie-talkie consists of a main unit and an earpiece. The main unit contains a radio frequency transceiver module, control circuitry, and a power supply module. The radio frequency transceiver module is used to transmit and receive wireless signals, the control circuitry contains a microprocessor, and the power supply module provides operating voltage to all functional units.

[0017] The earphones connect to the main unit via a spiral cable. The earphones include an ergonomically designed shell that conforms to the shape of the outer surface of the human ear. An array of piezoelectric thin-film sensors is embedded in the surface of the earphone shell. This array consists of six piezoelectric thin-film sensing units, each made of polyvinylidene fluoride (PVDF) piezoelectric thin film material with a thickness of 0.1 mm. These six units are positioned on the curved surfaces of the earphone shell corresponding to the tragus, antitragus, concha, and earlobe. Specifically, two piezoelectric thin-film sensing units are positioned in the area corresponding to the tragus, two in the area corresponding to the antitragus, one in the area corresponding to the concha, and one in the area corresponding to the earlobe.

[0018] The earphone housing houses a bone conduction sensor, an air conduction microphone, and an accelerometer. The bone conduction sensor is a piezoelectric accelerometer with its sensitive axis perpendicular to the skin-contacting surface of the earphone housing. It collects bone conduction vibration signals transmitted through the skull to the ear when the wearer speaks. The air conduction microphone is a MEMS microphone with its pickup port facing outwards from the earphone housing. It collects air conduction speech signals from the environment. The accelerometer is a triaxial MEMS accelerometer, used to detect the earphone's motion and output a motion acceleration signal.

[0019] The microprocessor in the control circuit is electrically connected to the piezoelectric thin-film sensor array, bone conduction sensor, air conduction microphone, and accelerometer via an analog-to-digital converter interface. The microprocessor internally stores a control program that executes the control method of this invention. Example 2

[0020] Step S1: Obtain contact pressure distribution data collected by the piezoelectric thin film sensor array, and generate fit characteristic values ​​based on the contact pressure distribution data.

[0021] Six piezoelectric thin-film sensing units output analog voltage values, each proportional to the contact pressure at its corresponding location. The microprocessor acquires the voltage values ​​from each sensing unit via an analog-to-digital converter and performs the following calculations: First, the voltage values ​​of each sensing unit are normalized to the range of 0 to 1 to obtain the contact pressure coefficient of each sensing unit. The normalization method is to divide the current voltage value by the full-scale voltage value of the sensing unit. When the contact pressure coefficient of a certain sensing unit is greater than 0.5, the position is determined to be in an effective contact state.

[0022] Then, the proportion of sensor units with a contact pressure coefficient greater than 0.5 to the total number of sensor units is counted, and this proportion is taken as the bonding coverage. If five out of six sensor units have a contact pressure coefficient greater than 0.5, then the bonding coverage is five-sixths, or 0.833.

[0023] Next, the arithmetic mean of the contact pressure coefficients of all sensing units is calculated, and this arithmetic mean is used as the bonding uniformity. The contact pressure coefficients of the six sensing units are 0.92, 0.88, 0.75, 0.82, 0.91 and 0.63, respectively, so the bonding uniformity is 0.818.

[0024] Finally, the product of the coverage and the uniformity of the fit is used as the feature value of the fit. In the example above, the feature value of the fit is 0.833 multiplied by 0.818, which equals 0.681.

[0025] Step S2: Compare the fit feature value with the first threshold.

[0026] The first threshold is set to 0.5. When the fit characteristic value reaches 0.5, it indicates that the contact area and contact pressure between the earphone and the human skin meet the basic usage conditions, and the control communication device enters the primary wake-up state. In the primary wake-up state, the microprocessor activates the power supply circuit of the bone conduction sensor and begins to collect bone conduction vibration signals.

[0027] When the fit characteristic value is below 0.5, it indicates that the earphone is not being worn or is being worn very loosely. The communication device remains in sleep mode, and the piezoelectric thin film sensor array performs intermittent sampling with a period of 1 second to reduce power consumption.

[0028] Step S3: In the initial wake-up state, perform voice activity detection on the bone conduction vibration signal.

[0029] The microprocessor acquires bone conduction vibration signals at a sampling rate of 16000 Hz. The bone conduction vibration signals are processed in frames, with each frame duration set to 20 milliseconds and a frame shift set to 10 milliseconds. The short-time energy value and short-time zero-crossing rate are calculated for each frame of the bone conduction vibration signal.

[0030] The short-time energy value is calculated by summing the squares of the amplitudes of all sampling points within the frame. The short-time zero-crossing rate is calculated by counting the number of sign changes at each sampling point within the frame.

[0031] The energy threshold is set to three times the average energy of the background noise, which is obtained by moving average of the short-time energy values ​​of the most recent 100 frames. The zero-crossing rate threshold is set to 0.3.

[0032] When the short-term energy value exceeds the energy threshold for three or more consecutive frames and the short-term zero-crossing rate is lower than the zero-crossing rate threshold, speech activity is detected. This judgment condition utilizes the characteristics of concentrated energy and relatively stable frequency components of bone conduction vibration signals when the wearer speaks, which is different from the high-energy but high-zero-crossing-rate signal characteristics generated by environmental vibration and impact.

[0033] Step S4: When voice activity is detected, control the communication device to enter the secondary wake-up state, and at the same time start the air conduction microphone to collect air conduction voice signals.

[0034] In the secondary wake-up state, the microprocessor synchronously acquires bone conduction vibration signals and air conduction speech signals. Both bone conduction vibration signals and air conduction speech signals are acquired synchronously at a sampling rate of 16000 Hz, and the sampling clock is provided by the same crystal oscillator source to ensure that the sampling time of the two signals is strictly synchronized.

[0035] Step S5: In the secondary wake-up state, perform time cross-correlation verification on the bone conduction vibration signal and the air conduction speech signal.

[0036] The microprocessor downsamples the bone conduction vibration signal and the air conduction speech signal to a uniform sampling rate of 8000 Hz. Then, it normalizes the amplitude of the two downsampled signals to unify the amplitude range of the two signals to the range of -1 to 1.

[0037] Using bone conduction vibration signal as the reference sequence and air conduction speech signal as the comparison sequence, the normalized cross-correlation function is calculated within a time delay search range of ±50 milliseconds. The normalized cross-correlation function is calculated as follows: multiply the corresponding points of the bone conduction vibration signal sequence and the air conduction speech signal sequence shifted by tau sampling points, sum the results, and then divide by the square root of the product of the squares of the two signal sequences.

[0038] Determine the maximum value of the normalized cross-correlation function and its corresponding time delay value, and use this time delay value as the peak cross-correlation time delay value. The time delay threshold is set to 5 milliseconds. When the peak cross-correlation time delay value is less than 5 milliseconds, the time cross-correlation check passes.

[0039] When bone conduction vibration signals and air conduction speech signals originate from the same vocalization event, they are highly synchronized in time, with peak cross-correlation delay values ​​ranging from 0 to 2 milliseconds. When the air conduction speech signal originates from other sound sources in the environment, there is no definite synchronization relationship between the bone conduction vibration signal and the air conduction speech signal in time, the peak cross-correlation delay values ​​are randomly distributed, and the cross-correlation peak value is significantly reduced.

[0040] Step S6: Perform frequency domain amplitude coherence verification on the bone conduction vibration signal and the air conduction speech signal.

[0041] The microprocessor performs Fast Fourier Transform (FFT) on the bone conduction vibration signal and the air conduction speech signal respectively to obtain their respective power spectral density functions. The FFT has 512 points, and a Hanning window is used for windowing to reduce spectral leakage.

[0042] Within the 200 Hz to 4000 Hz frequency band, the power spectral density function is divided into a sequence of equally spaced frequency points with a frequency interval of 31.25 Hz. The cross-power spectral density of the bone conduction vibration signal power spectral density and the air conduction speech signal power spectral density is calculated at each frequency point. Then, the magnitude of the cross-power spectral density is divided by the square root of the product of the bone conduction vibration signal power spectral density and the air conduction speech signal power spectral density to obtain the amplitude coherence coefficient at each frequency point.

[0043] The amplitude coherence coefficient is calculated by dividing the magnitude of the cross-power spectral density of the bone conduction vibration signal and the air conduction speech signal at frequency f by the square root of the product of the self-power spectral density of the bone conduction vibration signal and the self-power spectral density of the air conduction speech signal.

[0044] Calculate the arithmetic mean of the amplitude coherence coefficients at all frequency points, and use this arithmetic mean as the final amplitude coherence coefficient. The coherence threshold is set to 0.7. When the amplitude coherence coefficient is greater than 0.7, the frequency domain amplitude coherence check passes.

[0045] When the wearer speaks, the bone conduction vibration signal and the air conduction speech signal have similar spectral envelope structures in the 200 Hz to 4000 Hz frequency band, with an amplitude coherence coefficient above 0.85. There is no spectral structure correlation between the air conduction speech signal generated by environmental noise sources and the wearer's own bone conduction vibration signal, with an amplitude coherence coefficient below 0.3.

[0046] Step S7: When the peak cross-correlation delay is less than 5 milliseconds and the amplitude coherence coefficient is greater than 0.7, generate a wake-up command and start the call transmission function.

[0047] The microprocessor sends a wake-up command to the radio frequency transceiver module, which then enters the transmission state and modulates the voice signal collected by the air conduction microphone before transmitting it through the antenna. Example 3

[0048] Based on Embodiments 1 and 2, this embodiment further introduces acceleration sensor data as an auxiliary basis for judging the wearing status, and adaptively updates the judgment threshold.

[0049] The microprocessor acquires the triaxial motion acceleration signal output from the accelerometer and calculates the root mean square (RMS) value of the triaxial acceleration. The acceleration threshold is set to 0.5 m / s². When the fit characteristic value reaches the first threshold and the RMS value of the motion acceleration signal is less than the acceleration threshold, it indicates that the earphone is not only in good contact with the skin but is also in a relatively static wearing state. At this time, the control communication device enters the primary wake-up state. When the RMS value of the motion acceleration signal is greater than the acceleration threshold, it indicates that the earphone is being adjusted by the wearer or is in vigorous movement. In this case, even if the fit characteristic value meets the standard, it will not enter the primary wake-up state to avoid misjudgment caused by vibration signals generated by the friction between the bone conduction sensor and the skin during the wearing and adjustment process.

[0050] The microprocessor also performs adaptive updates to the decision threshold. It records the peak cross-correlation delay and amplitude coherence coefficient for each successful wake-up command generated by the communication device within the current call cycle. The call cycle is defined as the time interval from the first generation of a wake-up command to the last release of the call transmission function. After the call cycle ends, the average peak cross-correlation delay of all successful wake-up events within the current call cycle is calculated, and this average is used as the updated delay threshold; the average amplitude coherence coefficient of all successful wake-up events within the current call cycle is also calculated, and this average is used as the updated coherence threshold. Using the call cycle average instead of a fixed threshold can adapt to the physiological differences of different users and the changes in signal characteristics under different wearing conditions.

[0051] The microprocessor also performs a delayed update of the fit determination threshold. When the fit characteristic value drops from reaching the first threshold to below the second threshold, the communication device is controlled to exit the primary wake-up state and the power supply to the bone conduction sensor and air conduction microphone is turned off. The second threshold is set to 0.3, which is less than the first threshold of 0.5. The hysteresis interval between the first and second thresholds prevents the fit characteristic value from fluctuating repeatedly around the threshold due to slight displacement of the earphone, thus avoiding frequent start-up and shutdown of the voice detection function.

[0052] In the three-level progressive discrimination architecture of this invention, there is a mutually reinforcing closed-loop relationship between fit detection and bone conduction-air conduction dual-mode verification. Fit detection ensures good acoustic coupling between the bone conduction sensor and the skin, giving the bone conduction vibration signal a sufficient signal-to-noise ratio, providing a reliable signal foundation for dual-mode verification. The pass rate of dual-mode verification, in turn, reflects the actual effect of fit. When the average pass rate of dual-mode verification during a call cycle is lower than a preset value, the microprocessor automatically raises the first threshold from 0.5 to 0.6, requiring a tighter fit to activate the voice wake-up function. This adaptive adjustment mechanism allows the device to dynamically optimize its protection strategy based on actual usage, minimizing stringent requirements on wearing position while ensuring the reliability of accidental touch prevention, thus improving the user experience.

Claims

1. A method for preventing accidental triggering of voice wake-up call initiation control, applied to a communication device including an earpiece host, characterized in that, Includes the following steps: Acquire contact pressure distribution data collected by the piezoelectric thin film sensor array, and generate fit characteristic values ​​based on the contact pressure distribution data; The fitting feature value is compared with a first threshold. When the fitting feature value reaches the first threshold, the communication device is controlled to enter a primary wake-up state. In the initial wake-up state, the bone conduction sensor is activated to collect bone conduction vibration signals, and the bone conduction vibration signals are used to detect voice activity. When voice activity is detected, the communication device is controlled to enter a secondary wake-up state, and the air conduction microphone is activated to collect air conduction voice signals. In the secondary wake-up state, the time-domain feature sequence of the bone conduction vibration signal and the time-domain feature sequence of the air conduction speech signal are extracted, the time cross-correlation function between the two time-domain feature sequences is calculated, and the peak cross-correlation delay value is obtained. Extract the frequency domain feature vector of the bone conduction vibration signal and the frequency domain feature vector of the air conduction speech signal, and calculate the amplitude coherence coefficient of the two frequency domain feature vectors within a set frequency band. When the peak cross-correlation delay value is less than the delay threshold and the amplitude coherence coefficient is greater than the coherence threshold, a wake-up command is generated and the call transmission function is started.

2. The control method according to claim 1, characterized in that, The step of generating fit feature values ​​based on the contact pressure distribution data includes: The voltage values ​​of each sensing unit in the piezoelectric thin film sensor array are normalized to the range of 0 to 1 to obtain the contact pressure coefficient of each sensing unit. The proportion of sensor units with a contact pressure coefficient greater than 0.5 to the total number of sensor units is counted, and this proportion is used as the adhesion coverage. Calculate the arithmetic mean of the contact pressure coefficients of all sensing units, and use this arithmetic mean as the uniformity of the fit. The product of the adhesion coverage and the adhesion uniformity is used as the adhesion feature value.

3. The control method according to claim 1, characterized in that, The step of detecting speech activity using the bone conduction vibration signal includes: The bone conduction vibration signal is processed by framing, with each frame duration set to 20 milliseconds and the frame shift set to 10 milliseconds; Calculate the short-time energy value of each frame of bone conduction vibration signal; Calculate the short-time zero-crossing rate of the bone conduction vibration signal for each frame; When the short-time energy value exceeds the energy threshold for three or more consecutive frames and the short-time zero-crossing rate is lower than the zero-crossing rate threshold, it is determined that voice activity has been detected.

4. The control method according to claim 1, characterized in that, The steps of extracting the time-domain feature sequences of the bone conduction vibration signal and the air conduction speech signal, calculating the time cross-correlation function between the two time-domain feature sequences, and obtaining the peak cross-correlation delay value include: The bone conduction vibration signal and the air conduction speech signal are downsampled to a uniform sampling rate. Normalize the amplitude of the two downsampled signals respectively; Using the bone conduction vibration signal as a reference sequence and the air conduction speech signal as a comparison sequence, a normalized cross-correlation function is calculated within a time delay search range of ±50 milliseconds. Determine the maximum value of the normalized cross-correlation function and its corresponding time delay value, and use the time delay value as the peak cross-correlation time delay value.

5. The control method according to claim 1, characterized in that, The step of extracting the frequency domain feature vectors of the bone conduction vibration signal and the air conduction speech signal, and calculating the amplitude coherence coefficients of the two frequency domain feature vectors within a set frequency band includes: The bone conduction vibration signal and the air conduction speech signal are respectively subjected to fast Fourier transform to obtain their respective power spectral density functions; Within the 200 Hz to 4000 Hz frequency band, the power spectral density function is divided into an equally spaced sequence of frequency points; Calculate the cross-power spectral density of the bone conduction vibration signal and the air conduction speech signal at each frequency point; Divide the magnitude of the cross power spectral density by the square root of the product of the power spectral density of the bone conduction vibration signal and the power spectral density of the air conduction speech signal to obtain the amplitude coherence coefficient at each frequency point. Calculate the arithmetic mean of the amplitude coherence coefficients at all frequency points, and use this arithmetic mean as the final amplitude coherence coefficient.

6. The control method according to claim 1, characterized in that, The method further includes: When the fit feature value drops from the first threshold to below the second threshold, the communication device is controlled to exit the primary wake-up state and the power supply to the bone conduction sensor and air conduction microphone is turned off, where the second threshold is less than the first threshold. Record the peak cross-correlation delay value and the amplitude coherence coefficient when the communication device successfully generates a wake-up command in the current call cycle. Calculate the average peak cross-correlation delay value of all successful wake-up events in the current call cycle as the updated delay threshold. Calculate the average amplitude coherence coefficient of all successful wake-up events in the current call cycle as the updated coherence threshold.

7. A walkie-talkie with anti-accidental voice wake-up call activation, characterized in that, include: The host computer is equipped with an RF transceiver module, a control circuit, and a power supply module. The earphone is connected to the host via a cable or wireless link. The earphone shell surface is embedded with a piezoelectric thin film sensor array, and the earphone is equipped with a bone conduction sensor and an air conduction microphone. The piezoelectric thin film sensor array is used to detect the contact pressure distribution between the headphones and human skin and generate pressure distribution electrical signals. The bone conduction sensor is used to collect bone conduction vibration signals transmitted through the skull when the wearer speaks; The air conduction microphone is used to collect air conduction speech signals in the environment; The control circuit includes a microprocessor, which is electrically connected to the piezoelectric thin film sensor array, the bone conduction sensor, and the air conduction microphone. The microprocessor executes the control method according to any one of claims 1 to 6.

8. The walkie-talkie according to claim 7, characterized in that, The piezoelectric thin film sensor array consists of at least four piezoelectric thin film sensing units, which are respectively arranged in the curved areas where the earphone shell contacts the tragus, antitragus, concha, and earlobe.

9. The walkie-talkie according to claim 7, characterized in that, The earphone is also equipped with an accelerometer, which is used to detect the motion state of the earphone and output a motion acceleration signal; the microprocessor is also used to acquire the motion acceleration signal, and when the fit feature value reaches the first threshold and the root mean square value of the motion acceleration signal is less than the acceleration threshold, control the communication device to enter the primary wake-up state.

10. A headset for preventing accidental voice wake-up call activation, characterized in that, The headset is used in conjunction with the walkie-talkie main unit, and the headset includes: The earphone housing has a piezoelectric thin film sensor array embedded on its surface to detect the contact pressure distribution between the earphone and human skin and generate a pressure distribution electrical signal. A bone conduction sensor is installed inside the earphone housing to collect bone conduction vibration signals transmitted through the skull when the wearer speaks; An air conduction microphone is installed inside the earphone housing to collect air conduction voice signals from the environment; An accelerometer sensor is installed inside the earphone housing to detect the motion state of the earphone and output a motion acceleration signal; The controller, electrically connected to the piezoelectric thin-film sensor array, the bone conduction sensor, the air conduction microphone, and the accelerometer, is used for: Based on the pressure distribution electrical signal, a fit characteristic value is generated; When the fit feature value reaches the first threshold and the root mean square value of the motion acceleration signal is less than the acceleration threshold, the earphone is controlled to enter the primary wake-up state. The bone conduction sensor is activated to detect voice activity in the primary wake-up state. When voice activity is detected, the air conduction microphone is activated to collect air conduction voice signals; The bone conduction vibration signal and the air conduction voice signal are subjected to time cross-correlation verification and frequency domain amplitude coherence verification. When both verifications pass, a wake-up command is generated and sent to the walkie-talkie host to start the call transmission function.

11. The earphone according to claim 10, characterized in that, The controller is also used for: When performing time cross-correlation verification on the bone conduction vibration signal and the air conduction speech signal, the bone conduction vibration signal is used as the reference sequence and the air conduction speech signal is used as the comparison sequence. The normalized cross-correlation function is calculated within a time delay search range of ±50 milliseconds to obtain the peak cross-correlation time delay value. When performing frequency domain amplitude coherence verification on the bone conduction vibration signal and the air conduction speech signal, the amplitude coherence coefficient of the two signals is calculated in the frequency band from 200 Hz to 4000 Hz. When the peak cross-correlation delay is less than 5 milliseconds and the amplitude coherence coefficient is greater than 0.7, both checks are considered to have passed.

12. The earphone according to claim 10, characterized in that, The piezoelectric thin film sensor array includes six piezoelectric thin film sensing units, of which two piezoelectric thin film sensing units are arranged in the area corresponding to the tragus of the earphone shell, two piezoelectric thin film sensing units are arranged in the area corresponding to the antitragus of the earphone shell, one piezoelectric thin film sensing unit is arranged in the area corresponding to the concha of the earphone shell, and one piezoelectric thin film sensing unit is arranged in the area corresponding to the earlobe of the earphone shell.