A smart voice terminal with directional pick-up badge and bluetooth earphone cooperative interaction

CN122802854APending Publication Date: 2026-09-22SHENZHEN YUEHANGYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610925703.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]传统定向拾音工牌与蓝牙耳机协同交互时,声音采集端通常依赖内置拾音器件和常规音频编解码完成语音接收与信息通讯,蓝牙耳机播放提醒音频期间,耳机播放声容易沿耳侧空间扩散并被工牌麦克风重新拾取,采集端难以区分服务人员真实发声与播放漏入声,后台接收语音帧中混入非目标声音,语音识别和服务话术判断易受干扰,通信资源被无效语音占用,连续服务过程中的提醒响应和流程判断稳定性不足

Benefits of technology

[0015]本发明实施例提供的技术方案带来的有益效果至少包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802854A_ABST
    Figure CN122802854A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of voice sensing, in particular to an intelligent voice terminal with directional sound pickup service card and Bluetooth earphone cooperative interaction. The terminal comprises a calibration module, a playing interval module, a judgment module and a control module. The calibration module generates a mouth directional background reference and an environmental background reference. The playing interval module determines a playing influence time interval and an estimated leakage intensity. The judgment module calculates a playing leakage risk index and a wearer sound production confidence coefficient. The control module outputs uploaded voice frames, intercepts voice frames or temporarily stores voice frames. In the application, through joint judgment of playing time, earphone volume, link delay, mouth sound enhancement, environmental sound change and sound arrival order, the earphone playing leakage sound and the real sound production of service personnel are distinguished. During earphone playing, the real sound production uploading path is reserved, the recognition deviation caused by playing sound recovery is inhibited, resource occupation is inhibited, and the voice recognition and service speech judgment reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice sensing technology, and in particular to an intelligent voice terminal that allows directional sound pickup badges and Bluetooth headsets to interact collaboratively. Background Technology

[0002] The field of voice sensing technology mainly involves the acquisition of sound wave signals and the conversion and transmission of electrical signals. Its main purpose is to convert external sounds into digital signals for analysis and interaction, and it is widely used in industries such as communication equipment, smart homes, and IoT wearables.

[0003] Among them, the intelligent voice terminal that coordinates the interaction between the directional sound pickup badge and the Bluetooth headset refers to the traditional intelligent voice terminal. This terminal is a hardware carrier with sound acquisition and data interaction functions. It is used to capture the user's voice input and communicate with the background or surrounding nodes. It usually uses built-in sound pickup devices and conventional audio codec technology to achieve the purpose of sound reception and signal transmission.

[0004] When traditional directional microphone badges and Bluetooth headsets interact, the sound acquisition end typically relies on built-in microphones and conventional audio codecs to complete voice reception and information communication. During the Bluetooth headset's playback of reminder audio, the sound can easily diffuse along the ear side and be picked up again by the badge's microphone. The acquisition end has difficulty distinguishing between the service personnel's actual voice and the leaked sound. Non-target sounds are mixed in with the voice frames received in the background, making voice recognition and service script judgment susceptible to interference. Communication resources are occupied by invalid voice, resulting in insufficient stability in reminder response and process judgment during continuous service. Summary of the Invention

[0005] To address the technical problems existing in the prior art, this invention provides an intelligent voice terminal that allows directional sound pickup badges and Bluetooth headsets to interact collaboratively. The technical solution is as follows: On the one hand, a smart voice terminal is provided that allows directional sound pickup badges and Bluetooth headsets to interact collaboratively. This terminal includes: The calibration module is used to acquire the directional background sound intensity of the mouth and the ambient background sound intensity of the dual omnidirectional silicon microphones when the Bluetooth headset is not playing and the service personnel are silent, and to generate the directional background reference and the ambient background reference. The playback interval module is used to receive prompt audio, obtain playback start time, playback duration, Bluetooth headset volume level and Bluetooth link delay count value, determine the playback impact time interval based on the playback start time, the playback duration and the Bluetooth link delay count value, and determine the estimated leakage intensity based on the Bluetooth headset volume level. The determination module is used to acquire sound data from three microphones within the playback influence time interval, generate a speech frame, current mouth sound intensity, current ambient sound intensity, and sound arrival order based on the three microphone sound data, calculate a playback leakage risk index based on the proximity relationship between the current mouth sound intensity and the estimated leakage intensity, the time position of the speech frame within the playback influence time interval, and the sound arrival order, and calculate the wearer's voice credibility coefficient based on the enhancement relationship between the current mouth sound intensity and the mouth directional background reference, the change relationship between the current ambient sound intensity and the ambient background reference, and the sound arrival order. The control module is used to compare the playback leakage risk index with a preset interception threshold, compare the wearer's voice credibility coefficient with a preset credibility threshold, and output the uploaded voice frame, the intercepted voice frame, or the temporarily stored voice frame based on the comparison result.

[0006] As a further aspect of the present invention, the process of determining the silent calibration period specifically includes: acquiring the playback status of the Bluetooth headset, acquiring the sound intensity change status of the directional silicon microphone and the dual omnidirectional silicon microphone within a continuous sampling window, and determining the continuous sampling window as the silent calibration period when the Bluetooth headset is not playing and the sound intensity change status does not meet the preset voice trigger threshold.

[0007] As a further aspect of the present invention, the process of generating the mouth-oriented background reference specifically includes: acquiring multiple mouth-oriented background sound intensities collected by the directional silicon microphone during the silent calibration period, removing abnormal intensity values ​​that exceed a preset mutation threshold from the multiple mouth-oriented background sound intensities, and determining the mouth-oriented background reference based on the concentrated distribution results of the remaining mouth-oriented background sound intensities.

[0008] As a further aspect of the present invention, the process of generating the environmental background reference specifically includes: acquiring the first environmental background sound intensity collected by the first omnidirectional microphone and the second environmental background sound intensity collected by the second omnidirectional microphone; comparing the intensity difference between the first environmental background sound intensity and the second environmental background sound intensity; and when the intensity difference meets a preset environmental consistency threshold, determining the environmental background reference based on the first environmental background sound intensity and the second environmental background sound intensity.

[0009] As a further aspect of the present invention, the process of determining the playback impact time interval specifically includes: obtaining the playback start time and the Bluetooth link delay count value; performing delay compensation on the playback start time according to the Bluetooth link delay count value to obtain the actual sound start time; determining the actual sound end time according to the actual sound start time and the playback duration; and determining the continuous time period between the actual sound start time and the actual sound end time as the playback impact time interval.

[0010] As a further aspect of the present invention, the process of determining the estimated leakage intensity specifically includes: obtaining the Bluetooth headset volume level, reading a locally stored volume leakage mapping table, wherein the volume leakage mapping table records the mapping relationship between various Bluetooth headset volume levels and corresponding leakage intensities, and determining the leakage intensity corresponding to the Bluetooth headset volume level in the volume leakage mapping table as the estimated leakage intensity.

[0011] As a further aspect of the present invention, the calculation process of the playback leakage risk index specifically includes: calculating the intensity difference between the current sound intensity at the mouth and the estimated leakage intensity; normalizing the intensity difference to obtain a normalized leakage intensity difference; calculating the time difference between the center time of the speech frame and the center time of the playback influence time interval; normalizing the time difference to obtain a normalized playback time difference; determining the ear-side sound arrival state according to the sound arrival order; normalizing the ear-side sound arrival state to obtain a normalized ear-side arrival state value; calculating the normalized leakage intensity difference; and calculating the weighted average of the normalized playback time difference and the normalized ear-side arrival state value to obtain the playback leakage risk index.

[0012] As a further aspect of the present invention, the calculation process of the wearer's voice credibility coefficient specifically includes: calculating the difference between the current sound intensity of the mouth and the mouth intensity of the mouth directional background reference; normalizing the mouth intensity difference to obtain a normalized mouth enhancement value; calculating the difference between the current sound intensity of the environment and the environmental intensity of the environmental background reference; normalizing the environmental intensity difference to obtain a normalized environmental change value; determining the mouth sound arrival state according to the sound arrival order; normalizing the mouth sound arrival state to obtain a normalized mouth arrival state value; calculating the weighted sum of the normalized mouth enhancement value and the normalized mouth arrival state value; and calculating the weighted difference between the weighted sum and the normalized environmental change value to obtain the wearer's voice credibility coefficient.

[0013] As a further aspect of the present invention, the calculation process of the sound arrival order specifically includes: determining the sound intensity rise times of the directional microphone, the first omnidirectional microphone, and the second omnidirectional microphone within the same speech frame; calculating the first arrival time difference between the directional microphone and the first omnidirectional microphone; calculating the second arrival time difference between the directional microphone and the second omnidirectional microphone; determining the ear-side sound arrival state based on the first arrival time difference and the second arrival time difference; determining the mouth-side sound arrival state based on the first arrival time difference and the second arrival time difference; and obtaining the sound arrival order from the ear-side sound arrival state and the mouth-side sound arrival state.

[0014] As a further aspect of the present invention, the comparison result specifically includes: comparing the playback leakage risk index with a preset interception threshold, and comparing the wearer's voice credibility coefficient with a preset credibility threshold; when the playback leakage risk index does not reach the preset interception threshold and the wearer's voice credibility coefficient reaches the preset credibility threshold, an uploaded audio frame is output accordingly; when the playback leakage risk index reaches the preset interception threshold and the wearer's voice credibility coefficient does not reach the preset credibility threshold, an intercepted audio frame is output accordingly; when the playback leakage risk index does not reach the preset interception threshold and the wearer's voice credibility coefficient does not reach the preset credibility threshold, a temporarily stored audio frame is output accordingly.

[0015] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: By establishing a mouth-oriented background baseline and an environmental background baseline when the headphones are not playing and the service personnel are silent, and combining the playback start time, playback duration, and Bluetooth link latency count value when playing reminder audio, the playback impact time interval is defined, and the estimated leakage intensity is formed based on the Bluetooth headphone volume level. After the voice frame enters the judgment process, the proximity relationship between the current sound intensity of the mouth and the estimated leakage intensity, the time position, and the sound arrival order are fused into a playback leakage risk index. The mouth enhancement relationship, environmental change relationship, and sound generation order are fused into a wearer's voice credibility coefficient. Uploading, interception, and temporary storage are performed around the above results to reduce the uploading of non-target sounds and improve recognition reliability and resource utilization efficiency. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a diagram of the overall architecture of the terminal of the present invention; Figure 2 This is the background reference generation diagram for the calibration module of this invention; Figure 3 This is a diagram showing the time interval affected by the playback of the playback interval module in this invention and the estimated leakage intensity. Figure 4 This is a sound data determination diagram of the determination module of the present invention; Figure 5 This is a diagram showing the voice frame output of the control module of the present invention. Detailed Implementation

[0018] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0019] This embodiment provides an intelligent voice terminal that allows directional microphone badges and Bluetooth headsets to interact collaboratively. During continuous voice interaction when service personnel wear badges and receive prompt audio through Bluetooth headsets, the directional silicon microphone on the badge side faces the wearer's mouth to collect near-field sound, while dual omnidirectional silicon microphones are distributed on both sides of the badge to collect ambient sound. The Bluetooth headset's playback status, volume status, and link latency status are synchronously processed with the three microphones' acquisition status to prevent the Bluetooth headset's prompts from being mistakenly transmitted as the wearer's voice.

[0020] During the aforementioned operation, the prompt audio, played by the Bluetooth headset, is transmitted to the directional and omnidirectional silicon microphones via reflection from the wearer's ear, face, or clothing. The wearer's actual voice is preferentially amplified in the directional silicon microphone facing the mouth. The smart voice terminal establishes a silent background benchmark through a calibration module, determines the time range and leakage intensity reference that the prompt audio may affect microphone acquisition through a playback interval module, performs frame segmentation, intensity comparison, and arrival order identification of the three audio streams within the playback influence time interval through a judgment module, and outputs the corresponding voice frame processing result based on the leakage risk and the credibility of the wearer's voice through a control module.

[0021] Please see Figure 1 and Figure 2 The calibration module is used to acquire the directional background sound intensity of the mouth and the ambient background sound intensity of the dual omnidirectional microphones when the Bluetooth headset is not playing and the service personnel are silent, generating a directional background sound reference and an ambient background sound reference. The directional background sound intensity of the mouth refers to the non-speech background sound intensity acquired by the directional microphone when it is facing the wearer's mouth. Its source is the sound sampling results of the directional microphone during the silent calibration period, and its purpose is to serve as a reference in the subsequent judgment module to determine whether the sound in the mouth direction has increased. The ambient background sound intensity refers to the ambient sound intensity acquired by the first and second omnidirectional microphones during the same silent calibration period. It reflects the ambient noise state of the wearing location where the wearer is not speaking, and its purpose is to serve as a reference in the subsequent judgment module to determine whether the ambient sound changes with the speech frame.

[0022] In the calibration module, the silent calibration period refers to a continuous sampling window during which the Bluetooth headset is in a non-playing state and none of the three microphones exhibit voice triggering. The calibration module first obtains the Bluetooth headset's playback status, then reads the sound intensity changes of the directional silicon microphone, the first omnidirectional silicon microphone, and the second omnidirectional silicon microphone within the continuous sampling window. When the Bluetooth headset's playback status is marked as non-playing, and the sound intensity changes of the three microphones do not meet the triggering conditions corresponding to the preset voice trigger threshold, this continuous sampling window is defined as the silent calibration period. The preset voice trigger threshold is a configurable judgment condition used to distinguish between silent background fluctuations and voice activity. It is derived from silent background calibration records under the employee badge wearing environment or historical silent acquisition records. In this solution, it is not disclosed as fixed data but serves as a boundary condition for the calibration module to filter the silent window.

[0023] When generating the mouth-oriented background reference, the calibration module acquires multiple mouth-oriented background sound intensities collected by the directional microphone during the silent calibration period and performs abrupt changes checks on adjacent sampling states and the overall change state within the window. The preset abrupt change threshold is a configurable anomaly judgment condition used to exclude collision sounds, clothing friction sounds, and short external impact sounds. It originates from the device calibration results and historical anomaly sampling records of the directional microphone in silent mode. The calibration module marks the intensity values ​​exceeding the preset abrupt change threshold as abnormal intensity values ​​and excludes these abnormal intensity values ​​from the reference generation object. The remaining mouth-oriented background sound intensities are used to determine the mouth-oriented background reference according to their concentrated distribution state. The concentrated distribution state refers to the stable background interval formed by the remaining intensity values ​​within the silent window. The calibration module selects this stable background interval as the reference state for subsequent mouth enhancement judgment.

[0024] When generating an environmental background benchmark, the calibration module acquires the first ambient background sound intensity from the first omnidirectional microphone and the second ambient background sound intensity from the second omnidirectional microphone, and compares the intensity differences between the two omnidirectional acquisition results. A preset environmental consistency threshold is a configurable condition used to determine whether the two omnidirectional microphones jointly reflect the same environmental background. It is derived from the dual omnidirectional microphone assembly position, wearing posture calibration, and silent environment recording. When the aforementioned intensity difference meets the consistency condition corresponding to the preset environmental consistency threshold, the calibration module forms an environmental background benchmark based on the first and second ambient background sound intensities. When the intensity difference does not meet the consistency condition, the calibration module does not use the silent window to generate an environmental background benchmark and continues to wait for a new silent calibration period to prevent deviation of the environmental benchmark caused by unilateral obstruction, friction, or local sound sources.

[0025] During the calibration module startup phase, if the smart voice terminal has not yet established a mouth orientation background reference or environmental background reference, the calibration module marks the current background reference state as a state to be calibrated and provides this state to the playback interval module and the judgment module. In the state to be calibrated, the judgment module does not use the mouth enhancement relationship and environmental change relationship as the final output basis, but temporarily stores the three-channel audio data related to calibration until the calibration module establishes a callable background reference after meeting the silent calibration conditions. If the Bluetooth headset remains in playback mode for an extended period or the ambient sound continuously meets the voice triggering conditions during subsequent operation, the calibration module retains the previous valid background reference as a temporary reference and updates it the next time the silent calibration conditions are met.

[0026] In this embodiment, the calibration module generates a mouth direction background reference and an environmental background reference when the headphones are not playing and the service personnel are silent. This allows the subsequent sound data from the three microphones to enter the judgment module based on a reference formed under the same wearing environment, thereby avoiding the misinterpretation of fixed environmental noise in the name tag wearing scenario as voice enhancement for the wearer.

[0027] Please see Figure 1 and Figure 3 The playback interval module receives the prompt audio, obtains the playback start time, playback duration, Bluetooth headset volume level, and Bluetooth link delay count. Based on the playback start time, playback duration, and Bluetooth link delay count, it determines the playback impact time interval and the estimated leakage intensity based on the Bluetooth headset volume level. Specifically, the playback start time refers to the time when the prompt audio is scheduled to enter the Bluetooth headset playback process; the playback duration refers to the continuous playback length of the prompt audio from start to end; the Bluetooth headset volume level refers to the current loudness level of the Bluetooth headset; and the Bluetooth link delay count refers to the link delay status between the prompt audio being triggered and the actual sound being emitted by the Bluetooth headset.

[0028] When determining the playback impact time interval, the playback interval module first reads the playback start time and the Bluetooth link delay count value, and then compensates for the delay of the playback start time based on the Bluetooth link delay count value to obtain the actual sound start time. The actual sound start time refers to the initial state record of the actual acoustic output at the Bluetooth headset side after the prompt audio is transmitted via the Bluetooth link. The playback interval module then determines the actual sound end time based on the actual sound start time and playback duration, and defines the continuous time period between the actual sound start time and the actual sound end time as the playback impact time interval. The playback impact time interval refers to the time range within which the prompt audio may affect the three-microphone acquisition through ear-side leakage, facial reflection, or close-range propagation. It is output to the judgment module to limit the judgment objects of the playback leakage risk index.

[0029] When determining the estimated leakage intensity, the playback interval module reads the Bluetooth headset volume level and queries the locally stored volume leakage mapping table. The volume leakage mapping table is a set of lookup fields that records the relationship between the Bluetooth headset volume level and the corresponding leakage intensity level. It originates from the acoustic calibration records when the headset and employee ID are used together, and its stored content includes a volume level field, a leakage intensity level field, and a mapping validity status field. The playback interval module determines the leakage intensity corresponding to the current Bluetooth headset volume level in the volume leakage mapping table as the estimated leakage intensity. The estimated leakage intensity refers to the leakage reference intensity that the prompt audio may be picked up by the directional silicon microphone at the current volume level. It does not represent the wearer's vocal intensity but serves as a reference for subsequent judgment modules to determine whether the current lip sound intensity is close to the playback leakage state.

[0030] When the playback interval module cannot read the Bluetooth link latency count, it marks the playback impact time interval corresponding to the prompt audio as a latency pending confirmation state and uses the previous valid latency state as a temporary compensation basis. When there is no previous valid latency state, the playback interval module expands the range of pending determination states corresponding to the prompt audio and transmits the latency pending confirmation state along with the playback impact time interval to the determination module. After receiving the latency pending confirmation state, the determination module prioritizes the audio frames within that interval for temporary storage processing, and confirms them only after obtaining a valid latency state or playback end state. If the Bluetooth headset volume level cannot be matched in the volume leakage mapping table, the playback interval module marks the estimated leakage intensity as a mapping pending confirmation state and transmits this state to the control module, preventing the control module from directly outputting the uploaded audio frames.

[0031] In this embodiment, the playback interval module converts the playback start, duration, link delay, and volume level of the prompt audio into a playback impact time interval and estimated leakage intensity that can be called by the judgment module. This allows the acoustic leakage caused by the Bluetooth headset playback to form a clear time boundary and intensity reference before the voice frame is judged, thereby reducing the confusion between the prompt audio and the wearer's voice in the same acquisition interval.

[0032] Please see Figure 1 and Figure 4The judgment module is used to acquire sound data from three microphones within the playback influence time interval, and generate a speech frame, current mouth sound intensity, current ambient sound intensity, and sound arrival order based on the three microphone sound data. The three microphone sound data refers to the set of sound data collected by the directional microphone, the first omnidirectional microphone, and the second omnidirectional microphone under the same time reference. A speech frame refers to the sound segment to be judged, segmented from the three microphone sound data according to the same time window, carrying the frame start state, frame end state, three sound intensity states, and validity state. The current mouth sound intensity refers to the current intensity state corresponding to the directional microphone within the speech frame; the current ambient sound intensity refers to the ambient intensity state formed by the two omnidirectional microphones within the same speech frame; and the sound arrival order refers to the order in which the sound intensities of the directional microphone, the first omnidirectional microphone, and the second omnidirectional microphone increase within the same speech frame.

[0033] The decision module checks the validity of the audio data from the three microphones. Only audio data with aligned timestamps, complete sampling, no distortion markers on any of the three channels, and no consecutive missing segments is included in the speech frame generation. If a single microphone experiences a brief gap within a speech frame, the decision module marks that speech frame as incomplete and limits its subsequent processing to temporary or intercepted speech frames, preventing it from being directly output as an uploaded speech frame. If all three microphones are missing or the timestamps are misaligned, the decision module marks that time segment as invalid and outputs an invalid status to the control module, preventing the control module from uploading that segment.

[0034] When generating the sound arrival order, the determination module determines the sound intensity rise time of the directional microphone, the first omnidirectional microphone, and the second omnidirectional microphone within the same speech frame. The sound intensity rise time refers to the starting time when the sound intensity of a certain microphone transitions from a background state to a speech activity state or leaks into an active state; it is determined by the continuous change in the sound intensity of that microphone relative to the corresponding background state. The determination module compares the arrival order of the directional microphone and the first omnidirectional microphone, and also compares the arrival order of the directional microphone and the second omnidirectional microphone. Based on these two order relationships, it jointly determines the ear-side sound arrival state and the mouth-side sound arrival state. The ear-side sound arrival state refers to the state where one or both sides of the dual omnidirectional microphones experience a sound rise before the mouth-side directional path, indicating the possibility that the cue audio enters the microphone from the ear side or the environment side. The mouth-side sound arrival state refers to the state where the directional microphone experiences a sound rise before the dual omnidirectional microphones, or where the directional microphone's rise is more closely aligned with the near-field sound emission characteristics of the mouth, indicating the possibility that the wearer is emitting sound through their mouth. The determination module writes both the ear-side sound arrival state and the mouth-side sound arrival state into the sound arrival order field and outputs them to the subsequent risk and credibility determination process.

[0035] When calculating the playback leakage risk index, the judgment module first compares the proximity between the current sound intensity from the mouth and the estimated leakage intensity. This proximity refers to whether the sound intensity currently collected by the directional microphone falls near the leakage reference state corresponding to the current volume level. The closer the two are, the stronger the playback leakage correlation. When the current sound intensity from the mouth is far from the estimated leakage intensity and exhibits a near-field phonation state, the playback leakage correlation decreases. The judgment module then identifies the time position of the voice frame within the playback influence time interval. If the voice frame is located between the actual start and end times of phonation and is close to the acoustic output process of the prompt audio, the playback leakage correlation is enhanced. If the voice frame deviates from this interval or belongs to phonation activity outside the playback influence time interval, the playback leakage correlation decreases. The judgment module also calls the ear-side sound arrival state in the sound arrival sequence. When the ear-side sound arrival state is consistent with the Bluetooth headset leakage propagation path, the playback leakage correlation is enhanced. When the mouth sound arrival state occupies the priority state, the playback leakage correlation decreases.

[0036] The aforementioned playback leakage risk index is a dimensionless risk level field used to describe the degree to which a voice frame is affected by the leakage of Bluetooth headset prompts. Its inputs include the proximity relationship between the current sound intensity at the mouth and the estimated leakage intensity, the time position of the voice frame within the playback influence time interval, and the ear-side sound arrival status. When the judgment module processes different input factors with a unified standard, it does not directly mix the original sound intensity, time position, and arrival order. Instead, it first converts each input factor into the same risk level standard: the proximity relationship is converted into a leakage intensity level, the time position into a playback overlap level, and the ear-side arrival status into an ear-side arrival level. The boundaries of each level are derived from acoustic calibration, historical playback leakage records, and preset interception strategies. The closer the level is to the leakage state, the higher the corresponding risk contribution. The judgment module integrates each level according to locally preset weight configuration rules. These weight configuration rules are derived from the reliability calibration of different input factors in leakage identification and are stored as configurable fields. The integrated result is written into the playback leakage risk index field and transmitted to the control module.

[0037] When calculating the wearer's vocal credibility coefficient, the determination module first compares the enhancement relationship between the current sound intensity at the mouth and the directional background reference. This enhancement relationship refers to whether the directional microphone forms a near-field vocal state exceeding the silent mouth background within the current speech frame. When the enhancement relationship meets the near-field vocal condition, the wearer's vocal credibility is enhanced; when the enhancement relationship does not meet the near-field vocal condition, the wearer's vocal credibility is weakened. The determination module then compares the change relationship between the current ambient sound intensity and the ambient background reference. This change relationship is used to identify whether the current sound mainly comes from overall environmental changes. When environmental changes are synchronized with mouth enhancement and both omnidirectional microphones enhance the sound, the correlation of environmental interference is enhanced, suppressing the wearer's vocal credibility. When mouth-direction enhancement is significant and environmental changes do not form a co-directional interference state, the wearer's vocal credibility is enhanced. The determination module also calls the mouth sound arrival state in the sound arrival sequence. When the mouth sound arrival state is consistent with the wearer's near-field vocal path, the credibility is enhanced.

[0038] The wearer's vocal credibility coefficient is a dimensionless credibility level field used to describe whether the current speech frame originates from the wearer's mouth. Its inputs include mouth enhancement state, environmental change state, and mouth arrival state. The determination module performs unified grading on the mouth enhancement state, environmental change state, and mouth arrival state. The mouth enhancement state is converted to a mouth enhancement level, the environmental change state to an environmental interference level, and the mouth arrival state to a mouth arrival level. The mouth enhancement level and mouth arrival level contribute positively to the credibility level, while the environmental interference level imposes a negative constraint on the credibility level. The boundaries and weights of each level are derived from the speech acquisition calibration under the badge wearing state, environmental noise acquisition records, and preset credibility strategies, and are provided as configurable fields for the determination module to use. The determination module generates the wearer's vocal credibility coefficient based on the contribution direction of the above levels and outputs it to the control module.

[0039] When the judgment module processes consecutive speech frames, if the current speech frame lacks complete historical frames, the judgment module establishes a temporary judgment state using the valid frames already obtained within the current playback influence time interval. Once the number of valid frames can cover the complete speech activity segment, the judgment module switches to a stable judgment state. If a speech frame simultaneously exhibits both lip enhancement and ear-first arrival, the judgment module processes the conflict state in the order of priority based on the playback influence time interval, supplemented by the sound arrival order, and then verifies the background enhancement relationship, and transmits the conflict flag to the control module. When the control module receives the conflict flag, it does not directly use the speech frame as the upload speech frame, but instead determines the output flow direction by combining the comparison results of a preset interception threshold and a preset confidence threshold.

[0040] In this embodiment, the judgment module splits the sound data from the three microphones into speech frames, current intensity, background changes, and arrival order. It also converts the sound intensity proximity relationship, playback time position, and arrival status into a unified risk and credibility field, so that data with different physical meanings can be called by the control module under the same judgment criteria, thereby forming a basis for distinguishing between the leakage of the headphone prompt tone and the wearer's actual voice.

[0041] Please see Figure 1 and Figure 5 The control module compares the playback leakage risk index with a preset interception threshold and the wearer's voice credibility coefficient with a preset credibility threshold. Based on the comparison results, it outputs uploaded voice frames, intercepted voice frames, or temporarily stored voice frames. The preset interception threshold is a configurable boundary used to determine whether the playback leakage risk index meets the interception conditions. It is derived from Bluetooth headset prompt tone leakage calibration records, business upload strategies, and local erroneous upload suppression rules. The preset credibility threshold is a configurable boundary used to determine whether the wearer's voice credibility coefficient meets the upload conditions. It is derived from wearer's oral voice calibration records, silent background calibration status, and local voice upload strategies. Uploaded voice frames are those allowed to enter the subsequent voice upload process. Intercepted voice frames are those determined to be playback leakage or uncredible voice and not entered into the upload process. Temporarily stored voice frames are those retained for later confirmation due to insufficient evidence, state conflicts, or pending basic data confirmation.

[0042] When performing comparisons, the control module first reads the playback leakage risk index, the wearer's voice credibility coefficient, the preset interception threshold, and the preset credibility threshold. It then checks whether the judgment module has any pending confirmations (delay, mapping, incomplete frames, conflict flags, or calibration statuses). If the aforementioned pending statuses are absent or do not affect the comparison results, the control module performs a threshold comparison: when the playback leakage risk index does not reach the preset interception threshold and the wearer's voice credibility coefficient reaches the preset credibility threshold, the control module outputs the uploaded audio frame; when the playback leakage risk index reaches the preset interception threshold and the wearer's voice credibility coefficient does not reach the preset credibility threshold, the control module outputs the intercepted audio frame; when the playback leakage risk index does not reach the preset interception threshold and the wearer's voice credibility coefficient does not reach the preset credibility threshold, the control module outputs a temporarily stored audio frame.

[0043] When the playback leakage risk index reaches a preset interception threshold and the wearer's voice credibility coefficient simultaneously reaches a preset credibility threshold, the control module marks the voice frame as an overlapping pending judgment state. The overlapping pending judgment state refers to a frame state where both the prompt audio leakage feature and the wearer's voice feature exist simultaneously. This state arises from the fact that the risk field and credibility field output by the judgment module simultaneously meet the corresponding judgment conditions. In this state, the control module prioritizes checking the sound arrival order and the playback influence time interval: if the ear-side sound arrival state is highly consistent with the playback influence time interval, and the mouth enhancement relationship fails to form a stable continuous state, then an intercepted voice frame is output; if the mouth sound arrival state is continuous, and the environmental change state does not support environmental interference, then a temporary voice frame is output to wait for confirmation from adjacent voice frames. If the mouth voice credibility state remains after confirmation, the control module transfers the temporary voice frame to the upload voice frame; if the leakage risk state remains after confirmation, the control module transfers the temporary voice frame to the intercepted voice frame.

[0044] When the control module subsequently calls a temporarily stored voice frame, it retains the playback impact time interval identifier, risk field, confidence field, sound arrival order field, and abnormal status field of that voice frame. If subsequent adjacent voice frames form a continuous mouth-sounding state, the control module will jointly confirm the temporarily stored voice frame and the continuous mouth-sounding state before outputting the uploaded voice frame; if subsequent adjacent voice frames show that their playback process is consistent with the prompt audio, the control module will output an intercepted voice frame; if there is still a lack of confirmable evidence, the control module will maintain the temporary storage state until the voice frame no longer participates in the current voice upload process. The above temporary storage processing does not change the meaning of the fields already generated by the calibration module, playback interval module, and judgment module; it only changes the flow of the voice frame in the control output.

[0045] When the control module handles abnormal states, if the preset interception threshold or preset trusted threshold is unreadable, the control module outputs the current voice frame as a temporary voice frame and retains the threshold pending confirmation status. If the playback impact time interval is missing but the Bluetooth headset playback status shows that there is a prompt audio playback, the control module does not directly output the uploaded voice frame, but waits for the playback interval module to fill in the playback impact time interval. If the background reference is in a state of pending calibration, the control module uses the playback leakage risk index and the sound arrival order as temporary constraints and marks the trusted comparison result as a pending confirmation status. After the above abnormal handling is completed, the control module feeds back the output status to the judgment module, so that the judgment module continues the same status identification in subsequent voice frame processing until the calibration, delay, or mapping status is restored to usability.

[0046] In this embodiment, the control module compares the playback leakage risk index and the wearer's voice credibility coefficient with the corresponding preset judgment boundaries, and incorporates abnormal state, conflict state and temporary storage state into the output flow, so that the voice frame can form a definite control result between uploading, interception and temporary storage, so that the leakage voice frame when the Bluetooth headset plays the prompt tone will not be directly uploaded as the wearer's valid voice.

[0047] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A smart voice terminal that allows directional sound pickup badges and Bluetooth headsets to interact collaboratively, characterized in that, The terminal includes: The calibration module is used to acquire the directional background sound intensity of the mouth and the ambient background sound intensity of the dual omnidirectional silicon microphones when the Bluetooth headset is not playing and the service personnel are silent, and to generate the directional background reference and the ambient background reference. The playback interval module is used to receive prompt audio, obtain playback start time, playback duration, Bluetooth headset volume level and Bluetooth link delay count value, determine the playback impact time interval based on the playback start time, the playback duration and the Bluetooth link delay count value, and determine the estimated leakage intensity based on the Bluetooth headset volume level. The determination module is used to acquire sound data from three microphones within the playback influence time interval, generate a speech frame, current mouth sound intensity, current ambient sound intensity, and sound arrival order based on the three microphone sound data, calculate a playback leakage risk index based on the proximity relationship between the current mouth sound intensity and the estimated leakage intensity, the time position of the speech frame within the playback influence time interval, and the sound arrival order, and calculate the wearer's voice credibility coefficient based on the enhancement relationship between the current mouth sound intensity and the mouth directional background reference, the change relationship between the current ambient sound intensity and the ambient background reference, and the sound arrival order. The control module is used to compare the playback leakage risk index with a preset interception threshold, compare the wearer's voice credibility coefficient with a preset credibility threshold, and output the uploaded voice frame, the intercepted voice frame, or the temporarily stored voice frame based on the comparison result.

2. The intelligent voice terminal for collaborative interaction between a directional sound pickup badge and a Bluetooth headset as described in claim 1, characterized in that: The process of determining the silent calibration period specifically includes: obtaining the playback status of the Bluetooth headset, obtaining the sound intensity change status of the directional silicon microphone and the dual omnidirectional silicon microphone within a continuous sampling window, and determining the continuous sampling window as the silent calibration period when the Bluetooth headset is not playing and the sound intensity change status does not meet the preset voice trigger threshold.

3. The intelligent voice terminal for collaborative interaction between a directional sound pickup badge and a Bluetooth headset as described in claim 1, characterized in that: The process of generating the mouth-oriented background reference specifically includes: acquiring multiple mouth-oriented background sound intensities collected by the directional silicon microphone during the silent calibration period, removing abnormal intensity values ​​that exceed a preset mutation threshold from the multiple mouth-oriented background sound intensities, and determining the mouth-oriented background reference based on the concentrated distribution results of the remaining mouth-oriented background sound intensities.

4. The intelligent voice terminal for collaborative interaction between a directional sound pickup badge and a Bluetooth headset as described in claim 1, characterized in that: The process of generating the environmental background reference specifically includes: acquiring the first environmental background sound intensity collected by the first omnidirectional microphone and the second environmental background sound intensity collected by the second omnidirectional microphone; comparing the intensity difference between the first environmental background sound intensity and the second environmental background sound intensity; and determining the environmental background reference based on the first environmental background sound intensity and the second environmental background sound intensity when the intensity difference meets a preset environmental consistency threshold.

5. The intelligent voice terminal for collaborative interaction between a directional sound pickup badge and a Bluetooth headset as described in claim 1, characterized in that: The process of determining the playback impact time interval specifically includes: obtaining the playback start time and the Bluetooth link delay count value; performing delay compensation on the playback start time based on the Bluetooth link delay count value to obtain the actual sound start time; determining the actual sound end time based on the actual sound start time and the playback duration; and determining the continuous time period between the actual sound start time and the actual sound end time as the playback impact time interval.

6. The intelligent voice terminal for collaborative interaction between a directional sound pickup badge and a Bluetooth headset as described in claim 1, characterized in that: The process of determining the estimated leakage intensity specifically includes: obtaining the Bluetooth headset volume level, reading the locally stored volume leakage mapping table, the volume leakage mapping table recording the mapping relationship between various Bluetooth headset volume levels and corresponding leakage intensities, and determining the leakage intensity corresponding to the Bluetooth headset volume level in the volume leakage mapping table as the estimated leakage intensity.

7. The intelligent voice terminal for collaborative interaction between a directional sound pickup badge and a Bluetooth headset as described in claim 1, characterized in that: The calculation process of the playback leakage risk index specifically includes: calculating the intensity difference between the current sound intensity at the mouth and the estimated leakage intensity; normalizing the intensity difference to obtain a normalized leakage intensity difference; calculating the time difference between the center time of the speech frame and the center time of the playback influence time interval; normalizing the time difference to obtain a normalized playback time difference; determining the ear-side sound arrival state according to the sound arrival order; normalizing the ear-side sound arrival state to obtain a normalized ear-side arrival state value; calculating the normalized leakage intensity difference; and calculating the weighted average of the normalized playback time difference and the normalized ear-side arrival state value to obtain the playback leakage risk index.

8. The intelligent voice terminal for collaborative interaction between a directional sound pickup badge and a Bluetooth headset as described in claim 7, characterized in that: The calculation process for the wearer's voice credibility coefficient specifically includes: calculating the difference between the current sound intensity of the mouth and the mouth intensity of the mouth-oriented background reference; normalizing the mouth intensity difference to obtain a normalized mouth enhancement value; calculating the difference between the current sound intensity of the environment and the environmental intensity of the environmental background reference; normalizing the environmental intensity difference to obtain a normalized environmental change value; determining the mouth sound arrival state according to the sound arrival order; normalizing the mouth sound arrival state to obtain a normalized mouth arrival state value; calculating the weighted sum of the normalized mouth enhancement value and the normalized mouth arrival state value; and calculating the weighted difference between the weighted sum and the normalized environmental change value to obtain the wearer's voice credibility coefficient.

9. The intelligent voice terminal for collaborative interaction between a directional sound pickup badge and a Bluetooth headset as described in claim 7, characterized in that: The calculation process for the sound arrival order specifically includes: determining the sound intensity rise times of the directional microphone, the first omnidirectional microphone, and the second omnidirectional microphone within the same speech frame; calculating the first arrival time difference between the directional microphone and the first omnidirectional microphone; calculating the second arrival time difference between the directional microphone and the second omnidirectional microphone; determining the ear-side sound arrival state based on the first arrival time difference and the second arrival time difference; determining the mouth-side sound arrival state based on the first arrival time difference and the second arrival time difference; and obtaining the sound arrival order from the ear-side sound arrival state and the mouth-side sound arrival state.

10. The intelligent voice terminal for collaborative interaction between a directional sound pickup badge and a Bluetooth headset according to claim 1, characterized in that: The comparison results specifically include: comparing the playback leakage risk index with a preset interception threshold, and comparing the wearer's voice credibility coefficient with a preset credibility threshold. When the playback leakage risk index does not reach the preset interception threshold and the wearer's voice credibility coefficient reaches the preset credibility threshold, an uploaded audio frame is output. When the playback leakage risk index reaches the preset interception threshold and the wearer's voice credibility coefficient does not reach the preset credibility threshold, an intercepted audio frame is output. When the playback leakage risk index does not reach the preset interception threshold and the wearer's voice credibility coefficient does not reach the preset credibility threshold, a temporarily stored audio frame is output.