A wearable target speech extraction method based on electro-acoustic gate diagram and speech signal fusion

CN122531360APending Publication Date: 2026-08-07BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2026-05-12
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

针对现有技术的不足,本发明的目的在于提供一种基于电声门图和语音信号融合的佩戴式目标语音提取方法及系统,旨在解决现有技术中在高噪声、多人干扰、佩戴防护装备及移动作业等复杂场景下,单纯依赖空气传导麦克风采集方式导致目标语音易受环境噪声和非目标说话人干扰、传统目标语音提取方法依赖预先录制参考语音或适应语导致部署复杂,以及现有方案难以兼顾双模态辅助表征和低时延连续输出的问题,从而实现对对应佩戴者目标语音的稳定提取和流式输出

Benefits of technology

本发明将电声门图信号与语音信号进行融合处理,利用与佩戴者发声活动直接相关的电声门图信号作为辅助信息,相比仅依赖空气传导麦克风的技术方案,能够提高复杂环境下目标语音提取的稳定性和抗干扰能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531360A_ABST
    Figure CN122531360A_ABST
Patent Text Reader

Abstract

The application discloses a wearable target speech extraction method based on electroacoustic gate diagram and speech signal fusion, comprising the following steps: S1: collecting the electroacoustic gate diagram signal and the speech signal in the sound production process of a target user by a terminal; S2: generating double-mode input data to be transmitted from the electroacoustic gate diagram signal and the speech signal; S3: sending the double-mode input data to an inference device; S4: pre-processing the signal by the inference device to obtain electroacoustic gate diagram features and speech features; S5: inputting an auxiliary representation network to generate representation information; S6: inputting a streaming target speech extraction network; S7: outputting the target speech extraction result corresponding to the current data block and updating the historical cache state; and S8: splicing the target speech extraction results of continuous data blocks to form a continuous output target speech signal. The application can utilize the electroacoustic gate diagram signal directly related to the sound production activity of a wearer to assist target speech extraction and reduce the dependence on pre-recorded reference speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of speech signal processing, wearable intelligent sensing and edge computing technology, and specifically relates to a wearable target speech extraction method based on the fusion of electroglot diagram and speech signal. Background Technology

[0002] In multi-person collaborative work and complex environment voice communication scenarios, on-site personnel usually need to complete the transmission of instructions, status reporting and real-time collaboration under conditions such as continuous noise, personnel movement, multiple people speaking at the same time and personal protective equipment obstruction. Public information shows that high-level noise can not only cause hearing damage, but also interfere with communication and attention, and make it difficult for personnel to hear warning signals, thereby increasing the risk of accidents and injuries; when occupational noise exposure reaches a high level, voice communication itself will be significantly affected.

[0003] The aforementioned problems are not limited to a single industry but are widespread across multiple scenarios. For example, in collaborative communication scenarios where respirators or face shields are worn, respirators, masks, or hoods reduce speech intelligibility and increase the difficulty of verbal communication; in high-noise on-site operations such as construction, underground mining, tunnel maintenance, and heavy equipment maintenance, environmental noise makes it harder for personnel to hear warning signals and hinders communication; in mobile vehicles, cockpits, cabins, or other enclosed compartments, excessive environmental noise interferes with auditory feedback and voice communication; and in mobile response scenarios such as ambulance transport and on-site first aid, sirens, engine noise, and vehicle noise further reduce the quality of on-site communication.

[0004] To address the need for speech extraction in complex environments, existing technologies typically employ purely acoustic approaches to achieve speech enhancement, speech separation, or target speaker extraction. One approach relies on a single microphone or multi-microphone array to collect airborne speech, and then uses noise reduction, beamforming, or deep learning models to suppress background noise and interference with the speaker. Another approach identifies the target speaker through reference speech, speaker embedding, or other auxiliary cues. Relevant reviews indicate that target speaker extraction often requires additional target speaker cues.

[0005] Furthermore, electroglottography, as a non-invasive laryngeal signal detection technology, can detect changes in laryngeal impedance during phonation by placing electrodes on the neck, thereby indirectly reflecting the contact and vibration state of the vocal cords. Compared with traditional air-conducted speech signals, electroglottography signals are directly derived from the physiological activities of the vocal organs. In complex environments, it can provide auxiliary information for speech extraction of the target speaker that is different from pure acoustic signals. Therefore, combining electroglottography signals with speech signals acquired by microphones has become a noteworthy technical direction in the field of speech processing in complex environments.

[0006] However, existing technologies still have the following problems in practical applications: Limited ability to resist complex interference: Most existing speech extraction solutions mainly rely on air conduction microphones to collect speech. In scenarios with strong noise, multiple people speaking at the same time, obvious reverberation, or continuous movement, they are easily affected by environmental noise coupling, crosstalk from non-target speakers, etc., which leads to a decrease in the accuracy of target speech extraction.

[0007] It is highly dependent on additional auxiliary conditions: Some existing target speech extraction solutions require pre-recording of reference speech, speaker registration information or other external auxiliary information. In multi-person collaboration, dynamic operation and emergency scenarios, there are problems such as complex deployment, insufficient real-time performance and inconvenience of use.

[0008] There is a lack of integrated solutions suitable for wearable scenarios: most existing related technologies focus on a single algorithm or only on a single signal acquisition method. There is a lack of a complete target speech extraction solution that can adapt to neck-worn scenarios, simultaneously acquire electroglottic signals and speech signals, and work in collaboration with back-end intelligent inference through wireless links.

[0009] Insufficient adaptability to multi-person collaborative communication scenarios: In multi-person collaborative work scenarios, the system often needs to stably extract the corresponding voice of each wearer in order to improve the clarity and reliability of team communication. Existing technologies still lack the ability to stably extract the target voice of the corresponding wearer in such complex environments. Summary of the Invention

[0010] (1) Technical problems to be solved To address the shortcomings of existing technologies, the present invention aims to provide a wearable target speech extraction method and system based on electroglot diagram and speech signal fusion. This system addresses the problems in existing technologies, such as the susceptibility of target speech to environmental noise and interference from non-target speakers when relying solely on air-conducting microphones in complex scenarios like high noise, multi-person interference, wearing protective equipment, and mobile operations; the complexity of deployment due to pre-recorded reference speech or adaptation phrases in traditional target speech extraction methods; and the difficulty in simultaneously achieving dual-modal auxiliary representation and low-latency continuous output. The goal is to achieve stable extraction and streaming output of the target speech of the corresponding wearer.

[0011] (2) Technical solution To address the aforementioned technical problems, this invention provides a wearable target speech extraction method based on the fusion of electroglottic diagrams and speech signals, comprising the following steps: S1: The electroglottic signal and speech signal of the target user during the vocalization process are simultaneously acquired by the acquisition terminal worn around the user's neck. The electroglottic signal is used to characterize the change in the target user's laryngeal impedance or the state of vocal cord vibration. The speech signal is the original audio signal containing the target user's speech acquired by the microphone. S2: Sample, buffer, time-align, and encapsulate the electroglot signal and the speech signal to generate dual-modal input data to be transmitted; S3: The dual-modal input data is sent to the inference device via the wireless transmission module; S4: The inference device preprocesses the electroglottic signal and the speech signal respectively to obtain electroglottic features and speech features; S5: Input the electroglottic features into the auxiliary representation network to generate target speaker representation information, and fuse the target speaker representation information with the speech features; S6: Input the speech features of the current data block, the target speaker representation information, and the historical cache state into the streaming target speech extraction network to obtain the target speech extraction result corresponding to the current data block; S7: Output the target speech extraction result corresponding to the current data block, and update the historical cache state; S8: After splicing the target speech extraction results of continuous data blocks, a continuous output target speech signal is formed.

[0012] Preferably, the acquisition terminal in step S1 is a wearable acquisition device that can be attached to or surround the user's neck, and the electroglottic signal is acquired by electrode components disposed on both sides of the user's throat.

[0013] Furthermore, in step S2, time alignment of the electroglot diagram signal and the speech signal includes: synchronous sampling based on a unified clock trigger, or alignment processing of the two sampled signals based on timestamps.

[0014] Furthermore, the wireless transmission module in step S3 is a Wi-Fi module, a Bluetooth module, or other short-range wireless communication module.

[0015] Furthermore, in step S4, the preprocessing of the electroglottic signal includes one or more of the following: demodulation, filtering, envelope extraction, normalization, framing, time-frequency transformation, and fundamental frequency correlation feature extraction.

[0016] Furthermore, in step S5, the auxiliary representation network generates a target speaker representation vector based on the electroglottic features acquired synchronously with the speech signal. The target speaker representation vector is used to replace or partially replace the speaker prior information constructed based on pre-recorded reference speech or adaptive speech in traditional target speech extraction methods.

[0017] Furthermore, in step S5, the target speaker representation information is used to conditionally modulate at least one feature extraction layer in the streaming target speech extraction network to enhance the speech components of the target user and suppress interfering speech and noise components; the conditional modulation includes generating scaling factors and bias factors based on the target speaker representation information, and adaptively scaling and biasing the speech features of the feature extraction layer.

[0018] Furthermore, the electroglot signal and the speech signal are input into the streaming target speech extraction network in the form of time-aligned data blocks, and the streaming target speech extraction network outputs the target speech extraction result corresponding to the current data block when processing the current data block.

[0019] Furthermore, the streaming target speech extraction network calls the historical cache state retained in the previous data block when processing the current data block, and updates the historical cache state after outputting the target speech extraction result corresponding to the current data block; the historical cache state includes at least one of encoder cache state, convolutional layer cache state or feature memory state.

[0020] Furthermore, the inference device is an edge computing device, embedded processing device, or deep learning inference board with neural network inference capabilities, used to run the target speech extraction model and output the target speech.

[0021] Furthermore, the target voice signal output in step S8 is used for voice communication, voice recognition front-end enhancement, command recognition, playback, storage, or forwarding.

[0022] This invention also provides a wearable target speech extraction system based on the fusion of electroglottic diagrams and speech signals, characterized in that it includes: Wearable acquisition terminal, used to be worn around the user's neck, to simultaneously acquire electroglottic signals and voice signals; The synchronization processing module is used to sample, buffer, time-align, and encapsulate the acquired electroglottic and speech signals. The wireless transmission module is used to send the packaged dual-modal input data to the inference device; The inference device is used to receive the dual-modal input data, extract features from the electroglottic signal and the speech signal, and generate target speaker representation information based on the electroglottic features through an auxiliary representation network. The output module is used to output, play, store, or send the target speech.

[0023] Furthermore, the wearable data collection terminal includes: Electroglottic acquisition unit is used to acquire the user's throat impedance change signal; The voice acquisition unit is used to acquire the raw audio signal during the user's speech process; The control unit is used to perform sampling control, buffer management, and wireless transmission control.

[0024] Furthermore, the inference device includes a preprocessing module, an auxiliary representation module, a streaming fusion extraction module, a cache state management module, and a result output module; the auxiliary representation module is used to generate target speaker representation information based on electroglottic features; the streaming fusion extraction module is used to extract the target speech corresponding to the current data block based on the speech features of the current data block, the target speaker representation information, and the historical cache state; and the cache state management module is used to update and save the historical cache state.

[0025] Furthermore, the system includes multiple wearable acquisition terminals, which are worn by different users to acquire the electroglottic signal and speech signal of the corresponding user, and output the target speech of the corresponding user.

[0026] Beneficial effects Compared with the prior art, the beneficial effects of the present invention are as follows: This invention fuses electroglottic signals with speech signals, using electroglottic signals directly related to the wearer's vocalization as auxiliary information. Compared to technical solutions that rely solely on air conduction microphones, this invention improves the stability and anti-interference capability of target speech extraction in complex environments.

[0027] This invention generates target speaker representation information based on electroglottic features through an auxiliary representation network, thereby replacing or partially replacing the speaker prior information constructed based on pre-recorded reference speech or adaptation language in traditional target speech extraction methods. This reduces the system's dependence on pre-recorded reference speech and improves deployment convenience.

[0028] This invention introduces target speaker representation information into a streaming target speech extraction network and combines it with historical cache status to achieve continuous output of target speech in data blocks, thereby balancing target speech extraction accuracy and low-latency real-time processing capabilities.

[0029] This invention constructs a complete technical link for the collaborative operation of a wearable acquisition terminal, a wireless transmission link, and a back-end streaming inference device, which is suitable for voice communication scenarios in complex environments such as high noise, multiple interferences, wearing protective equipment, and mobile operations.

[0030] This invention can utilize electroglottic signals directly related to the wearer's vocal activity to assist in target speech extraction, reduce dependence on pre-recorded reference speech, and take into account both the accuracy of target speech extraction and low-latency continuous output capability. It is suitable for voice communication scenarios with high noise, multiple interference, and wearing protective equipment. Attached Figure Description

[0031] Figure 1 A schematic diagram of the overall structure of a wearable target speech extraction system based on the fusion of electroglottic diagram and speech signal is provided for an embodiment of the present invention; Figure 2 A flowchart illustrating a streaming target speech extraction method based on the fusion of electroglottic diagrams and speech signals, provided for an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a wearable data collection terminal provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of a process for synchronous acquisition and wireless transmission of electroglottic signal and voice signal provided in an embodiment of the present invention; Figure 5 A schematic diagram of a target speech extraction model structure including an auxiliary representation network, a streaming target speech extraction network, and historical cache state management is provided for an embodiment of the present invention. Figure 6 This is a schematic diagram of system deployment in a multi-person collaborative communication scenario provided by an embodiment of the present invention; Figure 7 This is a schematic diagram of experimental results for a wearable target speech extraction method based on the fusion of electroglottic diagram and speech signal, provided in an embodiment of the present invention. Detailed Implementation

[0032] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. Various modifications or equivalent substitutions made to the present invention by those skilled in the art without departing from the spirit and substance of the present invention should fall within the scope of protection of the present invention.

[0033] like Figure 1 As shown, this embodiment provides a wearable target speech extraction system based on the fusion of electroglottic diagrams and speech signals, including a wearable acquisition terminal, a synchronous processing module, a wireless transmission module, an inference device, and an output module.

[0034] The wearable acquisition terminal is worn around the user's neck and simultaneously acquires electroglottic and speech signals during the user's vocalization. The synchronous processing module samples, buffers, times-aligns, and encapsulates the two acquired signals. The wireless transmission module sends the encapsulated dual-modal input data to the inference device. The inference device preprocesses, extracts features, and fuses the received electroglottic and speech signals to output the target speech corresponding to the wearer. The output module plays, stores, displays, forwards, or provides the extracted target speech for subsequent speech recognition modules.

[0035] In this embodiment, as Figure 3 As shown, the wearable acquisition terminal includes an electrode assembly, a microphone assembly, a signal conditioning and sampling assembly, a control assembly, a wireless communication assembly, and a power supply assembly. The electrode assembly is located near the throat region and is used to acquire electroglottic signals. The microphone assembly is used to acquire raw audio signals containing target speech and ambient sound. The signal conditioning and sampling assembly is used to amplify, filter, perform analog-to-digital conversion, or perform other front-end processing on the electroglottic signals and speech signals. The control assembly is used to perform system timing control, data buffering, and transmission control. The wireless communication assembly is used to send the acquired data to an external inference device.

[0036] In some embodiments, the electrode assembly may employ surface electrodes disposed on both sides of the thyroid cartilage region. It acquires laryngeal impedance changes by applying a high-frequency AC excitation signal to the laryngeal tissue, and injects a constant-amplitude excitation current into the laryngeal tissue. The measured response voltage at both ends of the neck was Then the equivalent impedance at the corresponding time can be expressed as: in, The impedance sequence represents the equivalent impedance of the larynx as it changes over time. Since changes in the contact state of the vocal cords during phonation cause changes in the larynx impedance, the impedance sequence can be used to characterize the phonation activity of the corresponding wearer.

[0037] In some embodiments, to improve the purity of the excitation signal, an RC low-pass filter can be set in the excitation output stage, and its cutoff frequency satisfies: in, For filtering resistors, For filtering capacitors, This is the filter cutoff frequency. Adjustment is made... and The value of can suppress high-frequency spurious components in the excitation signal.

[0038] In some embodiments, the response signal detection module may employ a logarithmic ratio detection method to demodulate the amplitude difference between the reference carrier and the measurement carrier, and its output satisfies the following approximate relationship: in, Indicates the input amplitude of the measurement channel. Indicates the reference channel input amplitude. This represents the proportionality coefficient. Through this logarithmic ratio calculation, the carrier amplitude change caused by impedance variation can be converted into a low-frequency voltage output for sampling by subsequent digital processing modules.

[0039] In this embodiment, the inference device includes a preprocessing module, an auxiliary representation module, a streaming fusion extraction module, a cache state management module, and a result output module; wherein, the auxiliary representation module is used to generate target speaker representation information based on electroglottic features, the streaming fusion extraction module is used to extract the target speech corresponding to the current data block based on the speech features of the current data block, the target speaker representation information, and the historical cache state, and the cache state management module is used to update and save the historical cache state.

[0040] like Figure 2 and Figure 4 As shown, this embodiment provides a wearable target speech extraction method based on the fusion of electroglot diagrams and speech signals, including the following steps: Step S1: Synchronously acquire the electroglottic signal and voice signal of the target user during the vocalization process through the acquisition terminal worn around the user's neck.

[0041] Among them, the data collection terminal is a wearable data collection device that can be attached to or wrapped around the user's neck; Electroglottic signals are acquired by an electrode assembly placed in the larynx region and are used to reflect the dynamic changes in larynx impedance or vocal cord vibration state during the user's vocalization process. The voice signal is acquired by the microphone assembly and includes the target user's voice as well as the raw audio signal containing ambient noise, other people's voices, or reverberation components.

[0042] Step S2: Sample, buffer, time-align, and encapsulate the electroglot signal and the speech signal to generate bimodal input data to be transmitted.

[0043] In this step, time alignment of the electroglot signal and the speech signal includes: synchronous sampling can be triggered by a unified clock, or the two sampled data streams can be aligned using timestamps. The aligned dual-modal data is encapsulated into data packets suitable for wireless transmission for subsequent inference device reception and processing.

[0044] Step S3: Send the dual-modal input data to the inference device via the wireless transmission module.

[0045] The wireless transmission module can be Wi-Fi, Bluetooth, or other short-range wireless communication modules.

[0046] During the transmission of dual-modal input data, terminal identifiers, time sequence identifiers, or frame number information can be attached to different data packets so that the inference device can identify, sort, and reassemble the data.

[0047] Step S4: The inference device preprocesses the electroglottic signal and the speech signal respectively to obtain electroglottic features and speech features.

[0048] In this step, the preprocessing of the electroglot diagram signal includes one or more of the following: demodulation, filtering, envelope extraction, normalization, framing, time-frequency transformation, and fundamental frequency correlation feature extraction.

[0049] The preprocessing of the electroglottic signal may include demodulation, filtering, envelope extraction, normalization, framing, and extraction of time-domain or time-frequency features; the preprocessing of the speech signal may include framing, windowing, short-time Fourier transform, time-domain convolutional coding, or normalization.

[0050] In one embodiment, the original electroglottic signal After envelope extraction and normalization, the electroglottic features are obtained. , can be represented as: in, This indicates the envelope extraction operation. Represents the short-time Fourier transform. This indicates normalization processing.

[0051] Step S5: Input the electroglottic features into the auxiliary representation network to generate target speaker representation information, and fuse the target speaker representation information with the speech features.

[0052] Step S6: Input the speech features of the current data block, the target speaker representation information, and the historical cache state into the streaming target speech extraction network to obtain the target speech extraction result corresponding to the current data block.

[0053] In steps S5 and S6, the auxiliary representation network can convert the electroglottic map features into a target speaker representation vector related to the corresponding wearer's vocal state through convolutional networks, attention mechanisms, or other feature mapping structures. The target speaker representation vector can be expressed as: in, This indicates the electroglottic feature corresponding to the current data block. This represents the auxiliary representation network. This represents the target speaker representation vector corresponding to the current data block. The target speaker representation vector is used to replace or partially replace the speaker prior information constructed based on pre-recorded reference speech or adaptation language in traditional target speech extraction methods.

[0054] In one embodiment, the entire streaming target speech extraction model can be formally represented as: in, This indicates the mixed speech signal corresponding to the current data block. Indicates a voice encoder. This represents the target speaker representation vector corresponding to the current data block. This indicates the historical cache state retained from the previous data block. This represents a streaming target speech extraction network. Indicates decoder, This indicates the target speech extraction result corresponding to the current data block.

[0055] In some embodiments, the streaming target speech extraction network may employ a temporal causal network. The network first maps the mixed speech waveform of the current data block to speech features through an encoder, then combines the target speaker representation information of the current data block with the historical cache state retained from the previous data block to perform feature enhancement, and finally reconstructs the target speech waveform corresponding to the current data block through a decoder.

[0056] In one embodiment, in the streaming target speech extraction network... The layer will store the speech features of the current data block. Representation of the target speaker Perform adaptive fusion to obtain the fused features. It can be represented as: in, This represents the scaling factor generated by the target speaker's representation. This represents the bias factor generated by the target speaker's representation. This indicates element-wise multiplication.

[0057] In this way, while maintaining the core features of the speech, the prior information of the wearer's vocalization provided by the glottal diagram can be injected into the streaming target speech extraction process.

[0058] In step S5, the target speaker representation information is used to conditionally modulate at least one feature extraction layer in the streaming target speech extraction network to enhance the speech components of the target user and suppress interfering speech and noise components; the conditional modulation includes generating scaling factors and bias factors based on the target speaker representation information, and adaptively scaling and biasing the speech features of the feature extraction layer.

[0059] The electroglottic signal and the speech signal are input into the streaming target speech extraction network in the form of time-aligned data blocks.

[0060] Step S7: Output the target speech extraction result corresponding to the current data block and update the historical cache status; Step S8: After splicing the target speech extraction results of the continuous data blocks, a continuous output target speech signal is formed.

[0061] In steps S7 and S8, the streaming target speech extraction network outputs the target speech extraction result corresponding to the current data block after processing the current data block. At the same time, an updated historical cache state is generated. and the historical cache state The processing of passing the data to the next data block involves splicing the target speech extraction results of consecutive data blocks to form a continuously output target speech signal. The target speech signal can be directly used by speech playback, speech communication, speech recognition, speech transcription or other subsequent processing modules, as well as speech recognition front-end enhancement, instruction recognition, playback, storage or forwarding.

[0062] The streaming target speech extraction network calls the historical cache state retained from the previous data block when processing the current data block, and updates the historical cache state after outputting the target speech extraction result corresponding to the current data block; the historical cache state includes at least one of encoder cache state, convolutional layer cache state or feature memory state.

[0063] In this embodiment, the inference device is an edge computing device, embedded processing device, or deep learning inference board with neural network inference capabilities, used to run the target speech extraction model and output the target speech.

[0064] like Figure 6 As shown, this embodiment also provides a system deployment method suitable for multi-person collaborative communication scenarios.

[0065] The system includes multiple wearable acquisition terminals, which are worn by different users. Each wearable acquisition terminal acquires the electroglottic signal and speech signal of its corresponding wearer and transmits them to the same inference device or multiple inference devices via a wireless link. The inference device performs preprocessing, feature extraction and target speech extraction on the data sent by each terminal according to the received terminal identifier, thereby outputting the target speech of the corresponding wearer.

[0066] In this embodiment, although multiple people may speak at the same time in the work scenario, each wearable acquisition terminal corresponds to one wearer, and the electroglottic signal acquired by each terminal only corresponds to the vocal activity of its respective wearer. Therefore, the inference device can use the electroglottic features corresponding to each terminal to conditionally extract the speech of different wearers, thereby improving the clarity and stability of the respective voice links in the multi-person collaborative communication scenario.

[0067] Furthermore, in multi-person collaborative deployment, the wireless transmission module can attach a unique identifier to the data of different wearable acquisition terminals. The inference device establishes a data receiving buffer according to the identifier and executes the corresponding target voice extraction process. Thus, the present invention is not only applicable to single-person wearable voice enhancement scenarios, but also to target voice communication scenarios in multi-person teams, multi-person collaboration, and complex working environments.

[0068] like Figure 3 and Figure 5 As shown in the figure, this embodiment further provides an engineering implementation method.

[0069] The front-end wearable acquisition terminal is responsible for the real-time acquisition, synchronous processing and wireless transmission of electroglottic signals and speech signals, while the back-end inference device is responsible for auxiliary representation generation, streaming target speech extraction and continuous speech output.

[0070] Among them, such as Figure 3 As shown, the wearable acquisition terminal includes an electrode assembly, a microphone assembly, a signal conditioning and sampling assembly, a control assembly, a wireless communication assembly, and a power module; the electrode assembly is used to acquire electroglottic signals, the microphone assembly is used to acquire speech signals, and the control assembly is used to perform synchronous sampling control, data buffering, and wireless transmission control.

[0071] like Figure 5 As shown, the target speech extraction model includes an electroglottogram preprocessing module, an auxiliary representation network, a streaming target speech extraction network, and a historical cache state management unit. The auxiliary representation network is used to generate target speaker representation information based on the electroglottogram features of the current data block. The streaming target speech extraction network is used to receive the speech features of the current data block, the target speaker representation information, and the historical cache state corresponding to the previous data block, and output the target speech extraction result corresponding to the current data block. The historical cache state management unit is used to update and save the historical cache state for use in the next data block.

[0072] To meet the requirements of real-time applications, the front-end wearable acquisition terminal can adopt a low-power, miniaturized hardware architecture, and the back-end inference device can adopt an edge inference deployment method. With the above structure, the present invention can output the target speech result block by block for continuously arriving speech data blocks without waiting for the entire speech segment to end, thereby meeting the application requirements of low-latency voice communication scenarios.

[0073] To verify the effectiveness of the method described in this invention in extracting target speech in complex environments, tests were conducted on mixed speech input and the output of the method of this invention.

[0074] Test results are as follows Figure 7 As shown, the method of the present invention is superior to the mixed speech input state in terms of target speech clarity and interference suppression effect, thereby verifying the effectiveness of electroglottic auxiliary representation in target speech extraction. That is, the present application achieves target speech extraction by replacing additional training speech with electroglottic auxiliary representation, so that the present application does not require users to provide additional training speech at all compared with the prior art.

[0075] The above embodiments are preferred implementations of the present invention. In addition, the present invention can be implemented in other ways. Any obvious substitutions without departing from the concept of the present technical solution are within the protection scope of the present invention.

Claims

1. A wearable target speech extraction method based on the fusion of electroglottic diagrams and speech signals, characterized in that, Includes the following steps: S1: The electroglottic signal and speech signal of the target user during the vocalization process are simultaneously acquired by the acquisition terminal worn around the user's neck. The electroglottic signal is used to characterize the change in the target user's laryngeal impedance or the state of vocal cord vibration. The speech signal is the original audio signal containing the target user's speech acquired by the microphone. S2: Sample, buffer, time-align, and encapsulate the electroglot signal and the speech signal to generate dual-modal input data to be transmitted; S3: The dual-modal input data is sent to the inference device via the wireless transmission module; S4: The inference device preprocesses the electroglottic signal and the speech signal respectively to obtain electroglottic features and speech features; S5: Input the electroglottic features into the auxiliary representation network to generate target speaker representation information, and fuse the target speaker representation information with the speech features; S6: Input the speech features of the current data block, the target speaker representation information, and the historical cache state into the streaming target speech extraction network to obtain the target speech extraction result corresponding to the current data block; S7: Output the target speech extraction result corresponding to the current data block, and update the historical cache state; S8: After splicing the target speech extraction results of continuous data blocks, a continuous output target speech signal is formed.

2. The wearable target speech extraction method based on electroglottic diagram and speech signal fusion according to claim 1, characterized in that, The acquisition terminal in step S1 is a wearable acquisition device that can be attached to or surround the user's neck, and the electroglottic signal is acquired by electrode components placed on both sides of the user's throat.

3. The wearable target speech extraction method based on electroglottic diagram and speech signal fusion according to claim 1, characterized in that, In step S2, time alignment of the electroglot diagram signal and the speech signal includes: synchronous sampling based on a unified clock trigger, or alignment processing of the two sampled signals based on a timestamp.

4. The wearable target speech extraction method based on electroglottic diagram and speech signal fusion according to claim 1, characterized in that, The wireless transmission module in step S3 is a Wi-Fi module, a Bluetooth module, or other short-range wireless communication module.

5. The wearable target speech extraction method based on electroglottic diagram and speech signal fusion according to claim 1, characterized in that, In step S4, the preprocessing of the electroglot diagram signal includes one or more of the following: demodulation, filtering, envelope extraction, normalization, framing, time-frequency transformation, and fundamental frequency correlation feature extraction.

6. The wearable target speech extraction method based on electroglottic diagram and speech signal fusion according to claim 1, characterized in that, In step S5, the auxiliary representation network generates a target speaker representation vector based on the electroglottic features acquired synchronously with the speech signal. The target speaker representation vector is used to replace or partially replace the speaker prior information constructed based on pre-recorded reference speech or adaptive speech in traditional target speech extraction methods. In step S5, the target speaker representation information is used to conditionally modulate at least one feature extraction layer in the streaming target speech extraction network to enhance the speech components of the target user and suppress interfering speech and noise components. The conditional modulation includes generating scaling and bias factors based on the target speaker representation information, and adaptively scaling and biasing the speech features of the feature extraction layer.

7. The wearable target speech extraction method based on electroglottic diagram and speech signal fusion according to claim 1, characterized in that, The electroglot signal and the speech signal are input into the streaming target speech extraction network in the form of time-aligned data blocks. When processing the current data block, the streaming target speech extraction network outputs the target speech extraction result corresponding to the current data block. The streaming target speech extraction network calls the historical cache state retained in the previous data block when processing the current data block, and updates the historical cache state after outputting the target speech extraction result corresponding to the current data block; the historical cache state includes at least one of encoder cache state, convolutional layer cache state or feature memory state.

8. The wearable target speech extraction method based on electroglottic diagram and speech signal fusion according to claim 1, characterized in that, The inference device is an edge computing device, embedded processing device, or deep learning inference board with neural network inference capabilities, used to run the target speech extraction model and output the target speech.

9. A wearable target speech extraction system based on electroglottic diagram and speech signal fusion, used to implement the wearable target speech extraction method based on electroglottic diagram and speech signal fusion as described in any one of claims 1-8, characterized in that, include: Wearable acquisition terminal, used to be worn around the user's neck, to simultaneously acquire electroglottic signals and voice signals; The synchronization processing module is used to sample, buffer, time-align, and encapsulate the acquired electroglottic and speech signals. The wireless transmission module is used to send the packaged dual-modal input data to the inference device; The inference device is used to receive the dual-modal input data, extract features from the electroglottic signal and the speech signal, and generate target speaker representation information based on the electroglottic features through an auxiliary representation network. The output module is used to output, play, store, or send the target voice. The wearable acquisition terminal includes an electrode assembly, a microphone assembly, a signal conditioning and sampling assembly, a control assembly, a wireless communication assembly, and a power module; the electrode assembly is used to acquire electroglottic signals, the microphone assembly is used to acquire speech signals, and the control assembly is used to perform synchronous sampling control, data buffering, and wireless transmission control. The inference device includes a preprocessing module, an auxiliary representation module, a streaming fusion extraction module, a cache state management module, and a result output module. The auxiliary representation module is used to generate target speaker representation information based on electroglottic features. The streaming fusion extraction module is used to extract the target speech corresponding to the current data block based on the speech features of the current data block, the target speaker representation information, and the historical cache state. The cache state management module is used to update and save the historical cache state. The system includes multiple wearable acquisition terminals, which are worn by different users to acquire the electroglottic signal and speech signal of the corresponding user, and output the target speech of the corresponding user.

10. An electronic device, characterized in that, The device includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the electronic device to perform the wearable target speech extraction method based on electroglottic diagram and speech signal fusion as described in any one of claims 1 to 8.